Intelligent domain decision-making method and system based on knowledge base

By building a multi-source heterogeneous data fusion knowledge base and a lightweight large language model in the field of equipment operation and maintenance, combined with an improved anomaly detection mechanism, the problems of low data utilization and low efficiency of intelligent decision-making in the field of equipment operation and maintenance are solved, and efficient and accurate intelligent decision-making support is achieved.

CN120611795APending Publication Date: 2025-09-09BEIJING UNIV OF POSTS & TELECOMM
View PDF 0 Cites 2 Cited by

Patent Information

Application Number
CN202510737754.3
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-06-04
Publication Date
2025-09-09

AI Technical Summary

Technical Problem

Existing technologies in the field of equipment operation and maintenance have problems such as low utilization of multi-source heterogeneous data, insufficient data quality, high model complexity, large hardware resource requirements, difficulty in knowledge updating, and insufficient anomaly detection capabilities, resulting in low efficiency and accuracy of intelligent decision-making.

Method used

Build a maintenance knowledge base that integrates multi-source heterogeneous data, and achieve efficient intelligent decision support through a lightweight vertical field large language model and an improved inspection image fault point anomaly detection mechanism, combined with a multi-level retrieval-enhanced text maintenance strategy generation algorithm.

Benefits of technology

It improves the efficiency and accuracy of intelligent decision-making in the field of equipment operation and maintenance, reduces hardware resource requirements, improves the utilization of multi-source heterogeneous data, and enhances anomaly detection capabilities, ensuring timely knowledge updating and high-quality text maintenance strategy generation.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120611795A_ABST
    Figure CN120611795A_ABST
Patent Text Reader

Abstract

The invention discloses a knowledge base-based field intelligent decision-making method and system, and belongs to the field of equipment operation and inspection. The invention is oriented to the field of equipment operation and inspection, combines advanced technologies and ideas such as layout analysis, named entity recognition, text generation, large language model fine tuning, text classification and target detection, and constructs a knowledge base-based field intelligent decision method and system for efficient processing and decision support of equipment fault information. The method comprises the following steps: constructing and organizing multi-source heterogeneous data, and integrating high-quality information into a maintenance knowledge base; on the basis, generating a self-adaptive text maintenance strategy by utilizing domain knowledge and responding to a dynamic operation context; and finally, fusing image modal data through an anomaly detection method, accurately positioning a fault, and improving the generation quality of a text maintenance strategy. According to the method, a complete process from multi-source heterogeneous data to multi-dimensional strategy generation is realized, key problems in equipment maintenance are effectively solved, and efficiency and accuracy are remarkably improved.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention belongs to the technical field of equipment operation and maintenance, and particularly relates to a domain intelligent decision-making method and system based on a knowledge base. Background Art

[0002] With the continuous expansion of business, the number of specialized equipment is increasing, and the cost of daily operation and maintenance is increasing. Currently, maintenance strategies for specialized equipment rely primarily on manual experience, requiring professionals to conduct on-site inspections of equipment conditions and consult relevant standard documents before making decisions. As the concept of "AI+" continues to mature, intelligent decision-making based on massive amounts of multi-source, heterogeneous data has become a major trend in the field of equipment operation and maintenance. The reliance on manual labor for maintenance strategies in specialized fields is related to a lack of relevant knowledge and experience, the difficulty in extracting and retrieving effective information, and the difficulty in reasoning based on comprehensive analysis of equipment status to flexibly adapt multiple strategies to specific scenarios. This presents an opportunity for the application of large language models in specialized fields. In this process, the knowledge and experience for maintenance strategies are derived from massive amounts of multi-source data. However, this data on specialized maintenance strategies has been accumulated over the years but has been underutilized. Most industry maintenance strategy-related documents, such as equipment defect, failure, and maintenance records and historical cases, are typically written by different professionals. This results in long, redundant sections, diverse writing standards, and various storage formats, making them time-consuming and laborious for later generations to learn or consult. Furthermore, each domain document covers a wide range of factors, making it difficult to quickly migrate historical plans to specific maintenance scenarios. Therefore, how to extract information from unstructured, multi-source, heterogeneous data and construct a dense, concise maintenance knowledge base. Based on this knowledge base, large language models can be used to build vertical domain knowledge reasoning capabilities to achieve reliable intelligent decision-making while meeting lightweight requirements. This is the focus and difficulty of current research.

[0003] In recent years, deep learning technology has continued to develop. Large language models, characterized by "big data + large computing power + strong computing power", have performed well in various NLP tasks in general fields, and have played a positive role in promoting the development of industries with rich data accumulation. In the field of equipment operation and maintenance, the demand for using large language models for industry assistance is constantly expanding, which also puts forward new requirements for the quantity and quality of high-quality data. At present, the development level of document layout analysis technology has reached the level of being able to realize the unified extraction task of multi-type unstructured document information. At the same time, the application of target detection technology in image data related processing is also relatively mature, providing research directions for the integration and utilization of massive and diversified multi-source heterogeneous data in the industry.

[0004] The existing intelligent decision-making method closest to the present invention and its shortcomings are as follows:

[0005] (1) Improve BERT’s intelligent fault case matching method;

[0006] Yang Yi et al. (Yang Yi, Cui Qihui, Qin Jiafeng, Zheng Wenjie, Qiao Mu, Improved BERT-based Intelligent Matching Method for Fault Cases, Shandong Electric Power Technology, February 25, 2022) used the improved BERT model to apply it to intelligent matching of power grid fault cases to achieve intelligent decision-making. The method includes: first, starting from the power grid fault case data, analyzing the case characteristics, and extracting key information, such as the case process and case analysis; then, pre-training and fine-tuning based on the BERT model, combining the long short-term memory network (LSTM) to build a matching model to achieve deep semantic matching of power grid fault cases, and combining the conditional random field (CRF) model to further identify and classify the dispatch object entities in the power grid fault cases, thereby improving the model's fine-grained classification capabilities. By improving the BERT model, the accuracy and efficiency of power grid fault case matching can be significantly improved. However, the disadvantages of this method are high model complexity, the training process requires a large amount of computing resources and high-quality data, and limited by the capabilities of its classification model itself, the method cannot generate a complete text maintenance strategy with high readability.

[0007] (2) A power grid fault response plan matching method based on semantic enhancement;

[0008] Meng Fei et al. (Meng Fei, Li Jiangpeng, Li Tao, Xu Jianzhong, Gao Haiyang, and Qiao Yongtian, "A Semantically Enhanced Power Grid Fault Handling Plan Matching Method," China Electric Power, November 7, 2024) used a semantic enhancement model to adjust the hyperparameters of the BERT (Bidirectional Encoder Representations from Transformers) model, representing power grid fault handling plan text as computable word vectors. They then introduced a conditional random field (CRF) model to identify the entity categories of dispatch objects and calculated the semantic distance between power grid fault information and dispatch objects based on a residual vector-word embedding vector-encoding vector (RE2) model. The RE2 model improves its ability to capture semantic information by combining multiple layers of vector representations (residual vectors, word embedding vectors, and encoding vectors). While this method can improve the efficiency and accuracy of power grid fault handling plan matching, it requires high initial data quality, and data accuracy directly affects the model's performance and final results. Furthermore, its decision-making results can only match existing power grid fault handling plans and cannot independently organize comprehensive maintenance strategies for real-world scenarios.

[0009] Existing intelligent decision-making methods have the following main defects:

[0010] (1) The volume and quality of domain knowledge data are high;

[0011] When processing and analyzing text data related to equipment operation and maintenance, we face the challenge of accumulating large amounts of data due to its diverse nature. This data is organized in a variety of forms, including but not limited to accident reports, maintenance logs, operating manuals, technical specifications, and data records from various organizational systems. The diversity and complexity of this data requires more efficient and accurate methods to process and integrate it.

[0012] (2) Existing large language models have high requirements for data resources, hardware resources, and generation quality;

[0013] Most existing large language models possess a massive number of parameters. When addressing domain-specific tasks, they require model training to acquire a certain amount of domain knowledge. Furthermore, the quality of content generated by text maintenance strategies must maximize accuracy based on a given set of domain documents. For complex models with large parameter counts, such as large language models, training requires high data and hardware resources, making it difficult to directly modify inherent knowledge for a specific domain or to inject specific domain knowledge into the underlying large language model. While existing retrieval enhancement (RAG) methods linking to external maintenance knowledge bases are used to supplement model knowledge, while allowing the model to reference relevant documents when answering questions, they fail to effectively utilize known test documents during training. While the lightweight DeepSeek model is available for individual deployment, its knowledge utilization model is relatively primitive and its knowledge utilization capabilities are limited. Specifically, only a small number of documents can be uploaded per generation, making it incapable of supporting searches of large domain maintenance knowledge bases. Furthermore, the quality of the source evidence and content generated by text maintenance strategies is difficult to guarantee.

[0014] (3) The general domain large language model has hallucinations and lacks timely knowledge;

[0015] Large language models have achieved remarkable results in various natural language processing tasks in general domains. However, due to their limited domain knowledge, large language models can produce hallucinatory responses to facts related to equipment maintenance during their construction, which poses potential risks to their practical application in this field. Regularly training large domain models with new data to update domain knowledge within the model is extremely costly and prone to catastrophic forgetting. Furthermore, when learning new domain knowledge, large models suffer from severe memory loss, where old knowledge is replaced by new information. This makes continuous learning of large models extremely difficult.

[0016] (4) The utilization rate of multi-source heterogeneous data is low;

[0017] With the rapid development of specialized fields, the scale of infrastructure and the complexity of its operations and maintenance (OM) continue to grow. Traditional maintenance decision-making relies heavily on manual experience and static documentation, requiring personnel to conduct on-site inspections, consult standard guidelines, and analyze historical records. This approach is inefficient and highly subjective, making it difficult to cope with the complex and ever-changing challenges of real-world scenarios. At the same time, the field of equipment operation and maintenance has accumulated a large amount of multi-source heterogeneous data, including equipment fault logs, maintenance records, and on-site equipment images. The latter often serve as an important basis for maintenance strategies in case reports. Although these datasets contain rich information and offer enormous potential for knowledge mining, publicly available datasets for defective equipment target detection in the field are relatively insufficient. In particular, the limited number of defect categories involved and the substantial amount of raw image data are underutilized. The utilization of this multi-source heterogeneous data is far from ideal.

[0018] (5) Insufficient abnormality detection capabilities in the field of equipment operation and maintenance: The recognition accuracy of small targets at long distances commonly seen in inspection sites is low and the sensitivity is poor. Summary of the Invention

[0019] In order to realize "how to efficiently integrate multi-source heterogeneous data resources in the field of equipment operation and maintenance into a knowledge-driven reasoning framework to build an intelligent decision-making mechanism for complex scenarios and dynamic tasks?" the present invention provides a domain intelligent decision-making method and system based on a knowledge base.

[0020] The present invention focuses on multi-source heterogeneous data fusion and knowledge-driven dynamic adaptation. It systematically solves the limitations of existing technologies by constructing a high-density maintenance knowledge base, developing a multi-level retrieval-enhanced text maintenance strategy generation algorithm, and designing an improved inspection image fault point anomaly detection mechanism. First, by constructing and organizing multi-source heterogeneous data, high-quality information is integrated into the maintenance knowledge base. On this basis, adaptive text maintenance strategies are generated by utilizing domain knowledge and responding to dynamic operational contexts. Finally, image modal data is fused through anomaly detection methods to accurately locate faults and improve the quality of text maintenance strategy generation. The present invention realizes a complete process from multi-source heterogeneous data to multi-dimensional strategy generation, effectively solves key problems in equipment maintenance, and significantly improves efficiency and accuracy.

[0021] The technical solutions adopted by the present invention to solve the technical problems are as follows:

[0022] The domain intelligent decision-making method based on the knowledge base provided by the present invention mainly includes the following steps:

[0023] Step 1: Build a maintenance knowledge base for equipment operation and maintenance;

[0024] After data preprocessing, domain documents are first analyzed for layout using PP-PicoDet, followed by text recognition using the RARE model to complete layout analysis. Based on the layout analysis results, the PMRC model is used to transform the machine reading comprehension framework into template prompts to achieve named entity recognition. The MASKER model is used to classify text through contextual mask keyword reconstruction, and a multi-task instruction dataset is constructed.

[0025] Step 2: Generate a vertical field text inspection strategy based on a general large language model;

[0026] S2.1: Method for constructing lightweight large language models in vertical fields;

[0027] Collect and clean domain device text data, use a lightweight large language model as the backbone model, and perform modeling in an autoregressive manner after secondary pre-training to obtain a lightweight vertical domain large language model.

[0028] S2.2: instruction fine-tuning;

[0029] Construct a dataset in a multi-task instruction format and perform supervised training on a pre-trained lightweight vertical domain large language model to achieve instruction fine-tuning.

[0030] S2.3: RAT-based large language model maintenance strategy generation algorithm;

[0031] Using the RAT method and a fine-grained training strategy, a lightweight vertical domain language model is trained by introducing interference domain documents. Standard supervised learning techniques are then used to fine-tune the lightweight vertical domain language model. The lightweight vertical domain language model retrieves relevant knowledge records from the maintenance knowledge base based on the defect description text, and fills the corresponding prompt templates with multiple knowledge records and the original defect description text according to different stages. After multiple rounds of integration and adjustment, the maintenance strategy is obtained.

[0032] Step 3: Abnormal detection method for equipment defects at the operation and maintenance site;

[0033] The fault point is identified by using the faulty equipment anomaly detection method based on the improved YOLOv7 model using the ACmix attention mechanism.

[0034] Furthermore, the proposed PP-PicoDet learns parameterized hierarchical structure details based on the CSP-PAN feature fusion network and uses SimOTA as the label assignment algorithm to detect various components in domain documents.

[0035] Furthermore, the RARE model uses a spatial transformation network to correct text distortion and recognizes sequences through a sequence recognition network, thereby accurately reading input content.

[0036] Furthermore, the backbone model is the Baichuan2-13B model; in the Baichuan2-13B model, the second multi-head attention layer is deleted from the decoder structure of its Transformer, and the first Masked multi-head attention layer is retained, that is, the ability to mask the following information during model training is retained.

[0037] Furthermore, the training process of the Baichuan2-13B model is as follows:

[0038] (1) The words in the input text are vectorized by using the Transformer method of adding the word encoding and position encoding of the input text: first, the input text is uniquely encoded in units of words, each word is mapped into a unique vectorized representation, and then the word encoding mapping matrix W is used. e Mapping is performed to obtain the corresponding word encoding vector; for position encoding, the position encoding mapping matrix W is used p Mapping is performed, and after initializing the position encoding mapping matrix, the model itself is used for learning. The calculation formula is:

[0039] h0=UW e +W p #(2)

[0040] Where h0 represents the initial representation vector of each word in the input sequence, and U represents the learnable weight matrix used to map word embeddings to the hidden layer dimensions of the model;

[0041] Then h0 is passed into each block step by step to get the output h of the last block. n , h n It indicates the probability of predicting the next word from each word in the input text. The calculation formula is:

[0042]

[0043] Among them, h l Representation vector of the model output at layer l, h l-1 Represents the representation vector output by the model at layer l-1, n represents the number of block layers in the network structure, T represents the matrix transpose, P(u) represents the probability distribution of the model for each word in the vocabulary, and softmax represents the activation function;

[0044] (2) Use a large amount of unlabeled data for pre-training and perform fine-tuning on specific tasks on labeled data;

[0045] (3) For the task of generating fault repair strategies in the field of equipment operation and maintenance, the defect description in the field of equipment operation and maintenance is used as the prompt content to predict the repair strategy text: Assuming that T is the defect description text representation vector matrix, its prediction formula is:

[0046]

[0047] Among them, X represents the generated result, x k:n represents the generated word representation vector from position k to n, t 1:k-1 represents the hint word representation vector from position 1 to k-1, t i Represents the defect description text word representation vector.

[0048] Furthermore, the data set in the multi-task instruction format includes a common sense instruction data set and a maintenance strategy instruction data set; the maintenance strategy instruction data set is used to enhance the lightweight vertical domain large language model's ability to identify questions related to faulty equipment maintenance in the equipment operation and maintenance field; the common sense instruction data set is used to synchronously execute instruction fine-tuning, so that the lightweight vertical domain large language model can utilize the domain knowledge it has learned during instruction fine-tuning.

[0049] Furthermore, the RAT method introduces interference domain documents on the basis of RAG to train the model to ignore irrelevant information, while generating high-quality answers that contain chain thinking and direct references; collects a dataset containing questions, relevant documents and answers, and divides the dataset into two parts: one part contains documents with correct answers, and the other part contains interference documents; for each question in the dataset, constructs a training sample, including the question, a set of domain documents and an answer with a chain thinking style.

[0050] Furthermore, in the RAT method, through a step-by-step reasoning mode, linked thinking prompts encourage the model to perform multi-step reasoning when answering questions: first, the steps are analyzed from a single round of answers, and then the corresponding maintenance strategies are further retrieved based on the content or reasons of the analysis, and the existing answers are revised according to the specific content, so as to realize the generation and integration of multi-scenario inspection and maintenance strategies.

[0051] Furthermore, the ACmix attention mechanism module in the improved YOLOv7 model based on the ACmix attention mechanism first performs feature segmentation to form sub-features to capture feature information at different scales; the upper branch convolution path and the lower branch self-attention path calculation are performed separately; in the upper branch convolution path calculation, the sub-features are shifted and aggregated after passing through the fully connected layer, and the obtained features are further convolved to output H×W×C features; in the lower branch self-attention path calculation, the self-attention mechanism is used to calculate the similarity between features to enhance the expressiveness of the features; finally, feature fusion is performed to merge the output results of the upper branch convolution path and the lower branch self-attention path calculation. The strength of the merging is controlled by two learnable scalars α and β. The final output feature is:

[0052] F out =αF att +βF conv #(17)

[0053] Among them, F out Represents the final output feature map obtained by weighting, F att represents the feature map calculated by the attention mechanism, F conv Represents the feature map obtained by convolution operation.

[0054] The domain intelligent decision-making system based on the knowledge base provided by the present invention includes:

[0055] The front-end display layer adopts a responsive design and provides users with an intuitive operation interface through the information presentation module and the user operation module. The user operation module has user input, query, and data screening functions, and supports user-defined queries. The information presentation module visualizes the information query data and system analysis data of the intelligent decision-making process and generates intelligent decision results.

[0056] The service layer is used to receive user requests and route them to the corresponding services. It also supports cross-system communication and protocol conversion through the API gateway, providing a consistent interface for different clients. The service layer adopts a microservice architecture, with each microservice unit completing a specific business function.

[0057] The business layer is used to achieve the deep integration and application of domain knowledge and intelligent decision-making related models and algorithms; the business layer includes a maintenance knowledge base data query module, an intelligent decision-making module, and an anomaly detection module;

[0058] The data layer includes a maintenance knowledge base and a big data storage module; the maintenance knowledge base contains high-quality domain knowledge with high information density, supporting intelligent decision-making reasoning and rule matching; the big data storage module stores historical case data, operating procedures and domain-related materials, providing data support for the intelligent decision-making system.

[0059] The beneficial effects of the present invention are:

[0060] The present invention is aimed at the field of equipment operation and maintenance, and combines advanced technologies and ideas such as layout analysis, named entity recognition, text generation, large language model fine-tuning, text classification and target detection to construct a knowledge base-based domain intelligent decision-making method and system for efficient processing and decision support of equipment fault information. The present invention improves the level of intelligent decision-making in the field of equipment operation and maintenance from multiple levels, including building a high-quality domain maintenance knowledge base as a data source to assist in building a lightweight vertical domain large language model for reliable decision-making, and performing anomaly detection on equipment fault maintenance image information from a visual level. The overall intelligent decision-making process makes full use of the professional knowledge and historical maintenance practice experience in the field of equipment operation and maintenance to form a multi-dimensional intelligent decision-making generation mechanism. Through intelligent data analysis and processing, the present invention not only improves the professionalism and comprehensiveness of equipment fault maintenance, but also provides a strong guarantee for the safe and stable operation of equipment. In general, the present invention has practical significance in promoting intelligent equipment management and improving fault handling efficiency, and reflects the contribution and level of technological innovation and development in the field.

[0061] Compared with the prior art, the present invention has the following advantages:

[0062] (1) Integration of high-quality domain knowledge: The present invention constructs a maintenance knowledge base in the field of equipment operation and maintenance through layout analysis, text matching, named entity recognition and other technologies, solving the problem of lack of available data and low quality of data sources in the field of equipment operation and maintenance.

[0063] (2) Resource requirement optimization: The present invention uses technologies such as model quantization and model fine-tuning to significantly reduce the demand for data resources and hardware resources, allowing the model to be deployed on consumer-grade graphics cards to perform intelligent decision-making tasks for device failures, thereby reducing costs.

[0064] (3) The general domain large language model has hallucinations and lacks timely knowledge;

[0065] In order to improve the ability of large language models to understand and express equipment operation and maintenance-related knowledge in specific scenarios in the field of equipment operation and maintenance, so that the model can make intelligent decisions and reasoning based on multiple dimensions, the present invention constructs a variety of equipment operation and maintenance maintenance data sets. The data sets contain operation and maintenance knowledge from different sources, such as basic equipment concepts, daily operation and maintenance guidelines, and maintenance precautions. They are mainly used to assist the model in domain adaptation tuning, combined with self-supervision or supervised multi-stage tuning methods to achieve the effect of knowledge injection.

[0066] (4) Multi-level retrieval-enhanced text review strategy generation algorithm realizes reliable text review strategy generation: The present invention combines thinking chain, retrieval enhancement and other technologies to propose a multi-level retrieval-enhanced text review strategy generation algorithm, which improves the original knowledge utilization model, enhances the quality of text review strategy generation content, and realizes reliable intelligent assistance.

[0067] (5) Improved utilization of multi-source heterogeneous data: While providing textual maintenance strategies, the present invention can also identify on-site equipment images, providing intelligent decision-making assistance from a visual perspective, thereby improving the utilization efficiency of multi-source heterogeneous data and enhancing the accuracy and reliability of overall decision-making.

[0068] (6) Improved anomaly detection capabilities in the field of equipment operation and maintenance: In order to make full use of multi-source heterogeneous data for intelligent decision-making, the present invention proposes an anomaly detection method for equipment defects on the operation and maintenance site based on deep learning methods and attention mechanisms, which improves the sensitivity to small targets at long distances. After the system gives a text maintenance strategy for equipment abnormalities, it can mark the on-site equipment pictures, and further perform more accurate fault detection and intelligent decision-making from a visual perspective. BRIEF DESCRIPTION OF THE DRAWINGS

[0069] Figure 1 The present invention provides a domain intelligent decision-making method process based on a knowledge base.

[0070] Figure 2 Build a process for maintenance knowledge base in the field of equipment operation and maintenance.

[0071] Figure 3 It is the knowledge utilization mechanism (RAT).

[0072] Figure 4 A method flow for generating vertical domain text inspection strategies based on a general large language model.

[0073] Figure 5 To improve the YOLOv7 model structure.

[0074] Figure 6 To optimize the ELAN-H module structure. DETAILED DESCRIPTION

[0075] The present invention is further described in detail below with reference to the accompanying drawings.

[0076] In a first aspect, the present invention provides a domain intelligent decision-making method based on a knowledge base.

[0077] like Figure 1 As shown, the domain intelligent decision-making method based on the knowledge base of the present invention specifically includes the following steps:

[0078] Step 1: Construction method of maintenance knowledge base for equipment operation and maintenance;

[0079] (1) Overall design ideas for the maintenance knowledge base construction plan;

[0080] The present invention designs a method for constructing a maintenance knowledge base for the field of equipment operation and maintenance. The method extracts and condenses key information from unstructured data into structured knowledge records containing multidimensional attributes and stores them in a database. In order to extract relatively rich attribute information, the present invention comprehensively utilizes multiple natural language processing technologies and rationally applies them to different types of unstructured data features, thereby obtaining structured knowledge records containing multidimensional attributes. Among them, for long-sentence text attributes such as judgment basis and maintenance strategy, which have relatively obvious title prompts and fixed formats in the original text, the present invention uses layout analysis technology to extract them and use them as the information source for constructing a structured knowledge record; for vocabulary-type attributes such as faulty components and related status quantities, the present invention uses a named entity recognition model to extract them; for attribute information such as maintenance category and maintenance time, which are missing in the original text but are relatively critical in the maintenance process, the present invention uses text classification technology to generate them, further improving the information content and value of each structured knowledge record. The knowledge records of the maintenance knowledge base thus constructed contain multidimensional attributes such as judgment basis, maintenance strategy, faulty components, status quantities, maintenance category, and maintenance time. At the same time, in the process of extracting information from the original text, the present invention also uses the extracted structured knowledge record data to construct a corresponding text pair dataset for subsequent model training, such as a concept question answering dataset and a maintenance strategy dataset.

[0081] The present invention designs a maintenance knowledge base construction method for the equipment operation and maintenance field, and its specific implementation process mainly includes: data preprocessing, layout analysis, named entity recognition and text classification, such as Figure 2 shown.

[0082] S1.1: Layout analysis;

[0083] The original corpus mainly includes: case reports, operating procedures and domain-related materials. The present invention first uses PP-PicoDet to perform document layout analysis. PP-PicoDet is based on the CSP-PAN feature fusion network, learns parameterized hierarchical structure details, and uses SimOTA as the label assignment algorithm, so that it can accurately detect various components in the document, such as tables and title rows. In addition, the present invention uses the RARE model for text recognition. The RARE model plays a complementary role in this process because documents usually contain irregular and distorted text. The RARE model uses a spatial transformer network (STN) to correct text distortion and recognizes sequences through a sequence recognition network (SRN), enabling it to accurately read difficult-to-recognize input content, including handwritten marks or text on curves in scanned documents. Through the combined use of PP-PicoDet and the RARE model, the present invention can accurately extract the required key information from complex documents, which not only ensures the accurate extraction of long-sentence text attribute components such as judgment basis and maintenance strategy, but also lays an important foundation for downstream information extraction tasks. This invention leverages the collaborative work of PP-PicoDet and the RARE model to not only improve text recognition accuracy but also preserve the original structure of the document, making the extracted key information more organized and logically clear. This improves the accuracy of information extraction, reduces manual intervention, and improves work efficiency. Layout analysis can reveal faulty equipment, domain-specific terminology, and defect descriptions / repair strategies.

[0084] S1.2: Named Entity Recognition;

[0085] Based on the information extracted by the PP-PicoDet and RARE models, the present invention further uses a prompt-based machine reading comprehension PMRC model to extract key words related to defects, such as component names and status quantities. The PMRC model transforms the machine reading comprehension framework into an innovative entity recognition method implemented through template prompts, which is used to find terms such as equipment status and name. By designing special templates, it can complete knowledge extraction in the field of equipment operation and maintenance in the absence of large-scale data. In the field of equipment operation and maintenance, this low-resource flexibility is particularly powerful because professional terms and specialized vocabulary are almost ubiquitous in this field, while low-level output data sets are very scarce. Through a template-based prompt approach, the PMRC model can systematically adapt to new text data dimensions with minimal retraining cost, thereby ensuring the accuracy of key named entity recognition.

[0086] To address the current challenges of limited existing datasets and the difficulty of manual data annotation in the field of equipment operation and maintenance, this paper employs a defect keyword extraction method tailored for low-resource scenarios. This method, through a designed template, extracts key information from this field. For resource-limited equipment operation and maintenance scenarios, this method eliminates the need for efficient information extraction methods based on massive amounts of data and enables relatively accurate key information extraction under low-resource conditions, demonstrating strong practicality and potential for widespread adoption.

[0087] S1.3: Text classification;

[0088] This paper uses the MASKER model to classify text through contextual mask keyword reconstruction to obtain its maintenance category and maintenance time. The MASKER model uses the complete context to predict defect severity, allowing it to adapt to different syntactic structures and different grammars in more detail, improving the accuracy of text classification.

[0089] S1.4: Construct a multi-task instruction dataset, including a concept question answering dataset (which can be composed of domain-specific term explanations) and a repair strategy dataset (which can be composed of defect descriptions / repair strategies).

[0090] (2) The scale of the maintenance knowledge base and related datasets;

[0091] The maintenance knowledge base constructed for the equipment operation and maintenance field is stored in a relational database. Details of the maintenance knowledge base's storage structure and attributes are shown in Table 1. The maintenance knowledge base constructed for the equipment operation and maintenance field includes 1,933 records and 15,000 pieces of extracted high-quality information. The multi-task instruction dataset constructed by this invention includes a concept question-answer dataset with 4,251 text-question-answer pairs; the defect image dataset in the maintenance strategy dataset contains 8,840 examples.

[0092] Table 1

[0093] property describe device_name Faulty device name component_name Defective part name state_parameter Description of state quantities involved judge_reference Judgment basis diagnostic_analysis Diagnose and analyze possible causes of faults maintenance_strategy Maintenance strategy maintenance_level Maintenance category maintenance_time Maintenance time

[0094] Table 2

[0095]

[0096]

[0097] This invention implements a defect classification method tailored to domain context and constructs a maintenance knowledge base for equipment operation and maintenance based on unstructured power text data. This structured maintenance knowledge base not only improves information query and utilization efficiency but also provides a foundation for long-term knowledge accumulation in the field. Operations and maintenance personnel can accelerate the development of practical maintenance strategies based on historical data and analysis reports in the maintenance knowledge base, thereby improving the scientific nature and accuracy of decision-making.

[0098] Step 2: Generate a vertical domain text inspection strategy based on a general large language model;

[0099] (1) Overall design ideas for text maintenance strategy generation method;

[0100] This paper proposes a lightweight equipment operation and maintenance field-specific fine-tuning training strategy. In the pre-training stage, Baichuan2-13B is used as the backbone model of a lightweight vertical field large language model, and a large-scale original text dataset is used to inject basic domain knowledge into the model. By structuring the training instances into question-answer pairs in instruction format, the model is helped to generalize in unseen tasks and scenarios, ensuring that the model can accurately respond to a variety of maintenance-related queries while maintaining a lightweight computational burden.

[0101] Since the existing Retrieval-Augmented Generation (RAG) framework does not work well in noisy environments, this paper designs a novel Knowledge Exploitation Mechanism (RAT), such as Figure 3 As shown, the method aims to enhance the model's ability to generate powerful maintenance strategies in complex environments. Specifically, the method dynamically generates task-specific prompts to achieve multi-level retrieval enhancement. These task-specific prompts focus on fault cause diagnosis, troubleshooting suggestions, and maintenance strategy generation. At the same time, it combines the LangChain and LlamaIndex retrieval enhancement frameworks to achieve general large language model deployment and links with external large-scale information sources. These task-specific prompts are integrated into Chain-of-Thought (COT) reasoning to improve logical consistency and make it applicable to complex tasks. The output generated by COT is organized into a joint message style format together with the retrieved historical cases to achieve integration and optimization of decision generation. This integration enables the framework to dynamically combine retrieved knowledge with logical reasoning to create a powerful decision pipeline that strikes a balance between contextual completeness and domain-specific relevance.

[0102] (2) The vertical domain text inspection strategy generation method based on the universal large language model of the present invention;

[0103] The traditional vertical domain large language model construction process has high requirements for data resources, hardware resources and generation expertise, and is not suitable for the field of equipment operation and maintenance where the knowledge iteration and update speed is relatively fast. To address this problem, the present invention proposes a vertical domain two-stage fine-tuning method based on a lightweight large language model, and completes the construction of a multi-task instruction dataset of equipment domain knowledge. The present invention implements a lightweight fine-tuning method based on a large language model, especially in the field of equipment fault maintenance. By constructing a special common sense instruction dataset and a maintenance strategy instruction dataset, the model can significantly improve its performance at a relatively low computational cost through two-stage fine-tuning, thereby enhancing the model's adaptability and generalization ability in equipment fault handling. This fine-tuning method enables the model to quickly adapt and respond effectively when faced with new tasks or scenarios.

[0104] like Figure 4 As shown, the specific implementation process is as follows:

[0105] S2.1: Method for constructing lightweight large language models in vertical fields;

[0106] This paper first performs preliminary domain-specific optimization on the existing large language model used in general domains, and then uses a secondary pre-training approach to fine-tune the model in the equipment operation and maintenance domain and the Chinese language domain. The goal of this process is to enable the large language model to learn the basic structure of the language, grammatical rules, word meanings, and contextual relationships. In this way, the large language model can acquire a wide range of language knowledge in the field of equipment troubleshooting without specific task guidance, providing a foundation for downstream tasks. The specific implementation process is as follows:

[0107] S2.1.1: Data collection; In the data collection stage, the present invention collects domain-related original texts through crawler technology and pre-trains the field equipment text data used, which mainly includes seven types of texts: fault cases, standard documents, literature, encyclopedia entries, professional vocabulary, news, and inspection reports. Their quantity, source, and format are shown in Table 3.

[0108] Table 3 Statistics of unstructured text data of devices

[0109]

[0110]

[0111] S2.1.2: Data Processing: This method cleans unstructured text data from devices across various sources. This cleans the collected domain device text data, removing noise and redundant information, such as HTML tags, special characters, and sensitive information, to make the data more representative and consistent. Furthermore, data is segmented by setting the input window size, dividing the processed data into multiple small segments for input into the model to optimize computing resources and reduce the computational burden.

[0112] S2.1.3: Autoregressive training; This invention uses the lightweight large language model Baichuan2-13B as the backbone model for lightweight vertical domain large language models, performs secondary pre-training on a large amount of unstructured raw text data in the field of equipment operation and maintenance, and models the model through autoregression. The network structure of the Baichuan2-13B model adopts the decoder structure of the Transformer, but in order to adapt to the training method of the large language model, this invention makes some changes to the decoder structure of the Transformer: the second multi-head attention layer is deleted, and the first masked multi-head attention layer is retained, that is, the ability to mask the following information during model training is retained.

[0113] The following focuses on the training process of the Baichuan2-13B model. First, pre-training is performed using a large amount of unlabeled data to help the model understand human natural language. Then, fine-tuning for specific tasks is performed on labeled data. Self-supervised training is typically used for pre-training a language model, and its maximum likelihood objective function is as follows:

[0114]

[0115] Among them, p represents the conditional probability distribution of the next word given k elements, k represents the context window size, and u i Represents unlabeled training text, i represents the text word position, θ represents the model parameters used for conditional probability calculation, and the model parameters are updated using the gradient descent method.

[0116] Before model training, the words in the input text need to be vectorized. Here, the Transformer is used to perform word encoding and position encoding on the input text and then add them together. The specific approach is: first, the input text is one-hot encoded in units of words, each word is mapped into a unique vectorized representation, and then the word encoding mapping matrix W is used. e Mapping is performed to obtain the corresponding word encoding vector. The advantage of this is that it can solve the problem of excessive space waste after the one-hot encoding vector, and at the same time store the meaning of a single word in the process of vectorization. For position encoding, the position encoding mapping matrix W is also used p For mapping, the Transformer's sine and cosine calculation method is no longer used here. Instead, the position encoding mapping matrix is ​​initialized and then the model itself is used for learning. The specific calculation formula is as follows:

[0117] h0=UW e +W p #(2)

[0118] Here, h0 represents the initial representation vector of each word in the input sequence, and U represents the learnable weight matrix used to map the word embedding to the hidden layer dimension of the model. Then h0 is passed to each block step by step to obtain the output h of the last block. n , h n It represents the probability of predicting the next word from each word in the input text. The specific calculation formula is as follows:

[0119]

[0120] Among them, h l Representation vector of the model output at layer l, h l-1 Represents the representation vector output by the model at layer l-1, n represents the number of block layers in the network structure, T represents matrix transpose, P(u) represents the probability distribution of the model for each word in the vocabulary, and softmax represents the activation function.

[0121] In the present invention, the equipment operation and maintenance field fault case report used for model training contains both defect descriptions and text related to maintenance strategies. Therefore, the text maintenance strategy generation task can be used as a case report generation task for preliminary model fine-tuning, that is, the data label y is included in the input text data. At the same time, in view of the need to take into account the sufficient amount of training data, the present invention uses the same self-supervised training method as the language model pre-training for secondary pre-training to complete the model's domain adaptation to Chinese texts in the equipment operation and maintenance field.

[0122] The general forward pre-training approach for a one-way language model is to predict subsequent text given a sequence of length k as a prompt. For the task of generating fault repair strategies in the equipment maintenance domain, the defect description in the equipment maintenance domain can be used as a prompt to predict the repair strategy text. Here, we assume that T is the vector matrix representing the defect description text. The mathematical expression for the prediction method is as follows:

[0123]

[0124] Among them, X represents the generated result, x k:n represents the generated word representation vector from position k to n, t 1:k-1 represents the hint word representation vector from position 1 to k-1, t i Represents the defect description text word representation vector.

[0125] S2.1.4: Performance evaluation and adjustment: Based on the ROUGE series of indicators, the model was evaluated using a validation set from the equipment operation and maintenance field, and the model parameters were adjusted based on the evaluation results. This confirmed that the model's domain generation capability was enhanced after fine-tuning.

[0126] S2.2: instruction fine-tuning;

[0127] In the present invention, secondary pre-training based on a lightweight large language model improves the model's generalization capabilities in various tasks in the field of equipment operation and maintenance, accelerates the subsequent fine-tuning process, and enables the model to have basic language understanding capabilities in the field of equipment operation and maintenance, reducing training time and computing resource requirements. At the same time, the model acquires rich domain knowledge during the pre-training phase and can better adapt to new tasks. At the same time, the integrity of the knowledge is maintained. During the pre-training process, the model's weights are not significantly modified, and the knowledge obtained by the backbone model based on pre-training of general domain data can be retained, which helps to maintain the robustness of the model.

[0128] This method further fine-tunes the intermediate model after secondary pre-training, using a mixed multi-task dataset formatted with natural language descriptions for instruction fine-tuning. This allows the large language model to perform well on unseen tasks also described in the form of instructions. Through instruction fine-tuning, the large language model can follow the task instructions of new tasks without explicit examples, thus achieving better generalization capabilities. The specific implementation process is as follows:

[0129] S2.2.1: Construction of a dataset in a multi-task instruction format; the present invention constructs a common sense instruction dataset and a maintenance strategy instruction dataset. By constructing a maintenance strategy instruction dataset, the large language model's ability to discriminate questions related to faulty equipment maintenance in the field of equipment operation and maintenance is enhanced. By constructing a dataset in a multi-task instruction format, instruction fine-tuning can significantly improve model performance at a relatively small computational cost, achieving the goal of lightweight fine-tuning of a large language model of the present invention. At the same time, by training in a multi-task mode, the model can learn the characteristics of a variety of questioning patterns. This task diversity enables the model to better generalize and adapt when faced with new tasks. By using a consistent prompt format, when an unseen task uses a similar prompt format as the fine-tuning task, the model's performance will be significantly improved. At the same time, the present invention uses a common sense instruction dataset to synchronously perform fine-tuning, so that the model can use the domain knowledge it has learned during fine-tuning to more flexibly process new instructions and improve response accuracy.

[0130] S2.2.2: Supervised training; In essence, instruction fine-tuning is a method of fine-tuning a pre-trained lightweight vertical domain large language model on a set of formatted instances in the form of natural language, which is highly related to supervised fine-tuning and multi-task prompt training. In order to perform instruction fine-tuning, the present invention uses a self-constructed multi-task instruction format dataset in a supervised learning manner, and uses sequence-to-sequence loss for training to further fine-tune the pre-trained lightweight vertical domain large language model. After supervised training, the large language model can show superior generalization capabilities even in complex scenarios and few-sample scenarios.

[0131] S2.2.3: Performance evaluation and adjustment: The model performance is verified on an independent, manually constructed validation set from the field of equipment operation and maintenance. This validation set mainly includes real-world maintenance strategy case text pairs for various defective equipment. The effectiveness of instruction fine-tuning is verified using indicators such as the ROUGE series to ensure that the model is not overfitted and can generalize to unseen instructions and tasks.

[0132] S2.3: RAT-based large language model maintenance strategy generation algorithm;

[0133] The traditional RAG method retrieves all relevant information at once. When faced with complex maintenance scenarios in the field of equipment operation and maintenance, it is difficult to infer the key information needed, which leads to challenges in knowledge utilization. To address this problem, the present invention proposes a large language model maintenance strategy generation algorithm based on RAT. By linking external knowledge sources (such as equipment status, defect descriptions, and maintenance strategies, etc.), the accuracy of the model in tasks such as equipment fault diagnosis and text maintenance strategy generation is improved. Through retrieval enhancement and thinking chaining, the model can quickly obtain high-quality intelligent decision support based on specific defect descriptions, further improve the accuracy and reliability of intelligent decision-making, greatly improve the maintenance efficiency of equipment, reduce human errors, and improve the overall reliability of equipment maintenance work.

[0134] Based on the above process of strengthening the internal knowledge application capabilities of the large language model itself, in order to fully utilize the language organization capabilities of the large language model itself while avoiding the disadvantages brought by lightweighting, the present invention improves the knowledge utilization capabilities of the large language model based on technologies such as RAG and thinking chain, and solves the problems of high cost of building vertical field models and poor output professionalism.

[0135] The highly specialized nature of vertical domain knowledge places stringent demands on the quality of responses from large language models and the consistency of knowledge across multiple responses. While many leading large language models in vertical domains have demonstrated outstanding performance in incorporating multidimensional knowledge, adopting a multi-level development paradigm, enhancing Chinese text capabilities, and enhancing industry capabilities, and even making progress in multimodality, they still face practical challenges in accuracy and high construction costs when applied to domains lacking domain data. For example, in the field of equipment operation and maintenance, due to the highly specialized nature of the domain knowledge, the relatively limited knowledge base, diverse application scenarios, and complex domain issues, many key technical challenges remain to be addressed. These challenges are crucial research areas where large language model technology is urgently needed for breakthroughs. The equipment operation and maintenance sector has accumulated a vast amount of operational knowledge, and as this work progresses, the variety and scale of the accumulated data will continue to increase, laying a solid foundation for the industrial application of large language models for intelligent operation and maintenance. However, the limited accuracy and specialized nature of the knowledge limits the rapid implementation of large language model technology. Furthermore, transfer learning using large language models from other domains has not yet yielded satisfactory results, necessitating focused breakthroughs in specific domain issues.

[0136] In combination with the maintenance knowledge base constructed above for the field of equipment operation and maintenance, the present invention will further address key issues such as how domain knowledge can be utilized to improve the accuracy and reliability of maintenance strategies in the field of equipment operation and maintenance from the following three parts.

[0137] COT stands for "Chain-of-Thought" (COT). By providing examples of the reasoning process, it helps the model better perform logical reasoning and deduction when handling complex tasks. COT uses the zero-shot mode to enhance the model's reasoning ability. By using COT data during training, the model can learn how to combine multiple steps together for reasoning, which makes it perform better when faced with tasks that require logical reasoning. COT has great potential in small-scale models. Although small-scale models have fewer parameters, by providing clear steps and structured questions, small-scale models can imitate the reasoning process of larger language models to a certain extent, thereby improving their performance. Its specific implementation process is as follows:

[0138] S2.3.1: Retrieval enhancement and fine-grained training; This paper proposes the RAT method, which uses a fine-grained training strategy to enable the model to perform better in RAG tasks in specific domains. Specifically, the RAT method introduces interference documents on the basis of RAG to train the model to ignore irrelevant information, while generating high-quality answers that contain chain thinking and direct quotations. The model is fine-tuned using standard supervised learning techniques to enable it to generate answers from the provided documents. A dataset containing questions, relevant documents, and answers is collected, and the dataset is divided into two parts: one part contains documents with correct answers (key documents) and the other part contains interference documents. For each question in the dataset, a training sample is constructed, including the question, a set of documents (including key documents and interference documents), and answers in a chain thinking style. During the fine-tuning process, the model needs to learn to ignore interference documents and only focus on relevant information in key documents, thereby improving the model's ability to retrieve key information.

[0139] S2.3.2: Construct a thinking chain template; in order to improve the knowledge utilization ability of the model, the present invention encourages the model to perform multi-step reasoning when answering questions through a step-by-step reasoning model and linked thinking prompts. This method helps the model understand how to decompose the problem and draw conclusions step by step by providing some exemplary examples. Traditional RAG retrieves all relevant information at one time. However, when faced with complex scenarios or multi-step reasoning or asking vague questions, it is difficult to predict which "facts" or information are needed in the subsequent reasoning and generation steps, which leads to challenges in utilizing document knowledge. The RAT method is improved based on multiple rounds of thinking and integration models. Specifically, first, the steps are parsed from a single round of answers. In intelligent decision-making, this often corresponds to multiple potential causes of defects. Secondly, the corresponding maintenance strategies are further retrieved based on the parsed content or causes, and the existing answers are revised based on the specific content to achieve the generation and integration of multi-scenario inspection and maintenance strategies.

[0140] S2.3.3: Framework usage and model deployment; This invention builds a lightweight vertical domain large language model based on the LlamaIndex retrieval enhancement framework and the local lightweight large language model Baichuan2-13B to achieve intelligent decision-making, and provides temporary knowledge injection as a factual basis through an external maintenance knowledge base. The model retrieves relevant knowledge records based on the defect description text entered by the maintenance personnel, and fills the corresponding prompt templates with multiple knowledge records and the original defect description text according to different stages. After multiple rounds of integration and adjustment based on the above-mentioned RAT idea, the lightweight vertical domain large language model finally obtains a complete and highly readable maintenance strategy. At the same time, a similarity matching threshold for maintenance knowledge base retrieval is set, keyword inclusion retrieval is supported, and a Chinese prompt template in Refine mode is reasonably designed to optimize the professionalism of the generated results.

[0141] S2.3.4: Maintenance Strategy Generation and Evaluation; To verify that the method proposed in this invention can solve the problems of hallucination generation, insufficient knowledge timeliness, and difficulty in updating internal knowledge in vertical domain models, and can also achieve lightweight model construction, this invention compares the parameter scale and data cost required for training of the constructed lightweight vertical domain large language model with the existing general large language model. At the same time, text pairs consisting of equipment defect descriptions and corresponding maintenance strategies in real scenarios from hundreds of case reports are collected as a validation set in the field of equipment operation and maintenance. The professionalism, readability, and Chinese fluency of the maintenance strategy generation results of the constructed model and the general large language model are compared using multiple indicators such as the ROUGE series of indicators and BERTSCORE. The experimental results show that the lightweight vertical domain large language model constructed by this invention has certain advantages in both construction cost and generation effect, proving the effectiveness of the method proposed in this invention in RAG tasks in specific fields.

[0142] Step 3: Abnormal detection method for equipment defects at the operation and maintenance site;

[0143] (1) Overall design ideas of anomaly detection methods;

[0144] Currently, traditional target detection algorithms have difficulty accurately identifying small, long-range targets commonly found in inspection sites within the field of equipment maintenance, and the level of utilization of accumulated data within the field needs to be improved. To this end, the present invention proposes an innovative anomaly detection method for equipment defects at the inspection site, aiming to enhance the accuracy of maintenance strategy formulation through image analysis, effectively addressing the challenges of long-range object detection, small target identification, and abnormal target identification in complex backgrounds in the inspection environment. This method significantly enhances the accuracy of anomaly detection, further assists the maintenance strategy generation process, and improves the utilization rate of multi-source data, thereby providing image information support for intelligent decision-making.

[0145] (2) The present invention’s faulty equipment anomaly detection method based on the improved YOLOv7 model based on the ACmix attention mechanism;

[0146] In response to the problem that existing target detection algorithms are unable to cope with the recognition of small targets at long distances in the field of equipment operation and maintenance, the present invention proposes a method for detecting anomalies in faulty equipment based on an improved YOLOv7 model with the ACmix attention mechanism. By introducing the ACmix attention mechanism to optimize the model structure, the model's performance in identifying anomalies in equipment fault maintenance images is significantly improved. At different intersection-over-union thresholds, the model's precision and recall rate are improved, the phenomenon of mispick and missed detection is reduced, and positioning is more accurate. In addition, the improved YOLOv7 model can also identify small targets that were previously unrecognizable, such as floating objects and bird nests, greatly enhancing the reliability of the text maintenance strategy.

[0147] The specific implementation process is as follows;

[0148] S3.1: Domain dataset construction;

[0149] First, we extracted image data from a large collection of PDF case reports. These reports cover a variety of equipment, failure cases, and maintenance records, ensuring data diversity and representativeness. We then used layout analysis techniques to extract image-related descriptive information, such as equipment type and failure cause. After resizing and data augmentation, we used the LabelImg annotation tool to annotate the extracted and processed images, selecting object categories, drawing bounding boxes, and recording metadata. Finally, we converted the images into the dataset format required for model training.

[0150] S3.2: Fault point identification;

[0151] In the YOLOv7 model architecture, the ELAN-H module is a key component used to enhance feature extraction and fusion capabilities. It processes multiple convolutional layers of different kernel sizes through a hierarchical aggregation strategy to extract multi-scale features. It then fuses the output feature maps of multiple branches using splicing and linear superposition operations to form a comprehensive feature map. Its modular structure can be easily integrated into different neural network architectures and modified to enhance its feature representation capabilities. This paper improves the ELAN-H module using the ACmix attention mechanism to construct the ACmix-YOLOv7 model, thereby optimizing the model's feature representation capabilities. By using a multi-dimensional aggregation strategy combining standard convolutional attention mechanisms and self-attention mechanisms, it efficiently fuses features from different sources, focusing on the correlation information between input elements. This allows the model to better capture multi-scale information and is more conducive to small target detection.

[0152] S3.2.1: Standard convolutional attention mechanism feature calculation;

[0153] The standard convolutional attention mechanism uses an aggregation function on the local receptive field based on the convolution filter weights. These weights are shared across the entire feature map, and their inherent properties bring key inductive biases to image processing. Assuming the convolution stride is 1 and the standard convolution kernel is The standard convolution can be expressed as follows:

[0154]

[0155] Among them, k represents the convolution kernel size, R represents the vector matrix, C in 、C out Respectively represent the channel size of the input feature map and the output feature map, H and W represent the height and width of the input feature map, Represents the pixel tensor at position (i, j) in the input feature map and the output feature map, respectively, Indicates that the position in the input feature map f is The eigenvalues ​​of represents the weight at position (p, q) in the convolution kernel; the k×k convolution operation is regarded as a concatenation of k×k 1×1 convolutions, so the formula can be reconstructed as follows:

[0156]

[0157] By defining the shift operation Shift, the formula can be further simplified as follows:

[0158]

[0159] Among them, Δx and Δy represent the displacement in the horizontal and vertical directions respectively, f i+Δx,j+Δy Represents the eigenvalue of the corresponding position after the feature map f is shifted.

[0160] A standard convolution uses a k×k convolution kernel as a sliding window on the input feature map to perform feature transformation. After calculating and aggregating the feature values ​​within each window, the convolution kernel is shifted to the next position for convolution. The above formula can be rewritten to remodel it from a different perspective. The standard convolution operation can be viewed as first calculating the local feature map using a 1×1 convolution kernel, completing feature transformation at all positions, then shifting and aggregating. This restructures the entire process into two stages: transformation and offset aggregation. For example, a 3×3 standard convolution operation can be divided into two stages. The first stage is to divide the 3×3 convolution kernel into nine 1×1 convolutions, each of which performs convolution operations on the input feature map. The second stage is to shift and superimpose the nine output feature maps corresponding to different original convolution kernel positions. The point in the upper left corner of the 3×3 convolution kernel cannot be multiplied with the pixel in the upper right corner of the image. Instead, the pixel in the upper left corner of the image is multiplied only by the weight in the upper left corner of the convolution kernel. Therefore, the 1×1 convolution at the corresponding position should be offset one pixel to the upper left of the image and the pixels in the rightmost column and the bottom row should be discarded. In summary, the standard convolution can be summarized into the following two stages: Stage I and Stage II:

[0161]

[0162] in, It represents the output feature map obtained by translating and aggregating the projected feature map according to the kernel position. Represents the output feature map obtained by linearly projecting the input feature map along the kernel weights at position (p,q).

[0163] S3.2.2: Self-attention mechanism feature calculation;

[0164] Compared with the classic convolutional attention mechanism that establishes the association between the input feature map and the output feature map, the self-attention mechanism adopts a weighted average operation based on the input feature context, focusing more on exploring the degree of association between each element in the input feature map and using this as the main basis for feature calculation.

[0165] Among them, the self-attention weight is dynamically calculated through the similarity function between related pixel pairs. Its flexibility enables the self-attention module to adaptively focus on different areas and capture more information features. Assuming that the self-attention module has N attention heads, the mathematical expression of self-attention calculation is as follows:

[0166]

[0167] in, Represents the pixel tensor at position (i, j) in the input feature map and the output feature map, respectively, Represents the projection matrices of query, key and value respectively, N k (i, j) represents the local area with a pixel space range of k centered at position (i, j) in the input feature map, and d represents the feature dimension; f ab Represents the eigenvalue of the position (a, b) in the local area.

[0168] The process of constructing the query, key, and value matrices in the Transformer can be viewed as a 1×1 convolution operation. As part of feature learning, the same operation is shared by performing 1×1 convolutions to project features into a deeper space. In the process of calculating self-attention weights and aggregating the value matrix, this involves collecting local features and using the convolution calculation results. In summary, self-attention calculation can be summarized into the following two stages: Stage I and Stage II:

[0169]

[0170] in, Represents the eigenvalue at position (i, j) in the value vector of layer l, Represents the eigenvalue at position (a, b) in the value vector of layer l, represents the feature value at position (i, j) in the query vector of layer l, Represents the eigenvalue at position (i, j) in the key vector of layer l.

[0171] S3.2.3: ACmix attention mechanism module;

[0172] By interpreting the feature calculation processes of the above two attention mechanisms (standard convolutional attention mechanism features and self-attention mechanism) step by step, we can find that there is a close connection between the two. In particular, in the calculation process of stage one, the ACmix attention mechanism module integrates the similarities in the feature calculation process, uses the memo algorithm idea, and uniformly uses 1×1 convolution kernels to project the input feature map and save the calculation results, thereby obtaining richer intermediate features. The intermediate features are rearranged and aggregated according to the paradigms of the two attention mechanisms, achieving elegant integration while reducing the computational overhead of the projection operation. In the present invention, most of the computational overhead is performed in the 1×1 convolution, while the subsequent displacement and aggregation are lightweight.

[0173] Specifically, if Figure 5 (In the figure, C1-C5 represent convolutional blocks at different levels, used to progressively extract high-level features. Neck represents the neck network, used to fuse and enhance feature maps from different layers. Head represents the head network, used to generate the final detection results.) As shown, the ACmix attention mechanism module first performs feature segmentation. It projects the input feature map H×W×C through a 3×3 convolution, splitting it into N pieces. These sub-features are of size H×W×C / N. This captures feature information at different scales. It then performs computations in the upper convolutional path and the lower self-attention path. In the upper convolutional path, the network collects information from the local receptive field, similar to traditional convolutional attention. After passing through an N×K fully connected layer, the sub-features are shifted and aggregated. The resulting features are then convolved again, outputting H×W×C features. In the lower self-attention path, the network uses a self-attention mechanism to focus on the relationships between input features and enhances their expressiveness by calculating similarities between them. This process helps capture global information and improves the perception of small objects. Finally, feature fusion is performed. The output results calculated by the upper branch convolution path and the lower branch self-attention path are merged. The strength of the fusion is controlled by two learnable scalars α and β. The final output features are as follows.

[0174] F out =αF att +βF conv #(17)

[0175] Among them, F out Denotes the final output feature map obtained by weighting, F att represents the feature map calculated by the attention mechanism, F conv Represents the feature map obtained by convolution operation.

[0176] In this way, the ACmix attention mechanism module takes into account both global and local features, thereby improving the network's ability to detect small objects.

[0177] Based on the consideration of the differences and complementary characteristics of convolution and self-attention, the present invention found that by integrating these modules, small target detection tasks can benefit from both paradigms. Therefore, the present invention combines the advantages of the classic convolutional attention mechanism and the self-attention mechanism, introduces the ACmix attention mechanism module constructed by combining the two modules in parallel, optimizes the architecture of the original ELAN-H module, and refines the feature extraction granularity of the model. The optimized ELAN-H module architecture is shown in the figure below. Figure 6 As shown in the figure, CBS represents a composite key component, including convolution and pooling layers, which extracts local features and integrates global context information, thereby improving the representation ability of the model; cat represents the feature map splicing operation; ACmix represents the added ACmix module to improve the original ELAN-H module, strengthen the model's feature extraction ability, and enhance the information fusion ability.

[0178] S3.2.4: ACmix-YOLOv7 fault point identification;

[0179] The fault point identification is completed through the ACmix-YOLOv7 model constructed above.

[0180] Based on the existing deep learning method, the present invention realizes anomaly detection for various equipment defects in the equipment pictures in the operation and maintenance site, assists in identifying abnormal defect points in the equipment from the image level, performs target identification and positioning for equipment defects of different scales, and gives specific identification for existing or potential hidden dangers, thereby improving the capabilities of the current mainstream algorithms in field scenarios, reducing the phenomenon of false detection and missed detection, and at the same time expanding the detection scenario for abnormal defect points of equipment. According to the data collection situation, it covers detection contents such as foreign objects in bird's nests, broken insulators, abnormal silicone of respirators, damaged meters, abnormal closure of switch cabinets, etc., so as to provide factual support for the generation of text maintenance strategies, and provides a more intuitive intelligent decision-making assistance effect in the on-site operation and maintenance scenarios in the field of equipment operation and maintenance.

[0181] In a second aspect, the present invention provides a domain intelligent decision-making system based on a knowledge base, which is used to implement a domain intelligent decision-making method based on a knowledge base provided in the first aspect of the present invention.

[0182] This paper constructs a knowledge-based domain intelligent decision-making system. This system aims to help maintenance personnel and decision-makers make rapid, scientific, and rational decisions by deeply analyzing the vast amount of existing knowledge and big data within the industry. The main functions of this knowledge-based domain intelligent decision-making system include maintenance knowledge base management, text-based maintenance strategy generation, and fault point identification. Its overall architecture can be divided into four layers: front-end display layer, service layer, business layer, and data layer.

[0183] The front-end display layer is primarily responsible for user interaction. It employs a responsive design and provides an intuitive user interface through the information presentation module and the user operation module. The user operation module features user input, query, and data filtering, supporting user-defined queries and enabling fast execution of various tasks. The information presentation module visualizes information queries and system analysis data from the intelligent decision-making process in the form of reports, helping users intuitively understand relevant information and generate intelligent decision-making results.

[0184] The service layer is the entry point to the system, primarily responsible for receiving user requests and routing them to the appropriate services. Through intelligent routing strategies, the gateway can efficiently distribute requests to appropriate microservice instances, while also assuming load balancing functions to ensure efficient request processing and high system availability. The service layer supports cross-system communication and protocol conversion through the API gateway, providing a consistent interface format for different clients and reducing system complexity. At the same time, as the core of the system's external interactions, the service layer is primarily responsible for encapsulating and managing business logic. The service layer implements specific services such as information query, and handles various domain requirements through detailed business modules. The service layer adopts a microservices architecture, with each microservice unit fulfilling specific business functions, such as maintenance knowledge base data management and intelligent decision-making applications. Different microservice units are independent and decoupled from each other, allowing for flexible expansion and updating.

[0185] The business layer is the core of the entire system, responsible for the deep integration and application of domain knowledge with models and algorithms related to intelligent decision-making. Relying on a powerful maintenance knowledge base, the business layer uses large language models and deep learning algorithms to provide users with precise decision-making support. The business layer encompasses multiple submodules, such as maintenance knowledge base data query, intelligent decision-making, and anomaly detection. Each submodule performs data processing and reasoning based on specific business logic, integrating historical case data to provide users with optimized intelligent decision-making results. The submodules within the business layer are highly decoupled and operate independently, yet can collaborate through the service layer.

[0186] The data layer is the foundation of the system, consisting of a maintenance knowledge base and big data storage. The maintenance knowledge base contains high-quality domain knowledge with high information density, including specialized strategy questions and answers, classic maintenance defects and corresponding strategies, and definitions of domain-specific terminology, supporting intelligent decision-making reasoning and rule matching. The big data storage component stores a large amount of historical case data, operating procedures, and domain-related materials, providing data support for the intelligent decision-making system. The data layer uses a variety of storage methods, such as relational databases and NoSQL databases, to ensure efficient data storage and rapid retrieval. Through data mining and analysis technologies, the data layer provides high-quality data support for the business layer.

[0187] The knowledge base-based domain intelligent decision-making system constructed by the present invention has a core server built using the SpringBoot framework in combination with the Java language, while the deep learning models used in each core functional module are developed based on the Pytorch framework and Python language, and the commonly used Redis cache system is applied based on the C language.

[0188] The knowledge base-based domain intelligent decision-making system of the present invention can help operation and maintenance personnel quickly locate and extract key information from massive unstructured documents, extract and detect abnormal information from both images and text, and generate text maintenance strategies based on this, thereby providing the system with an accurate data source, promoting the intelligent development of operation and maintenance work, and promoting the intelligence and automation of operation and maintenance.

[0189] The above is only a preferred embodiment of the present invention. It should be pointed out that for ordinary technicians in this technical field, several improvements and modifications can be made without departing from the principles of the present invention. These improvements and modifications should also be regarded as within the scope of protection of the present invention.

Claims

1. A domain intelligent decision-making method based on a knowledge base, characterized by: The following steps are involved: Step 1: Build a maintenance knowledge base for equipment operation and maintenance; After data preprocessing, domain documents are first analyzed for layout using PP-PicoDet, followed by text recognition using the RARE model to complete layout analysis. Based on the layout analysis results, the PMRC model is used to transform the machine reading comprehension framework into template prompts to achieve named entity recognition. The MASKER model is used to classify text through contextual mask keyword reconstruction, and a multi-task instruction dataset is constructed. Step 2: Generate a vertical field text inspection strategy based on a general large language model; S2.1: Method for constructing lightweight large language models in vertical fields; Collect and clean domain device text data, use a lightweight large language model as the backbone model, and perform modeling in an autoregressive manner after secondary pre-training to obtain a lightweight vertical domain large language model. S2.2: instruction fine-tuning; Construct a dataset in a multi-task instruction format and perform supervised training on a pre-trained lightweight vertical domain large language model to achieve instruction fine-tuning. S2.3: RAT-based large language model maintenance strategy generation algorithm; Using the RAT method and a fine-grained training strategy, a lightweight vertical domain language model is trained by introducing interference domain documents. Standard supervised learning techniques are then used to fine-tune the lightweight vertical domain language model. The lightweight vertical domain language model retrieves relevant knowledge records from the maintenance knowledge base based on the defect description text, and fills the corresponding prompt templates with multiple knowledge records and the original defect description text according to different stages. After multiple rounds of integration and adjustment, the maintenance strategy is obtained. Step 3: Abnormal detection method for equipment defects at the operation and maintenance site; The fault point is identified by using the faulty equipment anomaly detection method based on the improved YOLOv7 model using the ACmix attention mechanism.

2. The domain intelligent decision-making method based on the knowledge base according to claim 1 is characterized in that: The proposed PP-PicoDet is based on the CSP-PAN feature fusion network, learns parameterized hierarchical structural details, and uses SimOTA as the label assignment algorithm to detect various components in domain documents.

3. The domain intelligent decision-making method based on the knowledge base according to claim 1 is characterized in that: The RARE model uses a spatial transformation network to correct text distortion and recognizes sequences through a sequence recognition network, so that it can accurately read the input content.

4. The domain intelligent decision-making method based on the knowledge base according to claim 1 is characterized in that: The backbone model is the Baichuan2-13B model. In the Baichuan2-13B model, the second multi-head attention layer is deleted from the decoder structure of its Transformer, and the first Masked multi-head attention layer is retained, that is, the ability to mask the following information during model training is retained.

5. The domain intelligent decision-making method based on the knowledge base according to claim 4 is characterized in that: The training process of the Baichuan2-13B model is as follows: (1) The words in the input text are vectorized by using the Transformer method of adding the word encoding and position encoding of the input text: first, the input text is uniquely encoded in units of words, each word is mapped into a unique vectorized representation, and then the word encoding mapping matrix W is used. e Mapping is performed to obtain the corresponding word encoding vector; for position encoding, the position encoding mapping matrix W is used p Mapping is performed, and after initializing the position encoding mapping matrix, the model itself is used for learning. The calculation formula is: h0=UW e +W p #(2) Where h0 represents the initial representation vector of each word in the input sequence, and U represents the learnable weight matrix used to map word embeddings to the hidden layer dimensions of the model; Then h0 is passed into each block step by step to get the output h of the last block. n , h n It indicates the probability of predicting the next word from each word in the input text. The calculation formula is: Among them, h l Representation vector of the model output at layer l, h l-1 Represents the representation vector output by the model at layer l-1, n represents the number of block layers in the network structure, T represents the matrix transpose, P(u) represents the probability distribution of the model for each word in the vocabulary, and softmax represents the activation function; (2) Use a large amount of unlabeled data for pre-training and perform fine-tuning on specific tasks on labeled data; (3) For the task of generating fault repair strategies in the field of equipment operation and maintenance, the defect description in the field of equipment operation and maintenance is used as the prompt content to predict the repair strategy text: Assuming that T is the defect description text representation vector matrix, its prediction formula is: Among them, X represents the generated result, x k:n represents the generated word representation vector from position k to n, t 1:k-1 represents the hint word representation vector from position 1 to k-1, t i Represents the defect description text word representation vector.

6. The domain intelligent decision-making method based on the knowledge base according to claim 1 is characterized in that: The dataset in the multi-task instruction format includes a common sense instruction dataset and a maintenance strategy instruction dataset; the maintenance strategy instruction dataset is used to enhance the lightweight vertical domain large language model's ability to identify questions related to faulty equipment maintenance in the equipment operation and maintenance field; the common sense instruction dataset is used to synchronously execute instruction fine-tuning, so that the lightweight vertical domain large language model can utilize the domain knowledge it has learned during instruction fine-tuning.

7. The domain intelligent decision-making method based on the knowledge base according to claim 1 is characterized in that: The RAT method introduces interference domain documents on the basis of RAG to train the model to ignore irrelevant information, while generating high-quality answers that contain chain thinking and direct references; collects a dataset containing questions, relevant documents and answers, and divides the dataset into two parts: one part contains domain documents with correct answers, and the other part contains interference domain documents; for each question in the dataset, constructs a training sample, including the question, a set of domain documents and an answer with chain thinking style.

8. The domain intelligent decision-making method based on the knowledge base according to claim 1 is characterized in that: In the RAT method, through a step-by-step reasoning mode, linked thinking prompts encourage the model to perform multi-step reasoning when answering questions: first, the steps are analyzed from a single round of answers, and then the corresponding maintenance strategies are further retrieved based on the content or reasons of the analysis. The existing answers are revised according to the specific content to achieve the generation and integration of multi-scenario inspection and maintenance strategies.

9. The domain intelligent decision-making method based on the knowledge base according to claim 1 is characterized in that: The ACmix attention mechanism module in the improved YOLOv7 model based on the ACmix attention mechanism first performs feature segmentation to form sub-features to capture feature information at different scales; the upper branch convolution path and the lower branch self-attention path calculation are performed separately; in the upper branch convolution path calculation, the sub-features are shifted and aggregated after passing through the fully connected layer, and the obtained features are further convolved to output H×W×C features; in the lower branch self-attention path calculation, the self-attention mechanism is used to calculate the similarity between features to enhance the expressiveness of the features; finally, feature fusion is performed to merge the output results of the upper branch convolution path and the lower branch self-attention path calculation. The strength of the merging is controlled by two learnable scalars α and β. The final output feature is: F out =αF att +βF ocnv #(17) Among them, F out Represents the final output feature map obtained by weighting, F att represents the feature map calculated by the attention mechanism, F conv Represents the feature map obtained by convolution operation.

10. A domain intelligent decision system based on a knowledge base for implementing the domain intelligent decision method based on a knowledge base according to any one of claims 1 to 9, characterized in that: include: The front-end display layer adopts responsive design and provides users with an intuitive operation interface through information presentation module and user operation module; The user operation module has user input, query, and data screening functions, and supports user-defined queries; the information presentation module visualizes the information query data and system analysis data of the intelligent decision-making process and generates intelligent decision-making results; The service layer is used to receive user requests and route them to the corresponding services. It also supports cross-system communication and protocol conversion through the API gateway, providing a consistent interface for different clients. The service layer adopts a microservice architecture, and each microservice unit completes a specific business function; The business layer is used to achieve the deep integration and application of domain knowledge and intelligent decision-making related models and algorithms; the business layer includes a maintenance knowledge base data query module, an intelligent decision-making module, and an anomaly detection module; The data layer includes a maintenance knowledge base and a big data storage module; the maintenance knowledge base contains high-quality domain knowledge with high information density, supporting intelligent decision-making reasoning and rule matching; the big data storage module stores historical case data, operating procedures and domain-related materials, providing data support for the intelligent decision-making system.

Citation Information

Cited By

  • Coal preparation plant full-category equipment anomaly identification method and system based on industrial large model

    CN122471304A

  • Coal preparation plant full-category equipment anomaly identification method and system based on industrial large model

    CN122471304B