Multi-modal large model knowledge retrieval and question-answering method and device for fault diagnosis of intelligent operation and maintenance teaching equipment
By using a multimodal large-scale model knowledge retrieval and question-answering method, the problems of low efficiency and information silos in the fault diagnosis of smart classroom teaching equipment are solved, the comprehensive utilization of multimodal information is realized, and the accuracy of fault location is improved.
Patent Information
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- BEIJING JINGYEDA TECH CO LTD
- Filing Date
- 2026-01-28
- Publication Date
- 2026-05-12
AI Technical Summary
In smart classrooms, the efficiency of fault diagnosis for teaching equipment is low, the diagnosis is not intuitive, knowledge is scattered, and it is difficult to quickly locate faulty knowledge. Traditional methods cannot effectively utilize multimodal information.
This method employs a multimodal large-scale model for knowledge retrieval and question answering. By acquiring a pre-defined vector database and user multimodal query data, it generates a comprehensive query vector and keyword set. It then retrieves highly relevant multimodal knowledge fragments from the database and uses the large-scale model to provide comprehensive teaching responses.
It improves the accuracy of fault location to over 90%, breaks through the limitations of traditional single text description, and realizes cross-validation of multi-dimensional information and accurate diagnosis.
Smart Images

Figure CN122019718A_ABST
Abstract
Description
Technical Field
[0001] This application relates to the field of data processing technology, specifically to a multimodal large-scale model knowledge retrieval and question-answering method and apparatus for fault diagnosis of intelligent operation and maintenance teaching equipment. Background Technology
[0002] The existing technology has the following disadvantages:
[0003] Low teaching efficiency: In the operation and maintenance of smart classrooms, when faced with the failure of teaching equipment (classroom podium control, server, etc.), it is difficult to quickly locate the knowledge from the massive amount of unstructured equipment manuals, drawings, and historical maintenance records. It relies on one-way guidance from teachers, resulting in a steep learning curve.
[0004] Diagnosis is not intuitive: Traditional methods mainly rely on text descriptions of faults, while equipment status is often reflected through multi-dimensional information such as indicator lights, abnormal sounds, and vibrations. Pure text question-and-answer cannot effectively utilize this multi-modal information, thus limiting the accuracy of diagnosis.
[0005] Knowledge silos: Equipment documents, sensor data, maintenance cases, and other knowledge are scattered across different systems and formats, forming information silos that prevent trainees from performing related queries and comprehensive analysis. Summary of the Invention
[0006] The purpose of this invention is to provide a multimodal large-scale model knowledge retrieval and question-answering method for fault diagnosis of intelligent operation and maintenance teaching equipment, which at least solves one of the above-mentioned technical problems.
[0007] One aspect of the present invention provides a multimodal large-scale model knowledge retrieval and question-answering method for fault diagnosis of intelligent operation and maintenance teaching equipment, the multimodal large-scale model knowledge retrieval and question-answering method for fault diagnosis of intelligent operation and maintenance teaching equipment includes:
[0008] Obtain the preset vector database;
[0009] Obtain the user's multimodal query data;
[0010] Generate a comprehensive query vector and keyword set based on the user's multimodal query data;
[0011] Based on the comprehensive query vector and keyword set, multimodal knowledge fragments highly relevant to the current fault situation are obtained from the vector database;
[0012] By inputting highly relevant multimodal knowledge fragments and user multimodal query data into a trained large model, comprehensive teaching response information can be obtained.
[0013] Optionally, the step of retrieving multimodal knowledge fragments highly relevant to the current fault situation from the vector database based on the comprehensive query vector and keyword set includes:
[0014] Generate a set of effective dimensions for fault scenarios and a confidence score table for each dimension based on the comprehensive query vector and keyword set;
[0015] Contextualized weighted query vectors and unimodal normalized weight tables are generated based on the comprehensive query vector and the confidence score tables for each dimension.
[0016] Based on the contextualized weighted query vector and keyword set, a semantic-keyword set bidirectional expansion retrieval and confidence filtering are performed in the vector database to obtain the effective set of bidirectional expansion retrieval.
[0017] Based on the effective set of bidirectional extended retrieval and vector database, the correlation strength and complementarity of cross-modal knowledge fragments are quantified and evaluated, thereby generating a set of highly correlated cross-modal knowledge groups;
[0018] Based on the set of highly correlated cross-modal knowledge groups and the confidence score table of each dimension, conflict reconciliation is carried out in a context-priority-oriented manner to obtain a set of conflict-free core knowledge groups.
[0019] Based on the set of conflict-free core knowledge groups and keyword sets, a multi-dimensional weighted ranking of teaching adaptability is performed to generate a set of high-quality multimodal knowledge fragments that are highly relevant to the current fault situation.
[0020] Optionally, the step of generating a set of effective dimensions for the fault scenario and a confidence score table for each dimension based on the comprehensive query vector and keyword set includes:
[0021] Using keyword set K as the search criteria, the top-100 relevant knowledge fragments are retrieved from the vector database, and their feature vectors are extracted. and related text descriptions ;
[0022] Description of associated text Semantic features are extracted using a BiLSTM model, and a candidate context dimension set is generated by combining it with the TF-IDF algorithm. ;
[0023] Based on the candidate scenario dimension set Generate a set of dimensional domain weights and a set of historical hit rates for each dimension;
[0024] Based on the feature vector Dimensional domain weight set, dimensional historical hit rate set, comprehensive query vector Keyword set The vector database generates a set of effective dimensions for fault scenarios and a confidence score table for each dimension.
[0025] Optionally, the step of using the feature vector Candidate Context Dimension Set Comprehensive query vector Keyword set The vector database generates a set of effective dimensions for fault scenarios and confidence scores for each dimension, including:
[0026] Based on the Top-100 relevant knowledge fragments, obtain the semantic similarity factor, dimensional text importance factor, and keyword-dimensional association factor for each knowledge fragment;
[0027] The confidence level for each candidate scenario dimension is calculated using the following formula. :
[0028] ;
[0029] in, Confidence level for each candidate scenario dimension To integrate query vectors, For feature vectors, For the j-th candidate scenario dimension Weighting coefficients in the field of operations and maintenance For the j-th candidate context dimension, Let K be the i-th text description, and K be the set of keywords. For the j-th candidate dimension Hit rate in historical fault diagnosis cases.
[0030] Optionally, the user's multimodal query data includes user query text, user-uploaded fault images, and user-uploaded abnormal audio.
[0031] The step of generating a comprehensive query vector and keyword set based on the user's multimodal query data further includes:
[0032] The user query text, user-uploaded fault images, and user-uploaded abnormal audio are encoded separately to obtain text feature vectors. Image feature vectors Abnormal audio feature vectors ;
[0033] The process of generating a contextualized weighted query vector and a unimodal normalized weight table based on the comprehensive query vector and the confidence score tables for each dimension includes:
[0034] Based on text feature vectors And keyword sets determine text integrity ;
[0035] Based on image feature vectors And keyword set to determine image sharpness ;
[0036] Based on image feature vectors And the keyword set determines the audio signal-to-noise ratio. ;
[0037] Based on text integrity Image clarity and audio signal-to-noise ratio Generate modal mass vector ;
[0038] Obtain the preset modality-dimension association mapping table;
[0039] Define a set of modal identifiers m∈{t,i,a}, where t represents text, i represents image, and a represents audio;
[0040] For each mode m, the following processing is performed:
[0041] Based on the modality-dimensional association mapping table, the number of dimensions strongly associated with a particular modality in the effective dimension set of fault scenarios is counted. ;
[0042] Calculate the context fit coefficient for this modality: ,in This represents the total number of valid dimensions in the fault scenario set; its value ranges from [0,1].
[0043] Integrate the context adaptation coefficients of various modalities to generate a context adaptation vector. ;in, Context adaptation coefficients for text modalities, Context adaptation coefficients for image modalities, Context adaptation coefficients for audio modalities;
[0044] For each modality m, calculate the mean modality-dimension semantic association degree:
[0045] Traversing the effective dimension set of fault scenarios Each dimension in The cosine similarity algorithm is used to calculate the single-mode feature vector. With dimensional feature vectors semantic similarity The value range is [−1, 1];
[0046] Combining dimensional confidence Calculate the weighted semantic relevance mean:
[0047] ;
[0048] Calculate the unnormalized weight for each modality based on the weighted semantic relevance mean. :
[0049] ;
[0050] Normalize each unnormalized weight to obtain the final normalized weight for each mode. ;
[0051] Based on the final normalized weights of each mode fusion of text feature vectors Image feature vectors Abnormal audio feature vectors Generate contextualized weighted query vectors And a single-modal normalized weight table.
[0052] Optionally, the step of performing semantic-keyword set bidirectional expansion retrieval and confidence filtering in the vector database based on the contextualized weighted query vector and keyword set to obtain the effective set of bidirectional expansion retrieval includes:
[0053] Obtain the trained WordNet semantic network;
[0054] Contextualized weighted query vectors are processed using a trained WordNet semantic network. Perform semantic extension to obtain a set of semantic extension vectors;
[0055] Obtain the trained BERT model;
[0056] For each keyword in the keyword set, input it into the trained BERT model, calculate the semantic similarity between that keyword and all terms in the operations and maintenance terminology database, and select the top 5 terms with the highest similarity as... Extended related words;
[0057] The cosine similarity algorithm is used to calculate the semantic similarity between each extended related word and the original keyword, and this similarity value is used as the weight of the related words. The weight values range from [0,1]. All original keywords and their related extended terms, along with their corresponding weights, are integrated to form an extended keyword set. ;
[0058] Using all vectors in the semantic expansion vector set as search criteria, cosine similarity matching is performed in the vector database to recall knowledge fragments with a similarity ≥ 0.6 to any expansion vector, denoted as the candidate knowledge fragment set. Each fragment in the set includes three attributes: feature vector, original resource, and similarity value with each extended vector.
[0059] Overall confidence score calculation: For each fragment in the candidate knowledge fragment set, substitute it into the following formula to calculate the overall confidence score. :
[0060] ;
[0061] in, The overall confidence level of any knowledge fragment s in the candidate knowledge fragment set; Contextualized weighted query vectors, For the feature vector of segment s, Let x be the weight value of the expansion vector at layer x. The higher the expansion layer x, the smaller the weight value. The default value is 0.95; To expand the keyword set; To expand the weight value of keyword k; This is the distance penalty coefficient, with a preset fixed value of 0.15; This is a set of effective dimensions for fault scenarios; A set of effective dimensions for fault scenarios The number of elements;
[0062] From the calculated overall confidence scores, segments exceeding the overall confidence threshold are selected to form a bidirectional expanded retrieval effective set. .
[0063] Optionally, the step of quantifying the correlation strength and evaluating the complementarity of cross-modal knowledge fragments based on the effective set of bidirectional extended retrieval and the vector database, thereby generating a set of highly correlated cross-modal knowledge groups, includes:
[0064] Obtain a preset fault topic classification dictionary, which includes topic tags such as fault type, equipment component, and processing flow;
[0065] Traversing the bidirectional extended search effective set For each knowledge fragment s in the data, extract the core information of the fragment, and label each fragment with one core fault topic tag according to the tags in the fault topic classification dictionary;
[0066] Knowledge fragments labeled with the same core fault topic tags are grouped together to form various topic group sets. where p is the number of topic groups;
[0067] For each topic group, compile all cross-modal knowledge fragments contained within the group, and denote them as follows: , where q is the number of segments in the topic group;
[0068] The total number of modalities is defined as three types: text modality, image modality, and audio modality. For each topic group, the actual number of different modal types contained within the group is counted, denoted as [missing information]. ; Calculate modal coverage The range of values is ;
[0069] Substitute the values into the following formula to calculate the modal complementarity coefficient Comp(g) of topic group g:
[0070] ;
[0071] Supplement modal complementarity coefficients for each topic group Thus, the set of topic groups G with accompanying modal complementarity coefficients is obtained;
[0072] For each topic group g, perform the following calculations:
[0073] Calculate the arithmetic mean of the overall confidence scores of all segments within the group. ;
[0074] Iterate through any two distinct segments within the group and calculate their correlation factor. ;
[0075] Calculate the overall association strength for each topic group g. ;
[0076] By setting a threshold for association strength, topic groups with a comprehensive association strength greater than the threshold are selected, forming a set of highly associated cross-modal knowledge groups. .
[0077] Optionally, the step of obtaining a conflict-free core knowledge group set by performing context-priority-oriented conflict reconciliation based on a set of highly correlated cross-modal knowledge groups and confidence scores for each dimension includes:
[0078] Obtain a fault probability database for the operations and maintenance (O&M) field, which includes the occurrence probability of different root cause types of faults. ;
[0079] From the effective dimensions of fault scenarios Extract confidence level Dimensions with a value ≥0.85 are used as the core dimension set, and the core dimensions themselves constitute the core dimension set. ;
[0080] Highly correlated cross-modal knowledge group set For each topic group, extract its core root cause of failure; calculate the relationship between each topic group and the core dimension set. The semantic relevance of each dimension is bound to the core dimension corresponding to each topic group;
[0081] By comparing the root causes of failures across all topic groups within the same core dimension, if the descriptions of the root causes are inconsistent, these groups are categorized as conflict groups, forming a set of conflict groups. Non-conflicting groups form a temporary core set. ;
[0082] Core dimensions of conflict group association Priority weights are generated by normalizing the confidence level, ensuring that the sum of the weights is 1.
[0083] For the set of conflict groups For each conflict group, calculate the topic group-dimension semantic matching degree, failure probability, mean modal quality within the group, and mean redundancy coefficient within the group.
[0084] The comprehensive reconciliation score for the conflict group is generated based on the topic group-dimension semantic matching degree, failure probability, mean modality quality within the group, and mean redundancy coefficient within the group.
[0085] For the set of conflict groups The groups are sorted in descending order of their comprehensive harmonization scores to obtain a set of conflict-free core knowledge groups.
[0086] Optionally, the set of conflict groups The groups are sorted in descending order of their overall harmonization scores to obtain a set of conflict-free core knowledge groups, which includes:
[0087] Preset score difference threshold;
[0088] If the difference between the overall reconciliation score of the highest-scoring group and the overall reconciliation score of the second-highest-scoring group is greater than or equal to the preset score difference threshold, then the highest-scoring group is retained, other conflicting groups are eliminated, and the retained group is assigned to the set of conflict-free core knowledge groups.
[0089] This application also provides a multimodal large-scale model knowledge retrieval and question-answering device for fault diagnosis of intelligent operation and maintenance teaching equipment, the multimodal large-scale model knowledge retrieval and question-answering device for fault diagnosis of intelligent operation and maintenance teaching equipment includes:
[0090] A vector database acquisition module, which is used to acquire a preset vector database;
[0091] A multimodal query data acquisition module, which is used to acquire the user's multimodal query data;
[0092] A comprehensive query vector and keyword set acquisition module is used to generate a comprehensive query vector and keyword set based on the user's multimodal query data;
[0093] A multimodal knowledge fragment acquisition module is used to acquire multimodal knowledge fragments that are highly relevant to the current fault situation from a vector database based on a comprehensive query vector and a keyword set.
[0094] The comprehensive teaching response information acquisition module is used to input highly relevant multimodal knowledge fragments and user multimodal query data into a trained large model to obtain comprehensive teaching response information.
[0095] This application overcomes the limitations of traditional single-text descriptions by integrating multi-dimensional information such as text, images, sound, and data. For example, the system can simultaneously analyze "error code E105," "indicator status in the startup video," and "abnormal friction sound," performing cross-validation to improve the accuracy of fault root cause localization to over 90% (based on tests on a specific device dataset), far exceeding the accuracy of traditional keyword retrieval (approximately 65%) or single large-model question answering (approximately 75%). Attached Figure Description
[0096] Figure 1 This is a flowchart illustrating a multimodal large-scale model knowledge retrieval and question-answering method for fault diagnosis of intelligent operation and maintenance teaching equipment according to an embodiment of this application;
[0097] Figure 2 This is another flowchart illustrating a multimodal large-scale model knowledge retrieval and question-answering method for fault diagnosis of intelligent operation and maintenance teaching equipment according to an embodiment of this application;
[0098] Figure 3 This is a recall illustration of an embodiment of this application. Detailed Implementation
[0099] To make the objectives, technical solutions, and advantages of this application clearer, the technical solutions in the embodiments of this application will be described in more detail below with reference to the accompanying drawings. In the drawings, the same or similar reference numerals denote the same or similar elements or elements having the same or similar functions throughout. The described embodiments are some, but not all, embodiments of this application. The embodiments described below with reference to the accompanying drawings are exemplary and intended to explain this application, and should not be construed as limiting this application. All other embodiments obtained by those skilled in the art based on the embodiments of this application without creative effort are within the scope of protection of this application. The embodiments of this application will be described in detail below with reference to the accompanying drawings.
[0100] like Figure 1 , Figure 2 as well as Figure 3 The multimodal large-scale model knowledge retrieval and question-answering method shown for fault diagnosis of intelligent operation and maintenance teaching equipment includes:
[0101] Obtain a pre-defined vector database. Specifically, first extract various types of structured and unstructured data (including books, industry standards, textbooks, operation manuals, circuit diagrams, maintenance procedures, historical cases, etc., containing relevant drawings and manual paragraph resources) from the knowledge base, along with associated drawing and manual paragraph resources. Use a text segmentation algorithm to split the data into fragments, and simultaneously split the association identifiers of non-text resources such as drawings. Then, use a vectorization (embedding) algorithm to convert the split data into feature vectors, and bind the drawing and manual paragraph resources to the corresponding feature vectors. Finally, store the vector data and resource association information into the vector database.
[0102] Obtain user's multimodal query data; specifically, user's multimodal query data includes multimodal forms such as images, sound, text, device model, and fault codes. The text contains keyword information, such as "E105" and "startup" in the error message E105 during startup.
[0103] Based on the user's multimodal query data, a comprehensive query vector and keyword set are generated. Specifically, a visual encoder is used to process image information, an audio encoder to process sound information, and a text encoder to process text information, converting various non-textual information into feature vectors. Then, an embedding algorithm is used to map the feature vectors of all modalities to a unified semantic space and fuse them to form a comprehensive query vector.
[0104] Based on the comprehensive query vector and keyword set, multimodal knowledge fragments highly relevant to the current fault situation are obtained from the vector database;
[0105] Multimodal knowledge fragments highly relevant to the current fault situation, along with users' multimodal query data, are input into a trained large model to obtain comprehensive teaching response information. Specifically, the input information is fed into a large model (Starry Sky Education, DeepSeek-R1), and through the model's causal reasoning and logical judgment algorithms, the causes of the fault are analyzed, troubleshooting steps are outlined, relevant principles are explained, and safety precautions are highlighted. The knowledge sources are also labeled (e.g., referencing Section 3.2 of the XX manual), generating instructional answers. Format conversion algorithms convert the answers into structured text, diagram annotations, and voice broadcasts; link mapping algorithms generate direct access links to relevant drawings and manual sections.
[0106] In this embodiment, the step of retrieving multimodal knowledge fragments highly relevant to the current fault situation from the vector database based on the comprehensive query vector and keyword set includes:
[0107] Generate a set of effective dimensions for fault scenarios and a confidence score table for each dimension based on the comprehensive query vector and keyword set;
[0108] Contextualized weighted query vectors and unimodal normalized weight tables are generated based on the comprehensive query vector and the confidence score tables for each dimension.
[0109] Based on the contextualized weighted query vector and keyword set, a semantic-keyword set bidirectional expansion retrieval and confidence filtering are performed in the vector database to obtain the effective set of bidirectional expansion retrieval.
[0110] Based on the effective set of bidirectional extended retrieval and vector database, the correlation strength and complementarity of cross-modal knowledge fragments are quantified and evaluated, thereby generating a set of highly correlated cross-modal knowledge groups;
[0111] Based on the set of highly correlated cross-modal knowledge groups and the confidence score table of each dimension, conflict reconciliation is carried out in a context-priority-oriented manner to obtain a set of conflict-free core knowledge groups.
[0112] Based on the set of conflict-free core knowledge groups and keyword sets, a multi-dimensional weighted ranking of teaching adaptability is performed to generate a set of high-quality multimodal knowledge fragments that are highly relevant to the current fault situation.
[0113] In this embodiment, the step of generating a set of effective dimensions for the fault scenario and a confidence score table for each dimension based on the comprehensive query vector and keyword set includes:
[0114] Using the keyword set K as the retrieval condition, perform semantic matching retrieval in the vector database, recall the top 100 operation and maintenance related knowledge fragments with the highest relevance to the keywords, and extract the corresponding feature vectors of these 100 knowledge fragments from the vector database, denoted as (i = 1, 2,..., 100) (the normalized feature vector of the i-th recalled knowledge fragment, and the unified feature vector generated after encoding the user's multi-modal query data are of the same dimension); synchronously extract the original text description corresponding to each knowledge fragment, denoted as ;
[0115] For the text description use the BiLSTM model to extract semantic features, and combine with the TF-IDF algorithm to generate a candidate context dimension set ;
[0116] In this embodiment, for the text description use the BiLSTM model to extract semantic features, and combine with the TF-IDF algorithm to generate a candidate context dimension set including:
[0117] Text preprocessing: Perform word segmentation on each text description perform stop word filtering (such as meaningless words like "of", "is", "then", etc.), remove punctuation marks and special characters, and obtain a normalized text sequence;
[0118] Semantic feature extraction: Input the normalized text sequence into the BiLSTM (Bidirectional Long Short-Term Memory Network) model, and through forward and backward recurrent calculations, capture the long-distance semantic association information hidden in the text, and output the deep semantic feature vector of each text;
[0119] Dimension screening and induction: Use the TF-IDF (Term Frequency-Inverse Document Frequency) algorithm to calculate the weight value of each term in the normalized text sequence, and screen out the top 20 core terms with the highest weights; according to the logical classification of fault diagnosis in the operation and maintenance field, classify the core terms into different context dimension categories to form a candidate context dimension set ;
[0120] Obtain the dimension domain weight set and the dimension historical hit rate set;
[0121] In this embodiment, the dimension domain weight set is obtained through the following method:
[0122] Build an operation and maintenance field fault dimension weight statistics library:
[0123] Based on a comprehensive historical fault diagnosis case library accumulated by enterprises / industries (containing at least 1000+ valid cases, each case must be labeled with fault dimensions, fault location success rate, and diagnostic expert evaluation information), all fault dimensions that appeared in historical cases are extracted. A statistical database linking dimensions to diagnostic value is established. The core fields of this database include: candidate dimensions. The percentage of times this dimension appears as the core diagnostic criterion in the case studies; the percentage of times this dimension effectively guides the localization of the root cause of the fault.
[0124] Quantization calculation of initial weights of dimensions :
[0125] For each candidate dimension The initial weights are calculated based on the statistical database data, using the following formula: =0.6 × frequency of occurrence + 0.4 × effective percentage;
[0126] Frequency of occurrence percentage refers to the dimension The number of cases used as the core diagnostic basis in all historical cases ÷ the total number of historical cases, with a value of [0,1];
[0127] Effective proportion refers to the dimension Number of cases that served as core diagnostic criteria and successfully located the fault ÷ Dimension The total number of cases, with values in [0,1].
[0128] Initial weights The value is [0,1], and is subsequently normalized to the preset range of [0.1,0.9].
[0129] Domain expert calibration yields the final weights :
[0130] Assemble a team of 3-5 domain experts with over 5 years of experience in operation and maintenance fault diagnosis. These experts will adjust the initial weights based on actual diagnostic scenarios. Conduct a rationality check, focusing on calibrating special dimensions (such as the "parameter range" dimension, which has extremely high guiding value in the diagnosis of precision equipment, but has a low statistical frequency).
[0131] Expert calibration employs a scoring system (calibration range ±0.05-±0.15). After calibration, the weights of all candidate dimensions are normalized and mapped to ensure the final weights. It strictly falls within the range of [0.1, 0.9], and the weight differences between dimensions match the actual diagnostic importance;
[0132] For example, if the initial weight of the component type dimension is 0.92, it is mapped to 0.9 (the highest weight in the operation and maintenance field) after expert calibration; and the initial weight of the fault impact range dimension is 0.25, it is 0.3 after calibration, which fully meets the value constraints.
[0133] The method for obtaining the historical hit rate set of a dimension is as follows:
[0134] Historical hit rate Representation candidate dimensions The probability of successfully guiding fault location in historical fault diagnosis is a purely objective statistical value, with a strictly constrained range of [0.01, 1.0]:
[0135] Hit data for targeted statistical candidate dimensions:
[0136] For each candidate dimension From the historical fault diagnosis case library in the operation and maintenance field, specific candidate dimensions are selected. For all historical cases covering the core diagnostic dimensions, two key data points were compiled:
[0137] Number of valid hit cases : By candidate dimension The number of cases where the root cause of the failure was successfully located, with the core diagnostic dimension being the number of cases.
[0138] Total number of cases in each dimension : Clearly define candidate dimensions The total number of cases (including successful and unsuccessful cases) for the core diagnostic dimensions;
[0139] If a certain candidate dimension Total number of cases (A completely new dimension with no historical data) will default to its hit rate. (Lower limit of the value range) to avoid null values affecting subsequent calculations.
[0140] Directly calculate historical hit rate :
[0141] Based on statistical data, the historical hit rate of each candidate dimension is calculated directly using the following formula:
[0142] ;
[0143] The calculation result naturally falls within [0,1]. If the result is 0 (no successful cases), it is corrected to 0.01 according to the value constraint. The remaining results are directly used as the final value.
[0144] For example, if the fault stage dimension was successfully located 850 times out of 1000 historical cases, then... If the environmental factor dimension is successfully located 10 times out of 50 cases, then... .
[0145] Based on the feature vector Dimensional domain weight set, dimensional historical hit rate set, comprehensive query vector Keyword set The vector database generates a set of effective dimensions for fault scenarios and a confidence score table for each dimension.
[0146] This approach breaks through the limitations of traditional methods that rely solely on frequency of occurrence or single similarity for dimension selection. This application uses a keyword set K to recall the Top-100 relevant knowledge fragments in a vector database, extracting the feature vector and text description for each fragment. A BiLSTM model is used to extract semantic features from the text, combined with the TF-IDF algorithm to generate a set of candidate context dimensions. For each candidate dimension, five key factors are calculated in tandem (cosine similarity between the query vector and the knowledge fragment vector, dimension domain weight, dimension's TF-IDF value in the text, semantic similarity between the keyword and the dimension, and historical hit rate of the dimension). Finally, the confidence score is calculated using a formula. This design deeply integrates objective semantic matching, domain experience weighting, and historical data verification, completely avoiding subjectivity in dimension selection and ensuring that the selection of each effective dimension is supported by clear data. The number of dimensions does not need to be preset. It is dynamically matched by the number of elements m in the candidate scenario dimension set. First, m candidate dimensions are obtained through text mining. Then, a set of dimension domain weights and a set of historical hit rates containing m elements are constructed. Each candidate dimension corresponds to a set of parameters to adapt to different fault query scenarios (such as different scenarios such as device power supply failure and heat dissipation failure will mine different numbers of candidate dimensions). This solves the rigidity problem of the traditional fixed dimension system being unable to adapt to multiple scenarios.
[0147] In this embodiment, based on the feature vector Dimensional domain weight set, dimensional historical hit rate set, comprehensive query vector Keyword set The vector database generates a set of effective dimensions for fault scenarios and confidence scores for each dimension, including:
[0148] Based on the Top-100 relevant knowledge fragments, obtain the semantic similarity factor, dimensional text importance factor, and keyword-dimensional association factor for each knowledge fragment;
[0149] Calculate the basic matching factor: for each candidate dimension Iterate through 100 knowledge fragments and calculate the following three factors in sequence:
[0150] Semantic similarity factor: Calculates the comprehensive query vector With the feature vector of the i-th knowledge fragment The cosine similarity is denoted as . The value ranges from [-1, 1]. The closer the value is to 1, the higher the semantic matching degree between the two.
[0151] Dimensional text importance factor: Calculating candidate dimensions The i-th text description The TF-IDF value in the data is denoted as... The value ranges from [0,1]. The higher the value, the more important the dimension is in the text.
[0152] Keyword-Dimension Relationship Factor: Calculates the relationship between the keyword set K and candidate dimensions using the Word2Vec model. The semantic similarity is denoted as The value ranges from [0,1]. The higher the value, the closer the relationship between the keyword and the dimension.
[0153] The confidence level for each candidate scenario dimension is calculated using the following formula. :
[0154] ;
[0155] in, Confidence level for each candidate scenario dimension To integrate query vectors, For feature vectors, For the j-th candidate scenario dimension Weighting coefficients in the field of operations and maintenance For the j-th candidate context dimension, Let K be the i-th text description, and K be the set of keywords. For the j-th candidate dimension Hit rate in historical fault diagnosis cases.
[0156] Set confidence threshold Filter out those that meet the requirements Candidate dimensions are used to form a set of effective dimensions for fault scenarios. ;
[0157] Compile the confidence scores for each valid dimension to generate a dimension confidence score table.
[0158] In this embodiment, the user's multimodal query data includes user query text, user-uploaded fault images, and user-uploaded abnormal audio.
[0159] The step of generating a comprehensive query vector and keyword set based on the user's multimodal query data further includes:
[0160] The user query text, user-uploaded fault images, and user-uploaded abnormal audio are encoded separately to obtain text feature vectors. Image feature vectors Abnormal audio feature vectors ;
[0161] The process of generating a contextualized weighted query vector and a unimodal normalized weight table based on the comprehensive query vector and the confidence score tables for each dimension includes:
[0162] Based on text feature vectors And keyword sets determine text integrity Specifically, extract the core keywords from the user's query text (which are consistent with the keyword set K), and count the actual number of keywords contained in the text. ;
[0163] Calculate text integrity: ,in The total number of keywords in the keyword set K; the value ranges from [0,1], and the closer the value is to 1, the more complete the description of the query intent in the text.
[0164] Based on image feature vectors And keyword set to determine image sharpness Specifically, the fault image is converted to grayscale to obtain a grayscale image matrix.
[0165] The Laplacian operator is used to perform convolution on the grayscale image to extract image edge information; the variance of the convolution result is calculated. ;
[0166] Normalize the variance values to the [0,1] interval to obtain the image sharpness: ,in , These are the minimum and maximum values of the Laplacian variance for a fault image dataset in the operations and maintenance field; the closer the value is to 1, the clearer the image and the more effective the visual features.
[0167] Based on image feature vectors And the keyword set determines the audio signal-to-noise ratio. Specifically, the abnormal audio segments are processed by frame segmentation (frame length 20ms, frame shift 10ms), and a fast Fourier transform is performed on each frame to obtain the spectrum matrix.
[0168] Distinguish between the signal frequency band (the frequency band that differs significantly from the audio spectrum of the device during normal operation) and the noise frequency band in the spectrum, and calculate the signal power separately. With noise power Calculate the signal-to-noise ratio and normalize it to the [0,1] interval: The closer the value is to 1, the higher the proportion of effective signal in the audio and the lower the noise interference.
[0169] Based on text integrity Image clarity and audio signal-to-noise ratio Generate modal mass vector ;
[0170] Obtain a pre-defined modality-dimension association mapping table (e.g., text modality corresponds to dimensions such as fault stage and parameter range; image modality corresponds to dimensions such as component type and abnormal features; audio modality corresponds to dimensions such as abnormal noise type and fault severity).
[0171] Define a set of modal identifiers m∈{t,i,a}, where t represents text, i represents image, and a represents audio;
[0172] For each mode m, the following processing is performed:
[0173] Based on the modality-dimensional association mapping table, the number of dimensions strongly associated with a particular modality in the effective dimension set of fault scenarios is counted. ;
[0174] Calculate the context fit coefficient for this modality: ,in This represents the total number of valid dimensions in the fault scenario set; its value ranges from [0,1].
[0175] Integrate the context adaptation coefficients of various modalities to generate a context adaptation vector. ;in, Context adaptation coefficients for text modalities, Context adaptation coefficients for image modalities, Context adaptation coefficients for audio modalities;
[0176] For each modality m, calculate the mean modality-dimension semantic association degree:
[0177] Traversing the effective dimension set of fault scenarios Each dimension in The cosine similarity algorithm is used to calculate the single-mode feature vector. With dimensional feature vectors semantic similarity The value ranges from [−1, 1], and the closer the value is to 1, the higher the semantic correlation between the two.
[0178] Combining dimensional confidence Calculate the weighted semantic relevance mean:
[0179] ;in, The weighted mean of semantic relevance; This is a set of effective dimensions for fault scenarios;
[0180] Calculate the unnormalized weight for each modality based on the weighted semantic relevance mean. :
[0181] ;in, Let m be the single-modal quality scalar index value corresponding to the m-th modality; where m∈{t,i,a} is the modality identifier (t=text, i=image, a=audio), which expands to: For text modality quality scalar (text integrity, a specific number from 0 to 1, such as 0.8, 0.65); This is a quality scalar for the image modality (image sharpness, a specific number between 0 and 1, such as 0.92 or 0.78). It is a quality scalar for the audio modality (audio signal-to-noise ratio, a specific number between 0 and 1, such as 0.85 or 0.7).
[0182] The context adaptation coefficient is the value corresponding to the m-th modality.
[0183] Normalize each unnormalized weight to obtain the final normalized weight for each mode. ;
[0184] Based on the final normalized weights of each mode fusion of text feature vectors Image feature vectors Abnormal audio feature vectors Generate contextualized weighted query vectors And a single-modal normalized weight table.
[0185] In this embodiment, the contextualized weighted query vector is obtained using the following formula:
[0186] ;
[0187] Organize the final normalized weights for each mode to generate a single-mode normalized weight table {Wt,Wi,Wa}.
[0188] Through the above steps, this application abandons the traditional fixed weight or similarity-only weighting method, and adopts a two-factor constraint: first, it calculates the objective quality indicators (text integrity, image clarity, and audio signal-to-noise ratio) of the three modalities through modality quality assessment; then, it calculates and statistically analyzes the proportion of strongly correlated dimensions between each modality and the effective dimension set of the fault scenario through the context fit coefficient, thus obtaining the context fit coefficient; traversing... For all dimensions, the semantic similarity between the single-modal vector and the dimensional feature vector is calculated, and the weighted semantic relevance mean is obtained by combining the dimensional confidence. The modality quality context adaptation coefficient is multiplied by 1 + the semantic relevance mean using a formula to obtain the unnormalized weights. Finally, the weights of each modality are obtained after normalization. This design allows the weight allocation to consider both the quality of the modality itself (e.g., reducing the weight of blurred images) and the needs of fault scenarios (e.g., increasing the weight of modalities strongly correlated with core dimensions), ensuring the targeted nature of multimodal fusion.
[0189] A two-step method of unnormalized weight calculation + global normalization is adopted. First, the unnormalized weights of text, image and audio are obtained through formula. Then, the normalized weights are obtained by dividing the unnormalized weights of each modality by the sum of the unnormalized weights of all modalities. This method not only preserves the relative importance differences of each modality (e.g., the weight of high-quality and highly adaptable image modality is higher than that of low-quality text modality), but also strictly ensures that the sum of the weights is 1. This makes the fusion process interpretable and quantifiable, and avoids the chaos of weight allocation.
[0190] In this embodiment, the step of performing semantic-keyword set bidirectional expansion retrieval and confidence filtering on the vector database based on contextualized weighted query vectors and keyword sets to obtain a valid bidirectional expansion retrieval set includes:
[0191] Obtain the trained WordNet semantic network;
[0192] Contextualized weighted query vectors are processed using a trained WordNet semantic network. Perform semantic extension to obtain a set of semantic extension vectors;
[0193] Specifically, semantic extension network construction: using the effective set of fault context dimensions Each effective dimension As a core node, perform hierarchical expansion within the WordNet semantic network:
[0194] Level 1 Expansion: Relevance and Effective Dimensions Directly synonymous or near-synonymous terms (such as power supply module failure associated with power module malfunction, power supply unit failure, etc.).
[0195] Level 2 Expansion: Further associate the terms obtained from Level 1 expansion with their subordinate conceptual terms (e.g., power module malfunction is associated with power module voltage instability, power module current overload, etc.).
[0196] The maximum number of expansion layers is set to 3 to avoid semantic drift caused by too many expansion layers.
[0197] Extended vector generation:
[0198] The terms obtained from each layer of expansion are transformed into their corresponding terms using the Sentence-BERT model. The feature vectors with consistent dimensions are denoted as the x-th layer expansion vector. (x=1,2,3, representing the number of expansion layers);
[0199] Define the original contextualized weighted query vector This is the expansion vector of layer 0 (x=0), i.e. ;
[0200] Introducing an expansion attenuation coefficient β=0.95, the weights of the expansion vectors in each layer are calculated. Where x is the extension layer number; the higher the layer number, the lower the weight, reflecting the principle that synonyms have higher weights and distant terms have lower weights; all extension vectors are integrated to form a semantic extension vector set. And record the number of expansion layers and weights corresponding to each vector;
[0201] Obtain the trained BERT model;
[0202] For each keyword in the keyword set Input the keyword into the trained BERT model, calculate the semantic similarity between the keyword and all terms in the operations and maintenance terminology database, and select the top 5 terms with the highest similarity as... Extended related words ;
[0203] The cosine similarity algorithm is used to calculate the extended related words. Related to the original keywords The semantic similarity is used as the weight of related words. , The weight values range from [0,1]. Higher values indicate a stronger semantic connection between the related words and the original keywords. All original keywords and their extended related words, along with their corresponding weights, are integrated to form an extended keyword set. , Each element in the collection contains three attributes: term text, its original keyword, and weight.
[0204] Using all vectors in the semantic expansion vector set as search criteria, cosine similarity matching is performed in the vector database to recall knowledge fragments with a similarity ≥ 0.6 to any expansion vector, denoted as the candidate knowledge fragment set. Each segment in the set includes a feature vector. The three attributes are: original resources, similarity values with each extended vector;
[0205] Overall confidence score calculation: For each fragment in the candidate knowledge fragment set, substitute it into the following formula to calculate the overall confidence score. :
[0206] ;
[0207] in, The overall confidence level of any knowledge fragment s in the candidate knowledge fragment set; Contextualized weighted query vectors, For the feature vector of segment s, Let x be the weight value of the expansion vector at layer x. The higher the expansion layer x, the smaller the weight value. The default value is 0.95; To expand the keyword set; To expand the weight value of keyword k; This is the distance penalty coefficient, with a preset fixed value of 0.15; This is a set of effective dimensions for fault scenarios; A set of effective dimensions for fault scenarios The number of elements;
[0208] Select those scores exceeding the overall confidence threshold from the calculated overall confidence scores. The fragments form a bidirectional extended retrieval effective set .for Each segment is supplemented with attributes such as comprehensive confidence score, semantic retrieval score, keyword retrieval score, and distance penalty value to improve the set information.
[0209] By employing the above method, this application addresses the semantic drift and redundancy issues caused by traditional extended retrieval methods that only expand without reducing content. The core approach focuses on the mid-dimensional semantic network, employing a three-layer extension using the WordNet semantic network. The weights of the extension vectors decrease with each layer; the weight of the first layer (original query vector) is 1, while the weight of the third layer extension vector is only 0.953 ≈ 0.857. This ensures that near-sense extension vectors have high weights and far-sense extension vectors have low weights, preventing semantic deviation from the core query intent. Distance penalty control is implemented: a distance penalty term is introduced into the overall confidence calculation, where... It is a fragment of knowledge and The average semantic distance across all dimensions is used to calculate the penalty for larger distances (weaker correlation with the core context), which directly reduces the overall confidence of redundant fragments, achieving expansion without chaos and balancing recall and accuracy.
[0210] In this embodiment, the step of quantifying the correlation strength and evaluating the complementarity of cross-modal knowledge fragments based on the effective set of bidirectional extended retrieval and the vector database, thereby generating a set of highly correlated cross-modal knowledge groups, includes:
[0211] Obtain a preset fault topic classification dictionary, which includes topic tags such as fault type, equipment component, and processing flow;
[0212] Traversing the bidirectional extended search effective set For each knowledge fragment s in the data, extract the core information of the fragment (extract keywords from text fragments, extract the fault type corresponding to visual features from image fragments, and extract the anomaly type corresponding to spectral features from audio fragments). Based on the tags in the fault topic classification dictionary, label each fragment with one core fault topic tag (such as unstable power supply module voltage, abnormal noise from cooling fan).
[0213] Knowledge fragments labeled with the same core fault topic tags are grouped together to form various topic group sets. where p is the number of topic groups;
[0214] For each topic group Organize all cross-modal knowledge fragments contained within the group, and denote them as follows: Where q is the number of fragments within the topic group; the modality type and feature vector of each fragment are recorded synchronously. Overall confidence score ;
[0215] The total number of modalities is defined as three types: text modality, image modality, and audio modality. For each topic group, the actual number of different modal types contained within the group is counted, denoted as [missing information]. ; Calculate modal coverage The range of values is ;
[0216] Substitute the values into the following formula to calculate the modal complementarity coefficient Comp(g) of topic group g:
[0217] ; This represents the number of modal types within the group, with values of 1, 2, or 3. The larger the value, the richer the modalities within the group, and the stronger the complementarity; when =1 (single mode only), Comp(g)=0, the complementarity is the weakest; when =2 (two modes), Comp(g)=0.5; when =3 (all three modes are complete), Comp(g)=32, with the strongest complementarity.
[0218] Supplement modal complementarity coefficients for each topic group Thus, the set of topic groups G with accompanying modal complementarity coefficients is obtained;
[0219] For each topic group g, perform the following calculations:
[0220] Calculate the arithmetic mean of the overall confidence scores of all segments within the group. The formula is as follows: Where Q represents the total number of segments within topic group g;
[0221] Iterate through any two distinct segments within the group. , Calculate the correlation factor between the two. The formula is as follows:
[0222] ; The cosine similarity between the feature vectors of two segments represents the semantic relevance. The predefined association weights of the two segments in the knowledge base represent the degree of logical association. The intermodal correlation coefficient between two segments is taken from a lookup table according to the modal type, and it represents the degree of cross-modal correlation.
[0223] Calculate the overall association strength for each topic group g. The formula is as follows:
[0224] ;in, The total number of segments within topic group g; This means averaging the correlation factors of all pairs of segments within a group to eliminate the influence of differences in the number of segments; The modal complementarity coefficient reflects the complementary value of cross-modal information within a group. This represents the average confidence score of the segments within the group, reflecting the overall query match rate of the segments within the group; value range: The higher the value, the stronger the correlation and complementarity of cross-modal segments within the topic group;
[0225] By setting a threshold for association strength, topic groups with a comprehensive association strength greater than the threshold are selected, forming a set of highly associated cross-modal knowledge groups. Specifically, a threshold for association strength is set. Filter out those that meet the requirements Thematic groups form a set of highly correlated cross-modal knowledge groups. ,for Supplementary information for each topic group: topic tags, cross-modal segment list, and overall association strength score. Modal complementarity coefficients Confidence level mean .
[0226] This application breaks through the limitations of traditional methods that only assess fragment similarity. Addressing the need for "multi-dimensional verification" in operational fault diagnosis, this application first groups valid fragments by fault theme (ensuring fragments within a group revolve around the same fault core). Then, it iterates through all pairwise fragments within a group, calculating the association factor as cosine similarity of the feature vectors of the two fragments × predefined association weights in the knowledge base × inter-modal association coefficient. The average of all association factors is then taken to obtain the average association strength of the fragments within the group. Secondly, complementarity assessment: The number of modal types included in each group (the actual number included in text, image, and audio categories) is counted. The modal complementarity coefficient is calculated using a formula; the more modal types included (e.g., simultaneously including text, images, and audio), the higher the complementarity coefficient, reflecting the complementary value of multimodal information. Finally, the comprehensive association strength is obtained through a formula, ensuring both close association of fragments within a group and complementary multimodal information, thus meeting the actual needs of cross-validation in fault diagnosis.
[0227] In this embodiment, the step of obtaining a conflict-free core knowledge group set by performing context-priority-oriented conflict reconciliation based on a highly correlated cross-modal knowledge group set and confidence score tables for each dimension includes:
[0228] Obtain a fault probability database for the operations and maintenance (O&M) field, which includes the occurrence probability of different root cause types of faults. Generated based on historical fault statistics;
[0229] From the effective dimensions of fault scenarios Extract confidence level Dimensions with a value ≥0.85 are designated as core dimensions, and these core dimensions form the core dimension set. ;
[0230] Highly correlated cross-modal knowledge group set Each topic group Extract the core root causes of failures; calculate the set of core dimensions for each topic group. The semantic relevance of each dimension is bound to the core dimension (i.e., the one with the highest relevance) corresponding to each topic group. );
[0231] Comparing the same core dimension If the root causes of failures in all topic groups are described inconsistently (e.g., within the same power supply module, one group points to unstable voltage, and another to short circuit), then these groups are classified as conflict groups, forming a conflict group set. Non-conflicting groups form a temporary core set. ;
[0232] Core dimensions of conflict group association Priority weights are generated by normalizing the confidence level, ensuring that the sum of the weights is 1.
[0233] For the set of conflict groups For each conflict group, calculate the topic group-dimension semantic matching degree, failure probability, mean modal quality within the group, and mean redundancy coefficient within the group.
[0234] In this embodiment, topic group-dimension semantic matching degree The method for obtaining the mean vector of feature vectors of all segments within the group is as follows: , and core dimensions eigenvectors The cosine similarity is calculated with a value in the range [−1, 1], where the closer the value is to 1, the stronger the association.
[0235] Failure probability Extract the preset probability corresponding to the root cause of the fault within the group and match it from the fault probability database;
[0236] Mean modal quality within the group Based on the modal quality of each segment within the group (text) ,image Audio ), take the arithmetic mean, with a value range of [0,1];
[0237] Mean of intragroup redundancy coefficient Calculate the redundancy coefficient of all segments within the group. The arithmetic mean of the , with values ranging from [0.8, 1.0];
[0238] The comprehensive reconciliation score for the conflict group is generated based on the topic group-dimension semantic matching degree, failure probability, mean modality quality within the group, and mean redundancy coefficient within the group.
[0239] In this embodiment, the comprehensive harmonization score is obtained using the following formula. :
[0240] ;
[0241] in, Priority weights for core dimensions; The semantic matching degree between the group and the core dimension is calculated based on cosine similarity;
[0242] For the set of conflict groups The groups are sorted in descending order of their comprehensive harmonization scores to obtain a set of conflict-free core knowledge groups.
[0243] In this embodiment, the set of conflict groups The groups are sorted in descending order of their overall harmonization scores to obtain a set of conflict-free core knowledge groups, which includes:
[0244] Preset score difference threshold;
[0245] If the difference between the overall reconciliation score of the highest-scoring group and the overall reconciliation score of the second-highest-scoring group is greater than or equal to the preset score difference threshold, then the highest-scoring group is retained, other conflicting groups are eliminated, and the retained group is assigned to the set of conflict-free core knowledge groups.
[0246] right All groups are organized and supplemented with information such as comprehensive reconciliation scores, conflict markers (no conflict / requires verification), and related core dimensions to form a set of conflict-free core knowledge groups. .
[0247] This application does not blindly compare the root causes of failures across all topic groups. Instead, it first selects a set of core dimensions with a confidence level ≥ 0.85 (these dimensions play a decisive role in determining the root causes of failures). Then, it calculates the semantic correlation between each topic group and each dimension in the set of core dimensions, binds the core dimensions corresponding to the group, and only compares the root causes of failures across all topic groups under the same core dimension. If the root causes are inconsistent, the group is determined to be in conflict. This ensures that the conflict determination focuses on the core and reduces ineffective reconciliation (such as differences in root causes under different core dimensions that do not need to be reconciled).
[0248] Conflict reconciliation is not based solely on similarity, but rather integrates four key factors: context priority weight, failure occurrence probability, mean modal quality within the group, and mean redundancy coefficient within the group. The arithmetic mean of the redundancy coefficients of all segments within the group is calculated (the lower the redundancy, the smaller the coefficient, and the higher the information uniqueness). These four factors are multiplied by a formula to obtain a comprehensive reconciliation score, which considers both theoretical matching and actual occurrence probability and data reliability, making the conflict resolution results more consistent with actual operation and maintenance scenarios.
[0249] In this embodiment, a multi-dimensional weighted ranking of teaching adaptability is performed based on the set of conflict-free core knowledge groups and the set of keywords, thereby generating a set of high-quality multimodal knowledge fragments that are highly relevant to the current fault situation:
[0250] For the set of conflict-free core knowledge groups All knowledge fragments s (text, image, audio multimodal) are input into the BERT+CRF model to identify four types of core elements of operation and maintenance teaching. The identification rules and labeling methods are as follows: ;
[0251] Generate a teaching element label vector F=( for each knowledge segment) , , , ), and label the corresponding teaching elements (such as "explanation of principles + operation steps + safety tips").
[0252] Based on the element tag values and preset weights, the matching degree of teaching elements for each segment is calculated using the following formula:
[0253] ;
[0254] For teaching element types (p=principle, s=step, safe=safety, trac=tracing the source); The feature label value is 1 / 0. To preset the weights of teaching elements and reflect the importance of different elements in operation and maintenance teaching (principles > steps > safety > traceability), you can set them as needed; For the matching degree of teaching elements;
[0255] After calculating all segments, a knowledge segment-teaching element matching table can be obtained (including segment ID, ...). (Tag value, Match(s) score);
[0256] The RankScore(s) of the teaching guidance is calculated segment by segment, using the following formula:
[0257] ;
[0258] Convert the fragment text / description into a Sentence-BERT feature vector, calculate the cosine similarity between this vector and the K-mean vector of the keyword set, with a value ranging from 0 to 1. The higher the value, the higher the relevance between the fragment and the user's query keywords. To ensure that values are assigned according to the preset mapping rules (Equipment Manual = 1.0, Enterprise Fault Cases = 0.85, Industry Standards = 0.75, Others = 0.7), the values are between 0.7 and 1.0. For cross-modal correlation strength; For overall confidence level;
[0259] Sort all fragments in descending order of RankScore(s) to generate a sorted set of high-quality multimodal knowledge fragments (including fragment content, resource index, RankScore(s), and teaching element tags).
[0260] In this embodiment, the comprehensive teaching response information includes a comprehensive and teaching answer with fault cause analysis, troubleshooting steps, principle explanation, safety precautions and knowledge source annotations, as well as corresponding drawing and manual section resource identifiers;
[0261] In this embodiment, the answer is converted into structured text, schematic diagram annotations, voice broadcast, etc. through a format conversion algorithm; based on resource identifiers, the drawing and manual paragraph resources bound in the vector database are called through a link mapping algorithm to generate direct access links.
[0262] This application also provides a multimodal large-scale model knowledge retrieval and question-answering device for fault diagnosis of intelligent operation and maintenance teaching equipment, the multimodal large-scale model knowledge retrieval and question-answering device for fault diagnosis of intelligent operation and maintenance teaching equipment includes:
[0263] A vector database acquisition module, which is used to acquire a preset vector database;
[0264] A multimodal query data acquisition module, which is used to acquire the user's multimodal query data;
[0265] A comprehensive query vector and keyword set acquisition module is used to generate a comprehensive query vector and keyword set based on the user's multimodal query data;
[0266] A multimodal knowledge fragment acquisition module is used to acquire multimodal knowledge fragments that are highly relevant to the current fault situation from a vector database based on a comprehensive query vector and a keyword set.
[0267] The comprehensive teaching response information acquisition module is used to input highly relevant multimodal knowledge fragments and user multimodal query data into a trained large model to obtain comprehensive teaching response information.
[0268] In this embodiment, the models mentioned above are not improvements; the specific training process for each model is as follows:
[0269] BiLSTM model (for semantic feature extraction):
[0270] Data preparation: Collect textual data such as fault cases and equipment manuals in the operation and maintenance field, and annotate semantic association information;
[0271] Data preprocessing: Perform word segmentation, stop word filtering, and special character removal on the text to generate a normalized text sequence;
[0272] Model training: Train a BiLSTM model based on preprocessed data, optimize model parameters to capture long-distance semantic associations in text, and output semantic feature extraction capabilities adapted to operation and maintenance scenarios.
[0273] 2. BERT+CRF model (for identifying teaching elements):
[0274] Data preparation: Collect operation and maintenance knowledge fragments containing teaching elements (principles, steps, safety tips, traceability), and manually label the element types;
[0275] Data preprocessing: Transform multimodal segments (images, audio) into text (OCR, ASR) and organize them into unified text training data;
[0276] Model fine-tuning: Based on the pre-trained BERT model, a CRF layer is added to build a sequence labeling model, which is then fine-tuned using labeled data to optimize the accuracy of teaching element recognition.
[0277] 3. BERT model (for keyword expansion):
[0278] Data preparation: Compile a terminology database for the operations and maintenance field, fault keywords, and related text corpora;
[0279] Domain adaptation: Fine-tuning the pre-trained BERT model using operational corpus to optimize the semantic understanding ability of domain terms;
[0280] Training objective: To optimize the semantic similarity between keywords and related words, ensuring that the top 5 highly relevant extended words can be selected.
[0281] 4. WordNet Semantic Network (for semantic extension):
[0282] Basic infrastructure: Import core terms from the operations and maintenance field (fault types, equipment components, etc.) and establish synonyms, near-synonyms, and hierarchical relationships;
[0283] Expert calibration: The terminology relationships are reviewed by experts in the field of operations and maintenance to correct semantic drift issues;
[0284] Data optimization: By combining historical fault diagnosis data, adjust the strength of term associations to adapt to the expanded needs of fault scenarios.
[0285] 5. Large models (such as Starry Sky Education, DeepSeek-R1, used for generating teaching responses):
[0286] Data preparation: Collect multimodal knowledge fragments, fault diagnosis cases, and operation and maintenance teaching materials (including principles, steps, and safety tips);
[0287] Model fine-tuning: Use the above data to fine-tune the pre-trained large model and optimize core capabilities such as root cause analysis, step-by-step analysis, and knowledge tracing.
[0288] Adaptation and optimization: Bind drawing / manual resource identification rules to train the model to generate structured responses (adaptation of text, schematic diagram annotations, and voice broadcast).
[0289] This application has the following advantages:
[0290] Multimodal information fusion for precise fault location: By integrating multi-dimensional information such as text, images, sound, and data, the limitations of traditional single-text descriptions are overcome. For example, the system can simultaneously analyze "error code E105," "indicator status in the startup video," and "abnormal friction sound," performing cross-validation to improve the accuracy of fault root cause location to over 90% (based on tests on a specific device dataset), far exceeding the accuracy of traditional keyword retrieval (approximately 65%) or single large-model question answering (approximately 75%).
[0291] Intelligent knowledge retrieval breaks down information silos: Semantic retrieval based on vector databases can quickly and accurately find the most relevant information fragments to the current fault situation from heterogeneous and unstructured massive knowledge bases (such as manuals, drawings, and case studies). The retrieval recall and precision are significantly higher than traditional database keyword searches. This reduces the average time for maintenance personnel to find information from 30 minutes to less than 1 minute.
[0292] Closed-loop reasoning capability reduces reliance on experts: The large model not only passively retrieves information, but also actively performs causal reasoning and logical judgment, providing a 24 / 7 online expert-level assistant, which greatly reduces reliance on senior engineers on-site.
[0293] Significantly improves teaching efficiency and learning experience: Students can interact with equipment through natural language and convenient multimedia methods, transforming abstract operation and maintenance knowledge into an intuitive and interactive exploration process, thus smoothing the learning curve. In practical training, the average time for teachers and operation and maintenance personnel to independently resolve typical faults is reduced by approximately 40%.
[0294] Strengthen knowledge traceability and systematic construction: The system-generated answers not only provide the steps, but also indicate the source of the knowledge (such as "refer to Section 3.2 of XX manual") and explain the underlying principles, guiding users to "know not only what, but also why", which helps trainees build a systematic operation and maintenance knowledge system, rather than fragmented memorization.
[0295] Although the present invention has been described in detail above with general descriptions and specific embodiments, modifications or improvements can be made to it, which will be obvious to those skilled in the art. Therefore, all such modifications or improvements made without departing from the spirit of the present invention fall within the scope of protection claimed by the present invention.
Claims
1. A multimodal large-scale model knowledge retrieval and question-answering method for fault diagnosis of intelligent operation and maintenance teaching equipment, characterized in that, The multimodal large-scale model knowledge retrieval and question-answering method for fault diagnosis of intelligent operation and maintenance teaching equipment includes: Obtain the preset vector database; Obtain the user's multimodal query data; Generate a comprehensive query vector and keyword set based on the user's multimodal query data; Based on the comprehensive query vector and keyword set, multimodal knowledge fragments highly relevant to the current fault situation are obtained from the vector database; By inputting highly relevant multimodal knowledge fragments and user multimodal query data into a trained large model, comprehensive teaching response information can be obtained.
2. The multimodal large-scale model knowledge retrieval and question-answering method for fault diagnosis of intelligent operation and maintenance teaching equipment as described in claim 1, characterized in that, The step of retrieving multimodal knowledge fragments highly relevant to the current fault situation from the vector database based on the comprehensive query vector and keyword set includes: Based on the comprehensive query vector and keyword set, generate a set of effective dimensions for fault scenarios and a confidence score table for each dimension; Contextualized weighted query vectors and unimodal normalized weight tables are generated based on the comprehensive query vector and the confidence score tables for each dimension. Based on the contextualized weighted query vector and keyword set, a semantic-keyword set bidirectional expansion retrieval and confidence filtering are performed in the vector database to obtain the effective set of bidirectional expansion retrieval. Based on the effective set of bidirectional extended retrieval and vector database, the correlation strength and complementarity of cross-modal knowledge fragments are quantified and evaluated, thereby generating a set of highly correlated cross-modal knowledge groups; Based on the set of highly correlated cross-modal knowledge groups and the confidence score table of each dimension, conflict reconciliation is carried out in a context-priority-oriented manner to obtain a set of conflict-free core knowledge groups. Based on the set of conflict-free core knowledge groups and keyword sets, a multi-dimensional weighted ranking of teaching adaptability is performed to generate a set of high-quality multimodal knowledge fragments that are highly relevant to the current fault situation.
3. The multimodal large-scale model knowledge retrieval and question-answering method for fault diagnosis of intelligent operation and maintenance teaching equipment as described in claim 2, characterized in that, The process of generating a set of effective dimensions for fault scenarios and a confidence score table for each dimension based on the comprehensive query vector and keyword set includes: Using keyword set K as the search criteria, the top-100 relevant knowledge fragments are retrieved from the vector database, and their feature vectors are extracted. and related text descriptions ; Description of associated text Semantic features are extracted using a BiLSTM model, and a candidate context dimension set is generated by combining it with the TF-IDF algorithm. ; Obtain the set of dimension domain weights and the set of historical hit rates for each dimension; Based on the feature vector Dimensional domain weight set, dimensional historical hit rate set, comprehensive query vector Keyword set The vector database generates a set of effective dimensions for fault scenarios and a confidence score table for each dimension.
4. The multimodal large-scale model knowledge retrieval and question-answering method for fault diagnosis of intelligent operation and maintenance teaching equipment as described in claim 3, characterized in that, The based on feature vector Candidate Context Dimension Set Comprehensive query vector Keyword set The vector database generates a set of effective dimensions for fault scenarios and confidence scores for each dimension, including: Based on the Top-100 relevant knowledge fragments, obtain the semantic similarity factor, dimensional text importance factor, and keyword-dimensional association factor for each knowledge fragment; The confidence level for each candidate scenario dimension is calculated using the following formula. : ; in, Confidence level for each candidate scenario dimension To integrate query vectors, For feature vectors, For the j-th candidate scenario dimension Weighting coefficients in the field of operations and maintenance For the j-th candidate context dimension, Let K be the i-th text description, and K be the set of keywords. For the j-th candidate dimension Hit rate in historical fault diagnosis cases.
5. The multimodal large-scale model knowledge retrieval and question-answering method for fault diagnosis of intelligent operation and maintenance teaching equipment as described in claim 4, characterized in that, The user's multimodal query data includes user query text, user-uploaded fault images, and user-uploaded abnormal audio. The step of generating a comprehensive query vector and keyword set based on the user's multimodal query data further includes: The user query text, user-uploaded fault images, and user-uploaded abnormal audio are encoded separately to obtain text feature vectors. Image feature vectors Abnormal audio feature vectors ; The process of generating a contextualized weighted query vector and a unimodal normalized weight table based on the comprehensive query vector and the confidence score tables for each dimension includes: Based on text feature vectors And keyword sets determine text integrity ; Based on image feature vectors And keyword set to determine image sharpness ; Based on image feature vectors And the keyword set determines the audio signal-to-noise ratio. ; Based on text integrity Image clarity and audio signal-to-noise ratio Generate modal mass vector ; Obtain the preset modality-dimension association mapping table; Define a set of modal identifiers m∈{t,i,a}, where t represents text, i represents image, and a represents audio; For each mode m, the following processing is performed: Based on the modality-dimensional association mapping table, the number of dimensions strongly associated with a particular modality in the effective dimension set of fault scenarios is counted. ; Calculate the context fit coefficient for this modality: ,in This represents the total number of valid dimensions in the fault scenario set; its value ranges from [0,1]. Integrate the context adaptation coefficients of various modalities to generate a context adaptation vector. ;in, Context adaptation coefficients for text modalities, Context adaptation coefficients for image modalities, Context adaptation coefficients for audio modalities; For each modality m, calculate the mean modality-dimension semantic association degree: Traversing the effective dimension set of fault scenarios Each dimension in The cosine similarity algorithm is used to calculate the single-mode feature vector. With dimensional feature vectors semantic similarity The value range is [−1, 1]; Combining dimensional confidence Calculate the weighted semantic relevance mean: ; Calculate the unnormalized weight for each modality based on the weighted semantic relevance mean. : ; Normalize each unnormalized weight to obtain the final normalized weight for each mode. ; Based on the final normalized weights of each mode fusion of text feature vectors Image feature vectors Abnormal audio feature vectors Generate contextualized weighted query vectors And a single-modal normalized weight table.
6. The multimodal large-scale model knowledge retrieval and question-answering method for fault diagnosis of intelligent operation and maintenance teaching equipment as described in claim 5, characterized in that, The step of performing semantic-keyword set bidirectional expansion retrieval and confidence filtering on the vector database based on contextualized weighted query vectors and keyword sets to obtain an effective bidirectional expansion retrieval set includes: Obtain the trained WordNet semantic network; Contextualized weighted query vectors are processed using a trained WordNet semantic network. Perform semantic extension to obtain a set of semantically extended vectors; Obtain the trained BERT model; For each keyword in the keyword set, input it into the trained BERT model, calculate the semantic similarity between that keyword and all terms in the operations and maintenance terminology database, and select the top 5 terms with the highest similarity as... Extended related words; The cosine similarity algorithm is used to calculate the semantic similarity between each extended related word and the original keyword, and this similarity value is used as the weight of the related words. The weight values range from [0,1]. All original keywords and their related extended terms, along with their corresponding weights, are integrated to form a set of extended keywords. ; Using all vectors in the semantic expansion vector set as search criteria, cosine similarity matching is performed in the vector database to recall knowledge fragments with a similarity ≥ 0.6 to any expansion vector, denoted as the candidate knowledge fragment set. Each fragment in the set includes three attributes: feature vector, original resource, and similarity value with each extended vector. Overall confidence score calculation: For each fragment in the candidate knowledge fragment set, substitute it into the following formula to calculate the overall confidence score. : ; in, The overall confidence level of any knowledge fragment s in the candidate knowledge fragment set; Contextualized weighted query vectors, For the feature vector of segment s, Let x be the weight value of the expansion vector at layer x. The higher the expansion layer x, the smaller the weight value. The default value is 0.95; To expand the keyword set; To expand the weight value of keyword k; This is the distance penalty coefficient, with a preset fixed value of 0.15; This is a set of effective dimensions for fault scenarios; A set of effective dimensions for fault scenarios The number of elements; From the calculated overall confidence scores, segments exceeding the overall confidence threshold are selected to form a bidirectional expanded retrieval effective set. .
7. The multimodal large-scale model knowledge retrieval and question-answering method for fault diagnosis of intelligent operation and maintenance teaching equipment as described in claim 6, characterized in that, The step of quantifying the correlation strength and evaluating the complementarity of cross-modal knowledge fragments based on the effective set of bidirectional extended retrieval and the vector database, thereby generating a set of highly correlated cross-modal knowledge groups, includes: Obtain a preset fault topic classification dictionary, which includes topic tags such as fault type, equipment component, and processing flow; Traversing the bidirectional extended search effective set For each knowledge fragment s in the data, extract the core information of the fragment, and label each fragment with one core fault topic tag according to the tags in the fault topic classification dictionary; Knowledge fragments labeled with the same core fault topic tags are grouped together to form various topic group sets. where p is the number of topic groups; For each topic group, compile all cross-modal knowledge fragments contained within the group, and denote them as follows: , where q is the number of segments in the topic group; The total number of modalities is defined as three types: text modality, image modality, and audio modality. For each topic group, the actual number of different modal types contained within the group is counted, denoted as [missing information]. ; Calculate modal coverage The range of values is ; Substitute the values into the following formula to calculate the modal complementarity coefficient Comp(g) of topic group g: ; Supplement modal complementarity coefficients for each topic group Thus, the set of topic groups G with accompanying modal complementarity coefficients is obtained; For each topic group g, perform the following calculations: Calculate the arithmetic mean of the overall confidence scores of all segments within the group. ; Iterate through any two distinct segments within the group and calculate their correlation factor. ; Calculate the overall association strength for each topic group g. ; By setting a threshold for association strength, topic groups with a comprehensive association strength greater than the threshold are selected, forming a set of highly associated cross-modal knowledge groups. .
8. The multimodal large-scale model knowledge retrieval and question-answering method for fault diagnosis of intelligent operation and maintenance teaching equipment as described in claim 7, characterized in that, The step of obtaining a conflict-free core knowledge group set by prioritizing contextual conflicts based on a highly correlated cross-modal knowledge group set and confidence score tables for each dimension includes: Obtain a fault probability database for the operations and maintenance (O&M) field, which includes the occurrence probability of different root cause types of faults. ; From the effective dimensions of fault scenarios Extract confidence level Dimensions with a value ≥0.85 are used as the core dimension set, and the core dimensions themselves constitute the core dimension set. ; Highly correlated cross-modal knowledge group set For each topic group, extract its core root cause of failure; calculate the relationship between each topic group and the core dimension set. The semantic relevance of each dimension is bound to the core dimension corresponding to each topic group; By comparing the root causes of failures across all topic groups within the same core dimension, if the descriptions of the root causes are inconsistent, these groups are categorized as conflict groups, forming a set of conflict groups. Non-conflicting groups form a temporary core set. ; Core dimensions of conflict group association Priority weights are generated by normalizing the confidence level, ensuring that the sum of the weights is 1. For the set of conflict groups For each conflict group, calculate the topic group-dimension semantic matching degree, failure probability, mean modal quality within the group, and mean redundancy coefficient within the group. The comprehensive reconciliation score for the conflict group is generated based on the topic group-dimension semantic matching degree, failure occurrence probability, mean modality quality within the group, and mean redundancy coefficient within the group. For the set of conflict groups The groups are sorted in descending order of their comprehensive harmonization scores to obtain a set of conflict-free core knowledge groups.
9. The multimodal large-scale model knowledge retrieval and question-answering method for fault diagnosis of intelligent operation and maintenance teaching equipment as described in claim 8, characterized in that, The set of conflict groups The groups are sorted in descending order of their overall harmonization scores to obtain a set of conflict-free core knowledge groups, which includes: Preset score difference threshold; If the difference between the overall reconciliation score of the highest-scoring group and the overall reconciliation score of the second-highest-scoring group is greater than or equal to the preset score difference threshold, then the highest-scoring group is retained, other conflicting groups are eliminated, and the retained group is assigned to the set of conflict-free core knowledge groups.
10. A multimodal large-scale model knowledge retrieval and question-answering device for fault diagnosis of intelligent operation and maintenance teaching equipment, characterized in that, The multimodal large-scale model knowledge retrieval and question-answering device for fault diagnosis of intelligent operation and maintenance teaching equipment includes: A vector database acquisition module, which is used to acquire a preset vector database; A multimodal query data acquisition module, which is used to acquire the user's multimodal query data; A comprehensive query vector and keyword set acquisition module is used to generate a comprehensive query vector and keyword set based on the user's multimodal query data; A multimodal knowledge fragment acquisition module is used to acquire multimodal knowledge fragments that are highly relevant to the current fault situation from a vector database based on a comprehensive query vector and a keyword set. The comprehensive teaching response information acquisition module is used to input highly relevant multimodal knowledge fragments and user multimodal query data into a trained large model to obtain comprehensive teaching response information.