Method and device for constructing fault dictionary of test system of test equipment
By using dynamic sliding windows and cross-document information entropy mutual information calculation, combined with a customized Word2Vec model, a high-quality fault dictionary suitable for experimental equipment testing systems was constructed. This solved the problem of low diagnostic efficiency in traditional methods and achieved more accurate equipment fault identification and diagnosis.
Patent Information
- Application Number
- CN202511173184.6
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-08-21
- Publication Date
- 2025-11-25
AI Technical Summary
Fault diagnosis of traditional testing equipment systems relies on human experience, which results in fragmented knowledge, weak correlation, and difficulty in accurately identifying and diagnosing equipment faults. Furthermore, existing technologies cannot effectively construct an adaptive fault dictionary.
A dynamic sliding window mechanism is used to combine word segmentation fragments, and cross-document information entropy and mutual information calculation are combined. Word vector representation is performed through a domain-customized Word2Vec model to construct a fault dictionary for the test equipment system.
It achieves high-quality, adaptive fault dictionary construction, improves the accuracy and efficiency of equipment fault diagnosis, adapts to different document and equipment types, and solves the problems of high professional requirements, high cost, and time and labor consumption in traditional methods.
Smart Images

Figure CN121009891A_ABST
Abstract
Description
Technical Field
[0001] This invention relates to the field of fault diagnosis technology, and in particular to a method and apparatus for constructing a fault dictionary for a test equipment testing system. Background Technology
[0002] For testing equipment systems, traditional equipment fault information faces challenges such as fragmented knowledge and weak correlation, making it difficult to identify and diagnose equipment faults in a timely and accurate manner. Currently, this task mainly relies on the knowledge reserves and long-term experience of staff. A large amount of information, such as equipment fault cases, recorded in text form, requires staff to repeatedly memorize and search for it. Therefore, there are prominent problems such as difficulties in collecting maintenance data, difficulty in ensuring the accuracy and standardization of fault handling, a single source of knowledge base data, and rules that are difficult to update. Meanwhile, with the increasing complexity of testing systems and the increase in the types and number of equipment, fault modes are becoming more diverse and hidden. Although intelligent technologies such as artificial intelligence, expert systems, and deep learning are being used to improve the performance of testing systems, they all depend on an accurate and comprehensive fault database or fault dictionary.
[0003] Knowledge graphs, as an emerging artificial intelligence technology, are based on the principle that constructing domain-specific dictionaries can distill information from faulty texts, transforming unstructured data into structured data, effectively improving knowledge storage and retrieval efficiency. They offer significant advantages in knowledge mining and summarization, decision generation, and assisting in fault location. For experimental equipment testing systems, the fault dictionary belongs to a specialized domain dictionary. Its construction typically involves three steps: professional vocabulary mining, word vector representation, and professional vocabulary similarity calculation. Professional vocabulary mining is the process of finding domain-specific terms not found in general dictionaries; its role is to preprocess the domain-specific text corpus through sentence segmentation and word segmentation. Word vector representation refers to converting the mined terms into numerical forms for easy representation and application in computation. The role of professional vocabulary similarity calculation is that the trained word vectors contain the inherent connections between domain-specific terms; that is, the degree of association between two words can be determined by calculating the word vector similarity.
[0004] Chinese patent application CN115238029A discloses a method and apparatus for constructing a power fault knowledge graph. The method includes the following steps: Step 1, acquiring data to be processed: acquiring preprocessed text data of power faults; Step 2, performing data preprocessing; Step 3, using a BERT-BiLSTM-CRF combined model to extract entities from the preprocessed data; Step 4, using a dependency parsing-based method to identify and extract relationships between entities, and analyzing the dependency relationships between sentence components by identifying and locating syntactic relationships; Step 5, knowledge storage and semantic triple representation; Step 6, constructing a power fault knowledge graph.
[0005] Chinese patent application CN118193745A discloses a method for constructing a standard knowledge graph and a search system for the field of pressure equipment. After segmenting the original data file, it uses sliding windows of different sizes to combine the segmented fragments to construct candidate words. Then, it uses a screening index that considers both adjacency information entropy and enhanced mutual information to filter the candidate words, thus constructing a domain-specific dictionary. This dictionary serves as the basis for segmenting the original data file, allowing for the extraction of common entities related to risk prevention and control. However, CN118193745A uses a discrete multi-size window combination of candidate words, which is essentially still a superposition of fixed windows and cannot adapt to the semantic coherence of the context.
[0006] The paper "An Improved New Word Synthesis Algorithm Based on Multi-Character Mutual Information and Adjacency Entropy" (by Wang Xin) addresses the issue of new words that have been incorrectly segmented into multiple words by word segmentation tools. It combines multi-character mutual information, left and right adjacency entropy, and the improved new word synthesis algorithm in this paper to merge adjacent words and obtain candidate new words. Then, it uses a filtering method that combines low-frequency words, first and last stop words, word formation rules, and common dictionary rules with statistics to remove strings that do not meet the requirements, and finally obtains a new word set. Summary of the Invention
[0007] To address the aforementioned issues, this invention proposes a method and apparatus for constructing a fault dictionary for experimental equipment testing systems. The domain embedding dictionary obtained by this invention contains the relative positions of specialized terms in the vector space, which can provide richer vocabulary information in named entity recognition of faults in experimental equipment testing systems to improve recognition performance.
[0008] The technical solution adopted in this invention is as follows: A method for constructing a fault dictionary for a test equipment testing system includes: Data related to the failure of the test equipment system is collected and preprocessed to obtain fine-grained word segments; based on a dynamic sliding window mechanism, the word segments are combined and spliced, and all the combined words are used as a candidate vocabulary set. The candidate vocabulary set is filtered by word frequency. The information entropy and mutual information of the filtered candidate vocabulary are calculated based on the multi-source fault text. The candidate vocabulary is then ranked based on the information entropy and mutual information to obtain a professional vocabulary set. Based on a set of professional vocabulary, word vectors are represented by a domain-customized pre-trained word vector model. Professional vocabulary is encoded into continuous dense vectors of fixed length, and the similarity between professional vocabulary is calculated. In this way, a fault dictionary for the test equipment testing system is established.
[0009] Furthermore, the method of combining and splicing the word segments based on the dynamic sliding window mechanism includes: adaptively adjusting the window size according to the context information of the fault description and the text length, and then combining and splicing the word segments to achieve non-continuous but semantically related word mining.
[0010] Furthermore, the step of filtering the candidate word set by word frequency includes: setting a preset word frequency threshold, filtering candidate words in the candidate word set, and eliminating irrelevant sparse words.
[0011] Furthermore, the calculation of information entropy and mutual information of candidate words after screening based on multi-source fault texts includes: based on information theory principles, quantitatively evaluating two key features of candidate words by analyzing statistical data from multiple fault documents: First, context boundary features are determined based on information entropy to measure cross-document boundary uncertainty. Secondly, internal structural features are determined based on mutual information to measure the solidification of internal structure across documents.
[0012] Furthermore, the cross-document boundary uncertainty measure includes: calculating cross-document information entropy. H D The cross-document information entropy H D Characterization candidate words in multi-source fault text sets D Boundary independence in; the cross-document information entropy H D Including left and right information entropy, when the values of left and right information entropy increase, it indicates that the candidate words have higher boundary independence in cross-document context.
[0013] Furthermore, the cross-document internal structure cohesion measurement includes: calculating cross-document mutual information based on the word-formation characteristics of specialized vocabulary through binary and ternary compound structures. MI D The cross-document mutual information MI D Characterizes the strength of association between components within candidate words.
[0014] Furthermore, the step of encoding professional vocabulary into a fixed-length continuous dense vector set by using a domain-customized pre-trained word vector model to represent word vectors based on a professional vocabulary set includes: performing distributed encoding through a Word2Vec model trained specifically for this purpose, and learning the vector representation of words through training a neural network to obtain a domain-embedded dictionary.
[0015] Furthermore, the method of learning the vector representation of words by training a neural network includes: learning the vector representation of words based on a composite structure of a CBOW model and a Skip-Gram model, wherein the context word vector of the target word is input into the CBOW model to infer the target word vector, and the inferred target word vector is input into the Skip-Gram model to infer the context word vector of the target word.
[0016] Furthermore, the calculation of the similarity between specialized terms includes: calculating the similarity between all words in the domain embedding dictionary and the target word using cosine similarity, selecting a number of words with the highest cosine similarity and returning them for viewing similar words of the target word.
[0017] A fault dictionary construction device for a test equipment testing system, comprising: The vocabulary mining module is configured to collect fault-related data from the test system of the experimental equipment, obtain fine-grained word segmentation fragments through preprocessing, and combine and splice the word segmentation fragments based on a dynamic sliding window mechanism, and use all the combined words as a candidate vocabulary set. The cross-document calculation module is configured to perform word frequency filtering on the candidate vocabulary set, calculate the information entropy and mutual information of the filtered candidate words based on multi-source fault text, and sort the candidate words based on the information entropy and mutual information to obtain a professional vocabulary set. The fault dictionary construction module is configured to use a domain-customized pre-trained word vector model to represent words based on a set of professional vocabulary, encode professional vocabulary into a continuous dense vector set of fixed length, and calculate the similarity between professional vocabulary, thereby establishing a fault dictionary for the test equipment testing system.
[0018] The beneficial effects of this invention are as follows: (1) Improvement and application of dynamic sliding windows Currently, equipment fault diagnosis in diagnostic testing systems often relies on the traditional method of manual input by staff, which is inefficient and presents significant problems such as high professional requirements, high costs, and time-consuming processes. Traditional knowledge graph construction methods use a fixed-size sliding window to piece together segmented fragments, but this method does not consider the text length and contextual changes in the fault description. This invention innovatively employs a dynamic sliding window, adjusting according to the contextual information and text length of the fault description, thus adapting to fault texts of varying lengths and solving the problem of mining non-continuous but semantically related words. For brief equipment fault reports, the sliding window may be small; however, for complex test logs, the window size can be increased accordingly to extract more contextual information.
[0019] (2) Improved cross-document computation of information entropy and mutual information Traditional information entropy and mutual information calculations are typically limited to single documents. However, this invention innovatively introduces cross-document information entropy and mutual information calculations. Specifically, it calculates the entropy and mutual information of words based on multi-source fault text, considering the contextual information of the same term in different documents, thus enhancing the accuracy and stability of word selection. This cross-document calculation method is applicable to processing test documents from different devices and of different types, enabling more precise filtering and sorting of specialized terms appearing in different documents.
[0020] (3) Further optimization of word vector representation and professional vocabulary To better adapt to the specialized terminology used in experimental equipment testing systems, this invention employs a customized trained Word2Vec model to generate word vectors for domain-specific vocabulary. This method effectively addresses the challenges faced by traditional methods in experimental equipment testing, such as polysemy and the imprecision of domain-specific terminology.
[0021] (4) Compared to CN115238029A, this invention creatively solves the fundamental and crucial prerequisite problem of automatically constructing a high-quality, standardized, and domain-adaptable core fault terminology database (dictionary) for experimental equipment testing systems. The dictionary construction process utilizes innovative techniques such as dynamic sliding windows, cross-document statistical calculations (information entropy, mutual information), and domain-customized word vectors, effectively overcoming the shortcomings of existing technologies in areas such as automated terminology mining, cross-document consistency, domain semantic representation accuracy, and dependence on labeled data. This high-quality fault dictionary not only has independent value but also lays a crucial foundation for the subsequent construction of a more accurate and robust fault knowledge graph for experimental equipment testing systems.
[0022] (5) Compared to CN118193745A, which uses discrete multi-size windows to combine candidate words, it is essentially still a superposition of fixed windows and cannot adapt to the semantic coherence of the context. The dynamic sliding window of this invention dynamically adjusts the window boundary based on the real-time text length and semantic relevance, which can solve the problem of the traditional window's fragmentation of continuous semantic units and realize the complete semantic encapsulation of fault description. At the same time, this solution only works on a single document and does not solve the terminology consistency of multi-source texts. This invention calculates and aggregates the statistical features of multi-source fault texts (maintenance manuals, test logs, cross-device reports) across documents, removes context-sensitive fluctuating terms through cross-document information entropy, and locks common core words in the domain through cross-document mutual information.
[0023] (6) Compared with the literature "An Improved Novel Word Synthesis Algorithm Based on Multi-Word Mutual Information and Adjacency Entropy", this invention breaks through the limitation of passively merging adjacent word segmentation errors in the literature at the candidate word generation level. It actively scans the original fragment stream through a dynamic sliding window and captures non-continuous semantic units across multiple word segmentation boundaries based on semantic relevance, achieving a revolutionary leap from mechanical error correction to continuous semantic complete extraction. At the terminology screening level, it abandons the fragile filtering mechanism of literature relying on manual rules and single-document statistics, and innovatively introduces cross-document mutual information aggregation and information entropy convergence analysis. It eliminates document-specific pseudo-terms and locks in the universal core vocabulary of the domain through the statistical evidence chain of multi-source fault texts, and constructs a verifiable consensus terminology set with cross-scenario stability. Attached Figure Description
[0024] Figure 1 This is a schematic diagram of a method for constructing a fault dictionary for a test equipment testing system according to Embodiment 1 of the present invention.
[0025] Figure 2 This is a schematic diagram of the candidate vocabulary set in Embodiment 1 of the present invention.
[0026] Figure 3 This is a schematic diagram of the Word2Vec model based on the composite structure of CBOW and Skip-Gram in Embodiment 1 of the present invention.
[0027] Figure 4 This is a diagram illustrating the word vector training and implementation process based on the Word2Vec model in Embodiment 1 of the present invention.
[0028] Figure 5 This is a schematic diagram of some word clusters obtained from the test in Embodiment 1 of the present invention, wherein sub-figure (a) is a word cluster related to "fault" and sub-figure (b) is a word cluster related to "temperature sensor". Detailed Implementation
[0029] To provide a clearer understanding of the technical features, objectives, and effects of the present invention, specific embodiments are now described. It should be understood that the specific embodiments described herein are merely illustrative of the invention and are not intended to limit the invention; that is, the described embodiments are only a part of the embodiments of the invention, not all of them. All other embodiments obtained by those skilled in the art based on the embodiments of the present invention without inventive effort are within the scope of protection of the present invention.
[0030] Example 1 like Figure 1 As shown, this embodiment provides a method for constructing a fault dictionary for a test equipment testing system, including: Data related to the failure of the test equipment system is collected and preprocessed to obtain fine-grained word segments. Based on the dynamic sliding window mechanism, the word segments are combined and spliced to solve the problem of mining non-contiguous but semantically related words, and all the combined words are used as a candidate word set. The candidate vocabulary set is filtered by word frequency. The information entropy and mutual information of the filtered candidate vocabulary are calculated based on the multi-source fault text. The candidate vocabulary is then ranked based on the information entropy and mutual information to obtain a professional vocabulary set. Based on a set of professional vocabulary, word vectors are represented by a domain-customized pre-trained word vector model. Professional vocabulary is encoded into a continuous dense vector set of fixed length, and the similarity between professional vocabulary is calculated. In this way, a fault dictionary for the test equipment testing system is established.
[0031] Specifically, the method for constructing the fault dictionary of the test equipment testing system in this embodiment can be implemented in the following way.
[0032] I. Terminology Mining: Dynamic Windows and Cross-Document Computation This embodiment takes fault-related data of the test equipment testing system as the research object, such as fault repair manuals, logs and other related materials in various text formats such as Excel spreadsheets, docx, and pdf. First, it mines professional terms for the test equipment testing system and proposes a professional term mining method based on information theory, as shown in Table 1.
[0033] Table 1 - Information Theory-Based Methods for Vocabulary Mining
[0034] The implementation methods of the above-mentioned dynamic sliding window mechanism include: 1) Initialize the sliding window range. This range includes the maximum and minimum values.
[0035] 2) Adjust the sliding window size based on text length. If the text length is within a certain threshold, adjust the maximum value of the sliding window to match that threshold.
[0036] 3) Adjust the sliding window size based on the domain dictionary matching degree. If the number of matched professional terms reaches a certain amount, the maximum value of the sliding window is increased.
[0037] 4) Adjust the sliding window size based on word repetition. If word repetition reaches a certain threshold, reduce the maximum value of the sliding window.
[0038] Specifically, the above method can use the following sliding window function to process the segmented word fragments:
[0039] In addition, the above-mentioned methods for combining and splicing word segments include: 1) Semantic association combination based on domain dictionary. If two non-contiguous words are both in the domain dictionary, then the two words are combined.
[0040] 2) Combinations based on predefined semantic patterns. If two consecutive words form a "noun + verb" structure, then the two words are combined.
[0041] The above method can be processed using the following combination function:
[0042] This embodiment performs initial filtering of candidate words based on a preset word frequency threshold (e.g., ≥2), effectively eliminating irrelevant sparse words. Subsequently, based on information theory principles, statistical data from multiple faulty documents is analyzed to quantify and evaluate two key features of the candidate words: first, contextual boundary features are determined based on information entropy for cross-document boundary uncertainty measurement; second, internal structural features are determined based on mutual information for cross-document internal structural solidification measurement. An example of the statistical data for candidate words is shown below. Figure 2 As shown.
[0043] (1) Measurement of uncertainty across document boundaries Define cross-document information entropy H D Used to characterize candidate words in a document set D Boundary independence in documents, cross-document information entropy H D It is divided into left information entropy and right information entropy.
[0044] Left information entropy H D (left|word): Calculates the uncertainty of the left-side neighboring words of a candidate word. (1) in, L D (word) represents a collection of documents. D The set of words adjacent to the left of the candidate word. P D Based on D The probability distribution, w left This represents the set of all the left-hand adjacent words of the candidate word "word" that actually appear in the document set D, where "word" represents the candidate word selected from the document set D through a dynamic sliding window.
[0045] Right information entropy H D(right|word): Calculate the uncertainty of the right-hand neighbor of a candidate word: (2) in, R D (word) is a collection of documents D The set of words adjacent to the right of a candidate word, w right This represents the set of all right-side adjacent words that actually appear in the document set D for the candidate word "word".
[0046] when H D (left|word) and H D When the value of (right|word) increases, it indicates that the candidate word has higher boundary independence in cross-document context, and the probability of it being identified as an independent semantic unit is significantly increased.
[0047] (2) Cross-document internal structural cohesion measurement Cross-document mutual information computation sets up binary and trigram compound structures based on the word-formation characteristics of specialized terms, and defines cross-document mutual information. MI D Used to quantify the correlation strength of components within candidate words: Binary structure: (3) in, MI D Used to quantify the correlation strength of components within candidate words. P D Based on D The probability distribution; D It is a collection of documents, an aggregated dataset containing all fault documents (maintenance manuals, test logs, etc.) of the high-altitude test platform system; x,y These are adjacent text segments within the candidate words; P D (x) and P D (y) represents the segments x and y exist D The probability of it appearing independently in the middle; P D ( x,y () is a fragment x and y exist D The probability of consecutive adjacent co-occurrences.
[0048] Cross-document mutual information calculation sets up binary and ternary compound structures for the word formation characteristics of professional terms: In the binary structure design, based on the linguistic features of the common double word combination pattern in domain professional terms (such as "fuel leak" and "sensor failure"), the intrinsic correlation between adjacent segments in cross-document context is quantified by formula (3); when the calculation result is significantly greater than zero, it indicates that the co-occurrence probability of segment combination exceeds random expectations and constitutes basic professional terms; for the high-frequency three-word long terms in the test equipment testing system (such as "turbine blade crack" and "inlet guide deformation").
[0049] Ternary structure: (4) in, x , y,z Three consecutive text segments within a candidate word; P D ( x,y,z ): fragment x,y,z exist D The probability of consecutive adjacent occurrences in a given context; P D (x), P D (y), P D (z): fragment x, y, z exist D The probability of it appearing independently in the middle; : Segmented aggregation mutual information value of ternary candidate words; MI D (x,y): Fragment x and y Binary mutual information (first paragraph association); MI D (y,z): Fragment y and z Binary mutual information (tail segment association); MI D (x,z): Fragment x and z Binary mutual information (cross-segment association).
[0050] The ternary structure provides two equivalent computational paths through formula (4): one is to directly capture complete semantic units using overall cohesion calculation, which is suitable for high-frequency candidate words; the other is to calculate the mean of the first and last segments and cross-segment associations separately through segmented cohesion aggregation. This innovative design effectively overcomes the data sparsity problem of low-frequency terms. MI D When the value increases, it indicates that the internal components of the candidate words have a strong correlation in cross-document statistics, and the probability of them constituting domain-specific vocabulary increases significantly.
[0051] in, P D (x) P D (x,y) P D The equal probability values (x, y, z) are all based on the document collection. D Overall statistics show that: ① The robustness of vocabulary boundary determination is enhanced, overcoming misjudgment caused by insufficient single-document samples; ② Improved accuracy of internal structure analysis, especially suitable for multi-source heterogeneous fault documents (such as scenarios where maintenance manuals and test logs coexist); ③ The stability of the screening results has been significantly improved, adapting to the terminology mining needs of different device types and document formats.
[0052] II. Word Vector Representation Methods The mined vocabulary needs to be further converted into numerical form. Given the characteristics of the testing equipment system—a high density of specialized terminology with domain-specific meanings and multiple meanings—this embodiment employs a customized Word2Vec model for distributed encoding. During customized training, a general Word2Vec model is first used as a foundation, and then retrained using specialized data from testing equipment fault documents (such as maintenance manuals and test logs). Simultaneously, through reinforcement training, the Word2Vec model memorizes the collocation patterns of specialized terms. The model adjusts the vector representations between corresponding terms, bringing them closer together in the digital space, thereby accurately capturing the semantic relationships between words within the domain.
[0053] The vector representation of words is learned by training a neural network. Figure 3 The model shown is based on a composite structure of CBOW and Skip-Gram. In this model, the context word vector of the target word is input into the CBOW model to infer the target word vector. The inferred target word vector is then input into the Skip-Gram model to infer the context word vector of the target word.
[0054] Specifically, the CBOW model infers target words from contextual vocabulary; Input: The context word vector of the target word (e.g., 1-2 words to the left and right of the target word); Output: Predicted target word vectors (e.g., "calibrate").
[0055] Skip-Gram model: Infers context words from target words; Input: The target vocabulary vector (e.g., “calibrate”) output by the CBOW model.
[0056] Output: Predict the context word vector of the target word (e.g., ["sensor", "parameter"]).
[0057] The target word is located from the context words using the CBOW model, and then the Skip-Gram model is used to verify the contextual consistency of the target word, ensuring that the vector distance of "sensor" in the fault domain is closer to words such as "calibration", "fault", and "parameter", rather than general device words.
[0058] Specifically, to accurately capture the semantic relationships and contextual patterns unique to the testing equipment system domain, this embodiment employs a customized word vector model. Specifically, the model is pre-trained or fine-tuned using collected documents specific to the testing equipment system (including fault manuals, logs, etc.). In terms of model structure, a customized word vector model can be adopted. Figure 3 The diagram illustrates a composite structure based on CBOW and Skip-Gram, or other effective structures (such as GloVe, fastText). The context words (vectors) of the target word are input into the CBOW model, and the output target word vector is fed into the Skip-Gram model to infer context word vectors. Figure 4 As shown. The core objective of training is to maximize the correlation between the target word and its context words; after the model is trained, the weights of its hidden layers (model parameters) serve as the word vector representation of the target word.
[0059] This deep domain-customized word vector representation method can effectively overcome the challenges faced by traditional general word vectors in this professional field, significantly alleviate the problems of polysemy (e.g., the same word has different meanings in general contexts and aero-engine testing) and the imprecise expression of professional terms in the field, thereby greatly improving the accuracy of subsequent word similarity calculation and knowledge graph construction.
[0060] III. Calculation of Similarity of Professional Terms For the word vectors of the transformed test equipment testing system, this embodiment uses the cosine similarity method to calculate the word similarity, as shown in Equation (5), to ensure the comprehensiveness and accuracy of the fault dictionary construction. That is, the similarity between two word vectors is measured by measuring the cosine value of the angle between them, which is in the range of [-1,1]. Here, -1 means that the two vectors point in opposite directions, which means that the meanings of the two corresponding words are also opposite; 1 means that their directions are exactly the same, and the meanings of the two words are also the same; 0 means that they are independent of each other, and the meanings of the two words are also unrelated.
[0061] By calculating the cosine similarity between all words in the domain-embedded dictionary and the target word, the words with the highest similarity can be selected as similar words to the target word. This helps to build a more comprehensive and relevant fault dictionary and provides a semantic foundation for fault diagnosis and maintenance suggestions based on knowledge graphs. Thanks to customized word vectors, this similarity calculation can more accurately reflect the professional semantic relationships in the testing field of experimental equipment.
[0062] (5) In the formula, A · B Representing vectors A sum vector B dot product, || A || represents a vector A The norm, | B || represents a vector B The norm of the word is then used. Furthermore, by calculating the similarity between all words in the domain embedding dictionary and the target word using cosine similarity, the words with the highest cosine similarity are selected and returned, allowing you to view the similar words of the target word.
[0063] To verify the correctness and effectiveness of the method in this embodiment, the content of books such as "High-altitude Simulation of Aircraft Engines" and "Aircraft Engine Testing and Experimentation Technology" is used as the research object. The Jieba word segmentation tool is used to extract text to construct a dataset, and after preprocessing such as sentence segmentation and data cleaning, the domain corpus used for the experiment is obtained.
[0064] Table 2 shows the vocabulary mined from the test equipment system faults using the information theory-based professional vocabulary mining method in Table 1. It can be seen that the higher the score ranking, the more accurate the vocabulary is. As the score of the vocabulary gradually decreases, the number of erroneous vocabulary mined will increase accordingly. Here, Table 2 is the domain corpus used for the experiment, which is determined based on a word frequency of 2.
[0065] Table 2 - Vocabulary for Fault Discovery in Test Equipment and Testing Systems
[0066] Based on the vocabulary scores, approximately the top 50 sets of vocabulary mining results were selected, and combined with manual screening to obtain relevant professional terms. Furthermore, all dictionary packages related to "high-altitude testing" were searched and downloaded from online thesaurus, converted to text format, and merged into a single file containing approximately 4000 test equipment-related terms, serving as the downloaded dictionary. Then, the screened vocabulary mining results were merged with the downloaded test equipment-related dictionary after deduplication to obtain approximately 5000 terms, collectively constructing a fault domain professional thesaurus. This embodiment's method can obtain professional terms not found in the downloaded dictionary; some examples are shown in Table 3.
[0067] Table 3 - A partial list of technical terms for testing equipment and system fault diagnosis.
[0068] The observations show that most of the domain-specific terms found in the vocabulary mining results but not in existing dictionaries are related to faults, operations, and parts. This is because this experiment used texts related to specific fault handling, and existing dictionaries rarely contain these three types of vocabulary. In addition, there are a few equipment-related terms. Although most of these terms already appear in existing dictionaries, this method can use statistical indicators to filter out specific-purpose or specific-type equipment terms as a supplement, such as "speed-regulating asynchronous motor," "explosion-proof motor," and "electric hoist crane."
[0069] Table 4 shows a comparison between the results of Jieba's default word segmentation method and the results after adding a domain-specific thesaurus as a custom dictionary. The results show that the word segmentation without the thesaurus is more granular, resulting in some professional terms being incorrectly segmented, such as "normal lifting," "aerial work platform," and "hydraulic system." Importing the thesaurus into Jieba's custom dictionary effectively reduces word segmentation errors in domain-specific texts.
[0070] Table 4 - Professional vocabulary obtained using the word segmentation tool Jieba
[0071] The Word2Vec model from the Gensim toolkit was used to train the segmented word sequence, with a vector dimension of 50, resulting in distributed representations of approximately 1000 domain-specific words. Some of the domain-specific words and their word vectors are shown in Table 5. These domain-specific words and their corresponding distributed vectors are used as a domain embedding dictionary, which can be applied to named entity recognition tasks to provide word sequence information for more accurate entity identification.
[0072] Table 5 - Word segmentation vectors trained based on the Word2Vec model
[0073] Next, this embodiment uses cosine similarity to calculate the similarity between word vectors, and selects the 10 words closest to the center word as its similar words to form a word cluster. Figure 5This is a schematic diagram of word clusters with similar words. Taking "fault" as an example, the cosine similarity results in "equipment fault," "sensor malfunction," "data loss," "signal interference," "environmental factors," "circuit fault," "structural damage," "detection," "fault diagnosis," and "short circuit fault," representing specific fault situations; "fault cause," "inspection," and "fault diagnosis" represent fault analysis and processing. The domain embedding dictionary obtained through the method in this embodiment contains the relative positions of professional terms in the vector space, which can provide richer lexical information in named entity recognition in the fault domain of experimental equipment testing systems to improve recognition performance.
[0074] Example 2 This embodiment provides a fault dictionary construction device for experimental equipment testing systems, including: The vocabulary mining module is configured to collect fault-related data from the test system of the experimental equipment, and obtain fine-grained word segmentation fragments through preprocessing; based on the dynamic sliding window mechanism, the word segmentation fragments are combined and spliced, and all the combined words are used as a candidate vocabulary set; The cross-document calculation module is configured to perform word frequency filtering on the candidate vocabulary set, calculate the information entropy and mutual information of the filtered candidate words based on multi-source fault text, and sort the candidate words based on the information entropy and mutual information to obtain a professional vocabulary set. The fault dictionary construction module is configured to use a domain-customized pre-trained word vector model to represent words based on a set of professional vocabulary, encode professional vocabulary into a continuous dense vector set of fixed length, and calculate the similarity between professional vocabulary, thereby establishing a fault dictionary for the test equipment testing system.
[0075] Example 3 This embodiment is based on embodiment 1: This embodiment provides a computer device, including a memory and a processor. The memory stores a computer program, and the processor executes the computer program to implement the steps of the test equipment test system fault dictionary construction method of Embodiment 1. The computer program can be in the form of source code, object code, executable file, or some intermediate form.
[0076] Example 4 This embodiment is based on embodiment 1: This embodiment provides a computer-readable storage medium storing a computer program that, when executed by a processor, implements the steps of the fault dictionary construction method for the test equipment testing system of Embodiment 1. The computer program can be in the form of source code, object code, executable file, or some intermediate form. The storage medium includes any entity or device capable of carrying computer program code, recording media, computer memory, read-only memory (ROM), random access memory (RAM), electrical carrier signals, telecommunication signals, and software distribution media, etc. It should be noted that the content contained in the storage medium can be appropriately added or removed according to the requirements of legislation and patent practice in the jurisdiction. For example, in some jurisdictions, according to legislation and patent practice, the storage medium does not include electrical carrier signals and telecommunication signals.
[0077] It should be noted that, for the sake of simplicity, the foregoing method embodiments are described as a series of actions. However, those skilled in the art should understand that this application is not limited to the described order of actions, as some steps may be performed in other orders or simultaneously according to this application. Furthermore, those skilled in the art should also understand that the embodiments described in the specification are preferred embodiments, and the actions and modules involved are not necessarily essential to this application.
Claims
1. A method for constructing a fault dictionary for a test equipment testing system, characterized in that, include: Collect fault-related data from the test equipment system and obtain fine-grained word segmentation fragments through preprocessing; Based on a dynamic sliding window mechanism, the word segments are combined and spliced to generate a candidate word set; The candidate vocabulary set is filtered by word frequency. The information entropy and mutual information of the filtered candidate vocabulary are calculated based on the multi-source fault text. The candidate vocabulary is then ranked based on the information entropy and mutual information to obtain a professional vocabulary set. Based on a set of professional vocabulary, word vectors are represented by a domain-customized pre-trained word vector model. Professional vocabulary is encoded into continuous dense vectors of fixed length, and the similarity between professional vocabulary is calculated. In this way, a fault dictionary for the test equipment testing system is established.
2. The method for constructing a fault dictionary for a test equipment testing system according to claim 1, characterized in that, The dynamic sliding window mechanism for combining and splicing word segments includes: adaptively adjusting the window size based on the context information and text length of the fault description, and then combining and splicing the word segments to achieve non-continuous but semantically related word mining.
3. The method for constructing a fault dictionary for a test equipment testing system according to claim 1, characterized in that, The step of filtering the candidate vocabulary set by word frequency includes: setting a preset word frequency threshold, filtering candidate words in the candidate vocabulary set, and eliminating irrelevant sparse words.
4. The method for constructing a fault dictionary for a test equipment testing system according to claim 1, characterized in that, The calculation of information entropy and mutual information of candidate words after screening based on multi-source fault texts includes: based on information theory principles, quantitatively evaluating two key features of candidate words by analyzing statistical data from multiple fault documents: First, context boundary features are determined based on information entropy to measure cross-document boundary uncertainty. Secondly, internal structural features are determined based on mutual information to measure the solidification of internal structure across documents.
5. The method for constructing a fault dictionary for a test equipment testing system according to claim 4, characterized in that, The cross-document boundary uncertainty measure includes: calculating cross-document information entropy. H D The cross-document information entropy H D Characterization candidate words in multi-source fault text sets D Boundary independence in; the cross-document information entropy H D Including left and right information entropy, when the values of left and right information entropy increase, it indicates that the candidate words have higher boundary independence in cross-document context.
6. The method for constructing a fault dictionary for a test equipment testing system according to claim 4, characterized in that, The cross-document internal structure cohesion metric includes: calculating cross-document mutual information based on the word-formation characteristics of specialized vocabulary through binary and ternary compound structures. MI D The cross-document mutual information MI D Characterizes the strength of association between components within candidate words.
7. The method for constructing a fault dictionary for a test equipment testing system according to claim 1, characterized in that, The method of encoding professional vocabulary into a fixed-length continuous dense vector set by using a domain-customized pre-trained word vector model to represent word vectors includes: performing distributed encoding through a targeted Word2Vec model, and learning the vector representation of words through training a neural network to obtain a domain-embedded dictionary.
8. The method for constructing a fault dictionary for a test equipment testing system according to claim 7, characterized in that, The method of learning the vector representation of words by training a neural network includes: learning the vector representation of words based on a composite structure of the CBOW model and the Skip-Gram model, wherein the context word vector of the target word is input into the CBOW model to infer the target word vector, and the inferred target word vector is input into the Skip-Gram model to infer the context word vector of the target word.
9. The method for constructing a fault dictionary for a test equipment testing system according to claim 7, characterized in that, The calculation of similarity between specialized terms includes: calculating the similarity between all words in the domain embedding dictionary and the target word using cosine similarity, selecting a number of words with the highest cosine similarity and returning them for viewing similar words of the target word.
10. A fault dictionary construction device for a test equipment testing system, characterized in that, include: The vocabulary mining module is configured to collect fault-related data from the test system of the experimental equipment and obtain fine-grained word segmentation fragments through preprocessing. Based on a dynamic sliding window mechanism, the word segments are combined and spliced to generate a candidate word set; The cross-document calculation module is configured to perform word frequency filtering on the candidate vocabulary set, calculate the information entropy and mutual information of the filtered candidate words based on multi-source fault text, and sort the candidate words based on the information entropy and mutual information to obtain a professional vocabulary set. The fault dictionary construction module is configured to use a domain-customized pre-trained word vector model to represent words based on a set of professional vocabulary, encode professional vocabulary into a continuous dense vector set of fixed length, and calculate the similarity between professional vocabulary, thereby establishing a fault dictionary for the test equipment testing system.
Citation Information
Patent Citations
Method and device for constructing power failure knowledge graph
CN115238029A
Construction method and search system for standard knowledge graph in field of pressure-bearing equipment
CN118193745A