Digital power distribution network equipment fault diagnosis method and system based on artificial intelligence voice recognition

By converting the voice signal into a Mel spectrum matrix and inputting it into a large language model, and combining the conditional probability model to generate a fault report in the power industry, the efficiency and accuracy of fault diagnosis in the digital distribution network are solved, and the operation safety of the power system and the work efficiency of grassroots personnel are improved.

CN120356486AInactive Publication Date: 2025-07-22NARI INFORMATION & COMM TECH
View PDF 7 Cites 0 Cited by

Patent Information

Application Number
CN202510867069.2
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-06-26
Publication Date
2025-07-22
Estimated Expiration
Not applicable · inactive patent

AI Technical Summary

Technical Problem

In the digital distribution network environment, the existing technology cannot effectively integrate voice input, standard fault library search and text generation capabilities of large language models, and cannot automatically generate structured and professional equipment fault reports, resulting in low efficiency and insufficient accuracy in power system fault diagnosis.

Method used

By obtaining the voice signal of equipment failure information, it is converted into a logarithmic Mel spectrum matrix, inputting a large language model for speech recognition, extracting keywords and entity recognition, using the conditional probability model to determine the fault type, and generating a fault report in the power industry standard.

Benefits of technology

It realizes fast and accurate equipment fault diagnosis, generates standardized fault reports, and improves the operation safety of the power system and the work efficiency of grassroots personnel.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120356486A_ABST
    Figure CN120356486A_ABST
Patent Text Reader

Abstract

The invention discloses a digital power distribution network equipment fault diagnosis method and system based on artificial intelligence voice recognition. The method comprises the following steps: acquiring a sound signal describing equipment fault information; converting the sound signal describing the fault information into a logarithm Mel frequency spectrum matrix; inputting the logarithm Mel frequency spectrum matrix into an encoder of a large language model, outputting a time sequence embedding sequence by the encoder of the large language model, and obtaining a speech recognition text by the time sequence embedding sequence through a corresponding text decoder; keywords related to the equipment fault are extracted from the voice recognition text, entity recognition is carried out on each keyword, and entities comprise an equipment name entity, a fault symptom entity and an environment condition entity; and adopting a conditional probability model to judge the correlation degree between the identified different entities and the fault types, and taking the fault type with the maximum fault probability as a diagnosis result.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention belongs to the field of equipment fault diagnosis, and specifically relates to a method and system for fault diagnosis of digital distribution network equipment based on artificial intelligence speech recognition. Background Art

[0002] With the continuous advancement of the digitalization and intelligentization process in the power industry, the types and quantities of equipment within the distribution network have increased significantly. The diagnosis and disposal of equipment faults have become a key link in ensuring the safe and stable operation of the power system. Traditional fault diagnosis methods mainly rely on the experience judgment of on-site professionals or troubleshooting according to established manuals. However, in the face of diverse equipment and complex fault environments, relying solely on manual analysis often takes a long time and places high requirements on the professional capabilities of operators.

[0003] At the same time, when grass-roots workers in the power industry perform maintenance or inspection tasks, they often need to operate equipment while obtaining fault information, making it very inconvenient to input or query data. If the fault details and diagnosis information cannot be understood in a timely and accurate manner, it will not only increase the fault handling time but also potentially affect the overall safe operation of the power system. Especially in the digital distribution network scenario, where there are numerous types of equipment faults, scattered substations, and a large amount of information, the traditional input and retrieval processes are difficult to meet the dual requirements of real-time and accuracy.

[0004] In recent years, the rapid development in the field of artificial intelligence, especially the technological breakthroughs represented by large language models, has provided new solutions for power equipment fault diagnosis. By combining speech recognition, multimodal retrieval, and enhanced generation technologies, two-way conversion between speech and text information can be achieved, greatly simplifying the operation process for front-line personnel. At the same time, by associatively retrieving the professional equipment fault database in the power industry with large language models, the fault point can be quickly located and detailed information can be extracted, making on-site diagnosis based on voice input possible. However, currently, most artificial intelligence applications for fault diagnosis of distribution network equipment are still limited to a single text input method, failing to fully utilize the advantages of multimodal retrieval and lacking in-depth integration of professional knowledge in the power industry. In addition, how to combine the retrieved fault information with the natural language generation ability of large language models to form a formatted, highly readable, and professional fault analysis report is another major challenge currently faced. Traditional large model training often requires a huge investment in data and computing power, while for power industry application scenarios, quickly and accurately producing diagnosis results is more practically significant than blindly pursuing large model scale. Only by taking into account resource consumption and generation quality and fully utilizing artificial intelligence retrieval enhanced generation technology can we truly empower grass-roots personnel in the power industry and improve the efficiency and accuracy of fault diagnosis work.

[0005] Based on this, how to effectively integrate voice input, standard fault library retrieval and text generation capabilities of large language models in a digital distribution network environment to automatically generate structured and professional equipment fault reports based on the on-site needs of front-line operators has become a key issue that needs to be urgently addressed in the power industry. Summary of the invention

[0006] Purpose of the invention: In order to solve the problem that the existing technology cannot effectively integrate voice input, standard fault library retrieval and text generation capabilities of large language models in a digital distribution network environment, and cannot automatically generate structured and professional equipment fault reports for on-site needs of front-line operators, the present invention proposes a digital distribution network equipment fault diagnosis method and system based on artificial intelligence speech recognition. By exploring and applying artificial intelligence retrieval enhanced generation technology, the present invention can significantly improve the efficiency and accuracy of fault handling, and provide more powerful protection for the safe operation of the power system.

[0007] Technical solution: A digital distribution network equipment fault diagnosis method based on artificial intelligence speech recognition, comprising the following steps: Step 1: Obtain a sound signal describing equipment failure information; Step 2: Convert the sound signal describing the fault information into a logarithmic Mel spectrum matrix; Step 3: Input the logarithmic Mel spectrum matrix into the encoder of the large language model, and the encoder of the large language model outputs a time-series embedding sequence, which passes through the corresponding text decoder to obtain the speech recognition text; Step 4: extracting keywords related to equipment failure from the speech recognition text, and performing entity recognition on each keyword, wherein the entities include: equipment name entity, fault symptom entity, and environmental condition entity; Step 5: Use the conditional probability model to determine the correlation between the different entities identified in step 4 and the fault type, and take the fault type with the highest fault probability as the diagnosis result.

[0008] Furthermore, in step 2, the sound signal describing the fault information is converted into a logarithmic Mel spectrum matrix, and the specific operations include: The sound signal x [n] After fast Fourier transform processing, the spectrum information of the sound signal is obtained f ; According to the spectrum information of the sound signal f , determine the spectrum information [ f min , f max ] range; among which, f min represents the lowest frequency,f max Denote the highest frequency; At f min , f max range design M a triangular Mel filter, and transform f min , f max to the Mel frequency axis, expressed as: ; ; Define that there are M + 2 frequency points on the Mel frequency axis, then the Mel frequency value corresponding to the i th frequency point is expressed as: ; Convert the Mel frequency value corresponding to each frequency point to the linear frequency according to the following formula, expressed as: ; Define the weight i of the f k th triangular Mel filter at the sound frequency sampling point in the time frame m, ; where m represents the time frame; Each triangular Mel filter covers a frequency range and performs weighted summation on the power spectrum within this frequency range to obtain the output energy of the i th triangular Mel filter: ; where, K represents the total number of frequency points obtained after performing the fast Fourier transform for each frame, k represents the frequency index, represents the power spectrum at the sound frequency sampling point f k in the time frame m; Take the logarithm of the output energy of each triangular Mel filter, expressed as: ; where, represents a positive number; Advancing in time, stack the of all time frames to form a logarithmic Mel spectrogram matrix.

[0009] Further, the sound frequency sampling points in time frame m f k The power spectrum at , is obtained according to the following steps: Multiply the sound signal x n with the moving time-domain window function shown in the following formula to intercept and obtain short-time signals of the sound signal x n at different time positions; ; In the formula, n represents the starting time index of the time-domain window function, N is the number of samples in the window length; Perform discrete Fourier transform on each short-time signal to obtain the STFT complex frequency spectrum , expressed as: ; Among them, H represents the hop size of the sliding in the time domain; Obtain the power spectrum according to the following formula: .

[0010] Further, in step 4, the keywords related to equipment failures are extracted from the speech recognition text, and entity recognition is performed on each keyword. The entities include: equipment name entity, fault symptom entity, and environmental condition entity; the specific operations include: Extract multiple keywords related to equipment failures from the speech recognition text; Match the extracted multiple keywords with the expert knowledge base in the power field to determine the category to which each keyword belongs; the expert knowledge base in the power field includes: equipment name library, fault type library, fault symptom library, and environmental condition library; the categories include equipment name, fault type, fault symptom, and environmental condition; According to the category to which it belongs, input each keyword into the corresponding named entity recognition model pre-trained based on existing guidelines to obtain the entity recognition result.

[0011] Further, in step 5, the conditional probability model is used to determine the association degree between different entities recognized in step 4 and the fault type, and the fault type with the highest fault probability is taken as the diagnosis result. The specific operations include: Use the following formula to convert the entity recognition result into a vector representation: ; Among them, represents the r ​​Entity recognition results, indicating the vectorized representation of the entity recognition results, f embed indicating the word embedding function, q indicating the total number of entity recognition results; The association degree between different entities and fault types is measured using the following conditional probability model: ; where FaultType represents the fault type, indicating the average probability of the occurrence of the fault type, is the cosine similarity, is the known entity in the fault type library; According to the calculated conditional probability, the fault type with the maximum conditional probability is taken as the diagnostic result.

[0012] Furthermore, after step 5, the following steps are also included: Step 6: Using the large language model and prompt words, generate an equipment fault diagnosis report that follows the pre-made common fault report template in the power industry based on the diagnostic result obtained in step 5.

[0013] The present invention also discloses a digital distribution network equipment fault diagnosis system based on artificial intelligence speech recognition, including: A sound signal acquisition module for acquiring sound signals describing equipment fault information; A signal conversion module for converting the sound signals describing fault information into a logarithmic Mel spectrogram matrix; A speech recognition text generation module for inputting the logarithmic Mel spectrogram matrix into the encoder of the large language model, and outputting a sequence of temporal embeddings by the encoder of the large language model, and obtaining the speech recognition text through the corresponding text decoder; An entity recognition module for extracting keywords related to equipment faults from the speech recognition text and performing entity recognition on each keyword, where the entities include: equipment name entities, fault symptom entities, and environmental condition entities; A fault type determination module for using a conditional probability model to determine the association degree between different entities recognized by the entity recognition module and fault types, and taking the fault type with the maximum fault probability as the diagnostic result.

[0014] Furthermore, in the signal conversion module, the specific operation of converting the sound signals describing fault information into a logarithmic Mel spectrogram matrix includes: The sound signal x [n] is processed by fast Fourier transform to obtain the spectral information of the sound signal f ; According to the spectral information of the sound signal f , determine the f min , f max range; where, f min represents the lowest frequency, f max represents the highest frequency; In the f min , f max range, design M triangular Mel filters, and transform f min , f max to the Mel frequency axis, expressed as: ; ; Define that there are M + 2 frequency points on the Mel frequency axis. Then, the Mel frequency value corresponding to the i th frequency point is expressed as: ; Convert the Mel frequency value corresponding to each frequency point to the linear frequency according to the following formula, expressed as: ; Define the weight i of the f k th sound frequency sampling point of the th triangular Mel filter in the time frame m, ; where, m represents the time frame; Each triangular Mel filter covers a frequency range and performs weighted summation on the power spectrum within this frequency range to obtain the output energy of the i th triangular Mel filter: ; where, K represents the total number of frequency points obtained after performing the fast Fourier transform for each frame, k represents the frequency index, represents the power spectrum at the sound frequency sampling point f k in the time frame m; Take the logarithm of the output energy of each triangular Mel filter, expressed as: ; Among them, represents a positive number; As time progresses, stack all time frames of to form a logarithmic Mel spectrogram matrix.

[0015] Furthermore, the sound frequency sampling points at time frame m f k The power spectrum at , is obtained according to the following steps: Multiply the sound signal x n with the moving time-domain window function shown in the following formula to intercept and obtain the short-time signals of the sound signal x n at different time positions; ; In the formula, n represents the starting time index of the time-domain window function, N is the number of samples in the window length; Perform a discrete Fourier transform on each short-time signal to obtain the STFT complex spectrum , expressed as: ; Among them, H represents the hop size sliding in the time domain; Obtain the power spectrum according to the following formula: .

[0016] Furthermore, extracting keywords related to equipment failures from the speech recognition text and performing entity recognition on each keyword, the entities include: equipment name entity, fault symptom entity, and environmental condition entity; the specific operations include: Extract multiple keywords related to equipment failures from the speech recognition text; Match the extracted multiple keywords with the expert knowledge base in the power field to determine the category to which each keyword belongs; the expert knowledge base in the power field includes: equipment name library, fault type library, fault symptom library, and environmental condition library; the categories include equipment name, fault type, fault symptom, and environmental condition; According to the category to which it belongs, input each keyword into the corresponding named entity recognition model pre-trained based on existing guidelines to obtain the entity recognition result.

[0017] ​​Further, the association degree between different entities identified by the entity recognition module is determined using a conditional probability model, and the fault type with the highest fault probability is taken as the diagnosis result. The specific operations include: Using the following formula, convert the entity recognition result into a vector representation: ; where, represents the r th entity recognition result, represents the vectorized representation of this entity recognition result, f embed represents the word embedding function, q represents the total number of entity recognition results; Use the following conditional probability model to measure the association degree between different entities and fault types: ; where FaultType represents the fault type, represents the average probability of the occurrence of the fault type, is the cosine similarity, is the known entity in the fault type library; According to the calculated conditional probability, take the fault type with the highest conditional probability as the diagnosis result.

[0018] Further, it also includes: An equipment fault diagnosis report generation module, which uses a large language model and prompt words to generate an equipment fault diagnosis report that follows a pre-made common fault report template in the power industry based on the diagnosis result obtained by the fault type determination module.

[0019] Beneficial effects: In the scenario of power grid equipment fault diagnosis, where grass-roots staff have inconvenient hands for information input and lack expert knowledge to quickly diagnose equipment faults, the present invention uses a large language model to perform text recognition on the input voice, inputs the text into the equipment fault standard library for retrieval and returns the detailed information of the equipment fault, and inputs it into the large language model to generate a formatted equipment fault analysis report. Compared with the prior art, the method of the present invention utilizes the powerful multi-modal retrieval, analysis, and text generation capabilities of the artificial intelligence retrieval-enhanced generation technology to generate a standardized equipment fault report according to the voice, effectively improving the work efficiency of grass-roots personnel in the power industry. Brief Description of the Drawings

[0020] Figure 1 is a flowchart of a digital distribution network equipment fault diagnosis method based on artificial intelligence speech recognition proposed by the present invention. Detailed Embodiment

[0021] The technical solution of the present invention will be further elaborated below in conjunction with the accompanying drawings and embodiments.

[0022] This embodiment proposes a digital distribution network equipment fault diagnosis method based on artificial intelligence speech recognition, which mainly includes the following steps: Step 1: Use a microphone or microphone hardware device to sample the speech describing the fault information f s and record it as an audio file. In this step, there is no restriction on the content of the externally input speech, and common statements and equipment defect sayings in the power industry are allowed to be input; Step 2: Convert the sound signal describing the fault information into a logarithmic Mel spectrogram matrix. The specific operations include: Convert the sound signal in the audio file x [n] through fast Fourier transform (FFT) processing to obtain the spectral information of the sound signal f ; According to the spectral information of the sound signal f , determine the f min , f max range; where f min represents the lowest frequency, which can be set to 0 or a certain low frequency, f max represents the highest frequency, f max and is generally set to half of the sampling rate.

[0023] In this embodiment, by converting the spectral information of the sound signal f into Mel spectral information and filtering, the non-linear perception of frequency by the human ear is simulated.

[0024] Therefore, in f min , f max range, design M triangular Mel filters, and convert f min , f max to the Mel frequency axis, which is expressed as: ; ; Define that there are M + 2 frequency points on the Mel frequency axis. Usually, M is 80, then the Mel frequency value corresponding to the i th frequency point is expressed as: ; Convert the Mel frequency value corresponding to each frequency point to the linear frequency according to the following formula, expressed as: ; In this step, the core is that the Mel frequency domain is more in line with the human ear's perception, that is, it is sensitive to low frequencies and insensitive to high frequencies. After processing in the Mel frequency domain, it needs to be restored to the natural frequency domain processed by FFT.

[0025] Define the sound frequency sampling points f k as: ; In the formula, k represents the frequency index after the fast Fourier transform, N represents the number of samples in the window length; Define the weight i of the i th triangular Mel filter ( M = 1, …, f k ) at the sound frequency sampling point as follows: ; where m represents the time frame, that is, the mth time frame of the entire audio segment.

[0026] The physical meaning of the parameter symbol i of the i th triangular Mel filter and the i th frequency point is the same.

[0027] Each triangular Mel filter covers a frequency range and weights and sums the power spectrum within this frequency range to obtain the Mel energy vector, that is, the output energy of the triangular Mel filter: ; where, K represents the total number of frequency points obtained after performing the fast Fourier transform on each frame, usually taking K = N / 2 + 1, k represents the frequency index, represents the power spectrum at the sound frequency sampling point f k at the time frame m, which can be calculated according to the following content: Multiply the sound signal x n by the moving time-domain window function shown in the following formula and intercept to obtain the sound signal x n ​​Short-time signals at different time positions; ; wherein, n represents the starting time index of the time-domain window function, which is an integer and slides with the analysis position, N is the number of samples in the window length, usually the number of samples corresponding to 20 - 40 milliseconds.

[0028] Perform discrete Fourier transform on each short-time signal to obtain the STFT complex spectrum , expressed as: ; wherein, H represents the hop size for sliding in the time domain, usually set H < N to produce overlap between frames.

[0029] The power spectrum is the energy distribution of the signal at different frequencies, so it is expressed as: .

[0030] Take the logarithm of the Mel energy vector according to the following formula to simulate the logarithmic perception of loudness by the human ear and also to enable the subsequent neural network to better process the dynamic range: ; wherein, is to prevent tending to negative infinity when the energy is very small, generally taking a small positive number, such as 10 −6 even smaller.

[0031] Stack all of all time frames according to the time progression to form a logarithmic Mel spectrum matrix.

[0032] In this step, the initial sound signal must be a time-domain signal, that is, the abscissa is the recording time. After the STFT transformation, if 100 frames are obtained, then m = 0~99. Each frame is then processed by Mel filtering to obtain 40 Mel filter outputs, so i = 0~39. The entire matrix has a size of 100 * 40. The rows are time frames, and each column vector is the Mel energy channel.

[0033] Step 3: Input the logarithmic Mel spectrum matrix into the encoder of the large language model. Use the encoder of the large language model to extract the high-level features of the audio, output a sequence of temporal embeddings, and then pass through the text decoder to form the complete speech recognition text. The encoder of the large language model used in this step is an existing encoder, and the corresponding text decoder is also an existing decoder.

[0034] Step 4: Extract keywords related to equipment failures from the speech recognition text, and perform entity recognition on each keyword. In this embodiment, the entity includes, but is not limited to: equipment name entity, fault symptom entity, and environmental condition entity. The specific operations include: Use the general byte-level byte pair encoding tokenization technique (Byte-level Byte Pair Encoding, BBPE) to tokenize the speech recognition text; According to the common stop word list, remove the stop words from the tokenized words to complete the filtering; the common stop word list refers to high-frequency and low-value words.

[0035] Use a natural language processing model to perform part-of-speech tagging and semantic extraction on the filtered words to determine multiple keywords related to equipment failures; the purpose of semantic extraction is, for example, for a power-specific word like neutral point, it is easily recognized as center point, so semantic extraction is needed.

[0036] Match multiple keywords with the power domain expert knowledge base respectively to determine the category to which each keyword belongs; in this embodiment, the power domain expert knowledge base at least includes: equipment name library, fault type library, fault symptom library, and environmental condition library; the categories to which each keyword belongs at least include: equipment name, fault type, fault symptom, and environmental condition. The purpose of this step is to classify the keywords, that is, to determine which library the keyword belongs to. For example, a certain keyword belongs to the fault type library.

[0037] According to the category to which it belongs, input each keyword into the corresponding named entity recognition (NER) model pre-trained based on existing guidelines and other documents to obtain the entity recognition result. Each library in the power domain expert knowledge base corresponds to a named entity recognition (NER) model. The named entity recognition (NER) model is pre-trained through the existing knowledge base to achieve entity recognition based on retrieval-augmented generation (RAG). Correspondingly, the entity recognition result at least includes: fault type entity, equipment name entity, fault symptom entity, and environmental condition entity. For example, if the keyword belongs to the fault symptom library and its category is fault symptom, then according to the category to which it belongs, input the keyword into the corresponding named entity recognition (NER) model to obtain the fault symptom entity.

[0038] Step 5: Use a conditional probability model to determine the association degree between different entities recognized in Step 4 and the fault type, and take the fault type with the highest fault probability as the diagnosis result. The specific operations include: Use the vectorization method shown in the following formula to convert the entity recognition result into a vector representation: ; where, Indicates the r th entity recognition result, represents the vectorized representation of this entity recognition result, f embed represents the word embedding function, q represents the total number of entity recognition results.

[0039] The following conditional probability model is used to measure the correlation between different entities and fault types: ; where FaultType represents the fault type, represents the average probability of the occurrence of the fault type, that is, the average probability of the occurrence of this fault type statistically obtained from existing guidelines, standards and other documents by humans in advance. For example, if the acetylene content in the oil chromatography exceeds 20%, then for the fault type of transformer overheating, its average probability is 40%, P base (Transformer overheating) = 40%. is the cosine similarity:

[0040] In the formula, is the known entity in the fault type library.

[0041] According to the calculated conditional probability, the fault type with the maximum conditional probability is taken as the diagnosis result.

[0042] In this step, the similarity reflects the similarity between the text in the speech and the text of the power professional vocabulary. For example, if the calculated similarity between "ferromagnetic resonance" and "resonant overvoltage" in the fault type library reaches >0.85, then these two can be classified into the same category with a probability of 85%.

[0043] For example: Suppose the keywords are phase B of the main transformer and the oil temperature is 78°C; match phase B of the main transformer and the oil temperature of 78°C with the expert knowledge base in the power field to determine that phase B of the main transformer belongs to the similarity of the equipment library, and the oil temperature of 78°C belongs to the fault symptom library.

[0044] Perform entity recognition on these two keywords respectively, and obtain the following entity recognition results: Equipment name entity: Phase B of the main transformer (similarity with the equipment library 0.93), category: transformer Fault symptom entity: Oil temperature 78°C (similarity with "overheating" in the fault symptoms 0.89), acetylene overlimit (similarity with "discharge" in the fault symptoms 0.91) Environmental entity: None.

[0045] According to the conditional probability model, the calculation gives: P(Arc discharge | Acetylene overlimit, oil temperature 78 overlimit model degree and similarity with "discharge" such as fault phenomena P base (0.25) = 0.72, so the most likely fault type is arc discharge.

[0046] Step 6: Use the large language model and prompts to perform context semantic understanding and segmentation on the speech recognition text, and combine the diagnostic results obtained in Step 5 to extract the key information and critical information required in the common fault report templates in the power industry; align and fill the critical information according to the placeholders, and use the large language model to rewrite, simplify or expand the filled text to meet the requirements of the power industry for report readability and professionalism.

[0047] All common fault report templates in the power industry used in this step need to define required fields and optional fields, including but not limited to equipment name, fault type, fault time, fault cause, and fault level. And the common fault report templates in the power industry contain expandable tags. Based on XML, JSON, HTML, or custom placeholders, placeholders are set for each keyword field in the common fault report templates in the power industry for the large language model to perform dynamic replacement. In this embodiment, different template versions or updates are recorded and managed to quickly switch or trace back template information in the fault diagnosis process.

[0048] Step 7: Export the text report output by the large language model into the required file formats, including but not limited to PDF, Word, HTML, Markdown, or JSON, etc. Push the generated report file to the operation and maintenance management system, cloud storage platform, or user terminal through HTTP interfaces, message queues, or other network communication methods, and record the timestamp and version information of the successfully output report to achieve traceable management of the fault diagnosis report.

[0049] Taking the inspection work of a certain provincial power grid of the State Grid as an example, when the inspection personnel of the distribution network substation are working, they find a suspected abnormal sound of the equipment. The sound is recorded on-site through a microphone device with a sampling rate set at 16 kHz and saved as an audio file. This audio file is first converted into a spectrum through the short-time Fourier transform, and then converted into a Mel spectrum to form a logarithmic Mel spectrum matrix. Subsequently, this logarithmic Mel spectrum matrix is input into the large language model encoder to extract high-level features and output a time series embedding sequence. The text decoder identifies the text content as "intermittent current noise in the main transformer in the southeast corner". Then, through byte pair encoding (BBPE) tokenization, stop word removal, part-of-speech tagging, and semantic extraction, keywords such as "main transformer", "current noise", and "southeast corner" are obtained. The named entity recognition model and the power industry knowledge base are used to match and identify the keywords. Through word embedding and cosine similarity calculation, the main transformer is identified as an equipment name entity, the current noise is identified as a fault symptom entity, and the southeast corner is identified as an environmental entity. Next, based on the power equipment fault conditional probability model, combined with each entity, the most likely fault type is determined as "poor winding contact", and the corresponding probability is output.

[0050] Finally, the "110kV Main Transformer Class" fault report template is selected, and fields such as equipment name, fault type, and time are identified and filled in to generate structured report content. After language optimization through the large language model, it is exported as a PDF format and pushed to the new generation of equipment asset lean management system (PMS 3.0). At the same time, the generation time and report version are recorded to achieve full-process fault diagnosis and traceable management. The model recognition accuracy of this method can be improved to about 90%.

Claims

1. A digital distribution network equipment fault diagnosis method based on artificial intelligence speech recognition, characterized in that: It includes the following steps: Step 1: Obtain a sound signal describing the device failure information; Step 2: Convert the sound signal describing the failure information into a logarithmic Mel spectrogram matrix; Step 3: Input the logarithmic Mel spectrogram matrix into the encoder of the large language model, and the encoder of the large language model outputs a sequence of temporal embeddings. This sequence of temporal embeddings passes through the corresponding text decoder to obtain the speech recognition text; Step 4: Extract keywords related to the device failure from the speech recognition text, and perform entity recognition on each keyword. The entities include: device name entity, fault symptom entity, and environmental condition entity; Step 5: Use a conditional probability model to determine the degree of association between different entities identified in Step 4 and the fault type, and take the fault type with the highest fault probability as the diagnosis result.

2. The digital distribution network equipment fault diagnosis method based on artificial intelligence speech recognition according to claim 1, characterized in that: In Step 2, the conversion of the sound signal describing the failure information into a logarithmic Mel spectrogram matrix specifically includes: The sound signal x [n] is processed by fast Fourier transform to obtain the spectral information of the sound signal f ; According to the spectral information of the sound signal f , determine the f min , f max range; wherein, f min represents the lowest frequency, f max represents the highest frequency; At f min , f max Range design M triangular Mel filters, and f min , f max go to the Mel frequency axis, expressed as: ; ; Define that there are M + 2 frequency points on the Mel frequency axis, then the Mel frequency value corresponding to the i th frequency point is expressed as: ; Convert the Mel frequency value corresponding to each frequency point to a linear frequency according to the following formula, expressed as: ; Define the i weight of the nth triangular Mel filter at the sound frequency sampling point in time frame m f k as follows: , denoted as: ; where m represents the time frame; Each triangular Mel filter covers a frequency range and performs a weighted summation of the power spectrum within that frequency range to obtain the output energy of the i th triangular Mel filter: ; Among them, K represents the total number of frequency points obtained after performing a fast Fourier transform on each frame, k represents the frequency index, represents the sound frequency sampling points at time frame m f k and the power spectrum at this position; Take the logarithm of the output energy of each triangular Mel filter, expressed as: ; Among them, represents a positive number; Stack all time frames as time progresses to form a log Mel spectrogram matrix.

3. The digital distribution network equipment fault diagnosis method based on artificial intelligence speech recognition according to claim 2, characterized in that: The sound frequency sampling points in time frame m f k The power spectrum at the position is obtained according to the following steps: The sound signal x n is multiplied by a moving time-domain window function shown by the following formula to intercept the short-time signals of the sound signal x n at different time positions;​​ ; wherein, n represents the starting time index of the time domain window function, N is the number of samples in the window length; Perform a discrete Fourier transform on each short-time signal to obtain the STFT complex spectrum , which is expressed as: ; Among them, H represents the step size of the hop that slides in the time domain; Obtain the power spectrum according to the following formula: 。 4. A digital distribution network equipment fault diagnosis method based on artificial intelligence speech recognition according to claim 1, characterized in that: In Step 4, the extraction of keywords related to the device failure from the speech recognition text and the entity recognition of each keyword. The entities include: device name entity, fault symptom entity, and environmental condition entity; specifically include: Extract multiple keywords related to the device failure from the speech recognition text; Match the extracted multiple keywords with the expert knowledge base in the power field to determine the category to which each keyword belongs. The expert knowledge base in the power field includes: device name library, fault type library, fault symptom library, and environmental condition library; the categories include device name, fault type, fault symptom, and environmental condition; According to the category to which it belongs, input each keyword into the corresponding named entity recognition model pre-trained based on existing guidelines to obtain the entity recognition result.

5. A digital distribution network equipment fault diagnosis method based on artificial intelligence speech recognition according to claim 4, characterized in that: In Step 5, the use of a conditional probability model to determine the degree of association between different entities identified in Step 4 and the fault type, and taking the fault type with the highest fault probability as the diagnosis result, specifically includes: Use the following formula to convert the entity recognition result into a vector representation: ; Among them, represents the r th entity recognition result, represents the vectorized representation of the entity recognition result, f embed represents the word embedding function, q represents the total number of entity recognition results; Use the following conditional probability model to measure the degree of association between different entities and the fault type: ; Among them, FaultType represents the fault type, represents the average probability of the occurrence of the fault type, is the cosine similarity, is the known entity in the fault type library; According to the calculated conditional probability, take the fault type with the highest conditional probability as the diagnosis result.

6. A digital distribution network equipment fault diagnosis method based on artificial intelligence speech recognition according to claim 1, characterized in that: After Step 5, it also includes the following steps: Step 6: Use the large language model and prompt words to generate a device fault diagnosis report that follows the commonly used fault report template in the power industry based on the diagnosis result obtained in Step 5.

7. A digital distribution network equipment fault diagnosis system based on artificial intelligence speech recognition, characterized in that: It includes: A sound signal acquisition module for obtaining a sound signal describing the device failure information; A signal conversion module for converting the sound signal describing the failure information into a logarithmic Mel spectrogram matrix; A speech recognition text generation module, which is used to input a logarithmic Mel spectrogram matrix into the encoder of a large language model, and the encoder of the large language model outputs a sequence of temporal embeddings. This sequence of temporal embeddings passes through the corresponding text decoder to obtain the speech recognition text; An entity recognition module, which is used to extract keywords related to equipment failures from the speech recognition text and perform entity recognition on each keyword. The entities include: equipment name entities, fault symptom entities, and environmental condition entities; A fault type determination module, which is used to use a conditional probability model to determine the correlation degree between different entities recognized by the entity recognition module and fault types, and take the fault type with the highest fault probability as the diagnostic result.

8. An artificial intelligence voice recognition-based digital distribution network equipment fault diagnosis system according to claim 7, characterized in that: In the signal conversion module, the conversion of the sound signal describing the fault information into a logarithmic Mel spectrogram matrix is as follows: The sound signal x [n] is processed by fast Fourier transform to obtain the spectral information of the sound signal f ; According to the spectral information of the sound signal f , determine the f min , f max range; wherein, f min represents the lowest frequency, f max represents the highest frequency; At f min , f max range design M triangular Mel filters, and f min , f max to the Mel frequency axis, expressed as: ; ; Define that there are M + 2 frequency points on the Mel frequency axis. Then, the Mel frequency value corresponding to the i th frequency point is expressed as: ; Convert the Mel frequency value corresponding to each frequency point to a linear frequency according to the following formula, expressed as: ; Define the i weight of the m-th triangular Mel filter at the sound frequency sampling point in time frame m f k as follows: , denoted as: ; where m represents the time frame; Each triangular Mel filter covers a frequency range and performs a weighted summation of the power spectrum within that frequency range to obtain the output energy of the i th triangular Mel filter: ; Among them, K represents the total number of frequency points obtained after performing a fast Fourier transform on each frame, k represents the frequency index, represents the sound frequency sampling point at time frame m f k is the power spectrum at this point; Take the logarithm of the output energy of each triangular Mel filter, expressed as: ; Among them, represents a positive number; Stack all time frames according to the time progression to form a log Mel spectrogram matrix.

9. An artificial intelligence voice recognition-based digital distribution network equipment fault diagnosis system according to claim 8, characterized in that: The sound frequency sampling points in time frame m f k The power spectrum at , is obtained according to the following steps: The voice signal x n is multiplied by a moving time-domain window function shown by the following formula and the voice signal x n at different time positions is intercepted to obtain short-time signals;​​ ; wherein, n represents the start time index of the time domain window function, N is the number of samples in the window length; Perform a discrete Fourier transform on each short-time signal to obtain the STFT complex spectrum , which is expressed as: ; Among them, H represents the step size that slides in the time domain; Obtain the power spectrum according to the following formula: 。 10. A digital distribution network equipment fault diagnosis system based on artificial intelligence speech recognition according to claim 9, characterized in that: The extraction of keywords related to equipment failures from the speech recognition text and the entity recognition of each keyword. The entities include: equipment name entities, fault symptom entities, and environmental condition entities; the specific operations include: Extract multiple keywords related to equipment failures from the speech recognition text; Match the extracted multiple keywords with the expert knowledge base in the power field to determine the category to which each keyword belongs; the expert knowledge base in the power field includes: equipment name library, fault type library, fault symptom library, and environmental condition library; the categories include equipment name, fault type, fault symptom, and environmental condition; According to the category to which it belongs, input each keyword into the corresponding named entity recognition model pre-trained based on existing guidelines to obtain the entity recognition result.

11. An artificial intelligence voice recognition-based digital distribution network equipment fault diagnosis system according to claim 10, characterized in that: The use of a conditional probability model to determine the correlation degree between different entities recognized by the entity recognition module and fault types, and take the fault type with the highest fault probability as the diagnostic result. The specific operations include: Use the following formula to convert the entity recognition result into a vector representation: ; Among them, represents the r th entity recognition result, represents the vectorized representation of the entity recognition result, f embed represents the word embedding function, q represents the total number of entity recognition results; Use the following conditional probability model to measure the correlation degree between different entities and fault types: ; Among them, FaultType represents the fault type, represents the average probability of the occurrence of the fault type, is the cosine similarity, is the known entity in the fault type library; According to the calculated conditional probability, take the fault type with the highest conditional probability as the diagnostic result.

12. A digital distribution network equipment fault diagnosis system based on artificial intelligence speech recognition according to claim 11, characterized in that: It also includes: An equipment fault diagnosis report generation module, which is used to use a large language model and prompting words to generate an equipment fault diagnosis report that follows the commonly used fault report template in the power industry based on the diagnostic result obtained by the fault type determination module.

Citation Information

Patent Citations

  • Built-in speech discriminating method based on sub-word hidden Markov model

    CN101030369A

  • Power grid fault handling plan analysis method based on neural regular expression

    CN114997168A

  • Intelligent interaction method and system for equipment fault knowledge

    CN117933249A

  • Fan sound fault feature detection method and system

    CN118335109A

  • Power grid fault information extraction and processing method and system based on voice recognition

    CN118609552A