An AI-based voice interaction system for power grid operation and maintenance

By adopting artificial intelligence interactive analysis module and model optimization module in the grid operation and maintenance site, combining voice and gesture recognition technology, the speech recognition accuracy and intelligence problems in the grid operation and maintenance site are solved, and more efficient voice and gesture interconnection recognition is achieved, and the applicability and intelligence of the system are improved.

CN118098233BActive Publication Date: 2025-07-11FOSHAN POWER SUPPLY BUREAU GUANGDONG POWER GRID
View PDF 2 Cites 0 Cited by

Patent Information

Application Number
CN202410334804.9
Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Filing Date
2024-03-22
Publication Date
2025-07-11
Estimated Expiration
2044-03-22

AI Technical Summary

Technical Problem

The existing grid operation and maintenance voice interaction system has low speech recognition accuracy at the grid operation and maintenance site, and has poor intelligence and interconnectivity, so it is unable to effectively deal with the influence of environmental factors such as mechanical operation sound and current sound.

Method used

The interactive analysis module and model optimization module based on artificial intelligence are adopted to receive voice commands through a microphone array, combine acoustic processing algorithms and acoustic models of multiple operation and maintenance personnel for voice recognition, and use gesture trigger units to perform supplementary recognition when the recognition fails, and combine the camera device to capture gesture commands to achieve comprehensive recognition of voice and gestures.

Benefits of technology

It improves the accuracy of receiving voice commands and the intelligence of gesture recognition, avoids the incorrect execution of voice commands, and enhances the applicability and interconnectivity of the system in the power grid operation and maintenance site.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN118098233B_ABST
    Figure CN118098233B_ABST
Patent Text Reader

Abstract

The present invention discloses a voice interaction system for power grid operation and maintenance based on artificial intelligence. The present invention receives a voice command issued by an operation and maintenance personnel and conducts preliminary comparison and matching. After successful matching, it further analyzes the reception error of the voice command, uses voice recognition technology to convert the currently received voice command into a text format, and matches the converted text format with the corresponding texts of multiple relevant voice commands to obtain the corresponding text with the highest matching degree for the currently converted text format. Based on the analysis results between the two texts and comprehensive consideration, the reception difference of the current voice command is obtained, thereby reflecting the difference degree between the currently received voice command and the matching corresponding text. The calculated reception difference is compared with the corresponding reception threshold, and based on the comparison result, the next response is executed, avoiding the wrong execution of the voice command and improving the accuracy of voice command reception and determination.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention relates to the technical field of power grid operation and maintenance voice interaction, and particularly to an artificial intelligence-based power grid operation and maintenance voice interaction system. Background Art

[0002] With the continuous development of the power industry, the scale of the power grid has been expanding, and the operation and maintenance management work has become increasingly complex. There are many deficiencies in the traditional operation and maintenance management methods, such as low efficiency, many misoperations, and untimely information transmission. With the continuous development of artificial intelligence technology, human-computer interaction has become an indispensable part of daily life. The traditional interaction methods mainly rely on input devices such as keyboards and mice, which cannot meet people's needs for a more natural and convenient interaction method.

[0003] However, the existing power grid operation and maintenance voice interaction systems still have the following defects:

[0004] In the actual process of power grid operation and maintenance, due to the particularity of the power grid operation and maintenance site itself, such as mechanical operation sounds, current sounds, and other environmental factors that can affect voice reception, the accuracy of voice interaction recognition used by operation and maintenance personnel in actual applications is relatively low, and the currently received voice command cannot be converted into text, and the accuracy of the voice command cannot be determined based on the analysis result of the characters in the text;

[0005] At the same time, the intelligence level and connectivity of the voice interaction system are relatively poor, and it cannot use the camera device on the operation and maintenance site and the preset gesture commands to further realize the interaction function in the case of voice interaction failure.

[0006] Therefore, an artificial intelligence-based power grid operation and maintenance voice interaction system is introduced. Summary of the Invention

[0007] In view of this, the present invention provides an artificial intelligence-based power grid operation and maintenance voice interaction system to solve the problems raised in the above background art.

[0008] The object of the present invention can be achieved by the following technical solutions: including an interaction analysis module and a model optimization module;

[0009] A voice processing unit and a gesture trigger unit are arranged in the interaction analysis module;

[0010] The voice processing unit receives the voice command issued by the operation and maintenance personnel through the microphone array arranged in the operation and maintenance environment and conducts preliminary comparison and matching. After successful matching, it further analyzes the reception error of the voice command and issues the next system response based on the analysis result. The specific steps are as follows:

[0011] Step 1: The operation and maintenance personnel wake up the interaction system through a wake-up voice command and issue relevant voice commands; the relevant voice commands include but are not limited to "query the current power grid status", "handle power grid alarms", "turn on or off the device No. 3", "query the historical data of the power grid", "trigger the substation maintenance task".

[0012] Step 2: The system records the received voice command, extracts the decibel level of the audio corresponding to the currently received voice command, calculates the decibel level of the current audio using an acoustic processing algorithm, compares the calculated decibel level with a preset decibel threshold. If the decibel level is greater than the preset decibel threshold, it is determined that the current voice command receiving environment is too noisy and further processing of the sound is required. When the sound decibel is less than the threshold, no further processing is performed, and the reception error of the current voice command content is directly analyzed.

[0013] Step 3: When the decibel level is greater than the preset decibel threshold, through steps of anti-aliasing filtering, analog-to-digital conversion, framing, and pre-emphasis processing, three characteristic parameters of the Mel-frequency cepstral coefficients, formants, and zero-crossing rate corresponding to the current sound data are extracted, and the extracted characteristic parameters are combined to obtain a complete acoustic feature representation.

[0014] Step 4: Using the acoustic models of multiple operation and maintenance personnel that have been trained and set up in advance, the extracted acoustic features are matched with each model to obtain the matching probability of each model; the maximum probability value among the current matching probabilities of each model is extracted, and this probability value is used as comparison data and compared with a preset matching threshold. If it is greater than the preset matching threshold, it is determined that the matching is successful, and the reception error of the current voice command content is further analyzed; if it is less than the preset matching threshold, it is determined that the matching fails, and the system issues a response "Voice command recognition failed, please say it again", and at the same time wakes up the gesture recognition function set in the gesture trigger unit.

[0015] The analysis of the reception error of the current voice command content is as follows:

[0016] Using speech recognition technology, the currently received voice command is converted into a text format, and the converted text format is matched with the texts corresponding to multiple relevant voice commands to obtain the text with the highest matching degree corresponding to the current converted text format.

[0017] Count the number of different characters in the current converted text format and the matched corresponding text and mark it as A1; count the total number of characters in the current converted text and the corresponding text respectively, and mark them as A2 and A3. Based on the comparison result of the two, calculate the proportion JG of the number of unmatched characters; that is, through the formula Calculate and obtain;

[0018] Perform exact word segmentation based on the thesaurus, and segment the current converted text format and the matching corresponding text;

[0019] Perform word-level comparison on the segmented text in the following manner:

[0020] Count the number of words that appear in the converted text format but do not appear in the matching corresponding text to obtain the new value XZ;

[0021] Count the number of words that appear in the matching corresponding text but do not appear in the converted text format to obtain the deletion value XY;

[0022] Count the number of words with corresponding positions but inconsistent contents in the converted text format and the matching corresponding text to obtain the replacement value XM;

[0023] Substitute the parameters obtained above into the formula , and calculate to obtain the error condition value JM; where a1, a2, and a3 are the influence weight factors of the new value XZ, the deletion value XY, and the replacement value XM respectively;

[0024] Calculate the Levenshtein distance between the converted text and the matching corresponding text to obtain the difference value JW; the specific steps are as follows:

[0025] Initialization: Create a matrix of (K + 1) * (H + 1), where K represents the length of the matching corresponding text and H represents the length of the converted text format; the rows of the matrix represent the characters of the matching corresponding text, and the columns represent the characters of the converted text format;

[0026] Filling: Fill the first row with numbers from 0 to H; fill the first column with numbers from 0 to K;

[0027] Calculate the edit distance: Use the dynamic programming method to fill the created matrix; for each position (t, b) in the created matrix, which represents the edit distance between the first t characters of the matching corresponding text and the first b characters of the converted text, calculate based on the following different situations:

[0028] If the t-th character of the matching corresponding text is the same as the b-th character of the converted text, the value of this position is equal to the value of the previous position in the matrix;

[0029] If the t-th character of the matching corresponding text is different from the b-th character of the converted text, the value of this position is equal to the minimum of the value of the previous position in the matrix plus 1, the value of the adjacent position in the previous row plus 1, and the value of the adjacent position in the previous column plus 1;

[0030] Obtain the Levenshtein distance: The value in the lower right corner of the finally created matrix is the Levenshtein distance between the matching corresponding text and the transformed text, representing the difference value JW between them;

[0031] Substitute the occupancy ratio JG, error condition value JM, and difference value JW in the above parameters into the formula , and calculate to obtain the received difference JSC of the current voice command; where F1, F2, and F3 respectively represent the maximum occupancy ratio, maximum error condition value, and maximum difference value allowed for the matching corresponding text; va1, va2, and va3 are the influence weight factors of the occupancy ratio JG, error condition value JM, and difference value JW respectively;

[0032] Compare the calculated received difference JSC with the corresponding received threshold. If it is greater than the corresponding received threshold, it is determined that the reception accuracy of the current voice command is low, and a response "Voice command recognition failed, please say it again" is sent through the system, and at the same time, the gesture recognition function set in the gesture trigger unit is awakened;

[0033] If after the matching fails, the operation and maintenance personnel simultaneously make a gesture command and a new voice command, then analyze the two groups of commands respectively, and match the analysis results of the two groups of commands. If the matching fails, the analysis result of the gesture command is taken as the main one.

[0034] After the gesture trigger unit is awakened by the built-in gesture recognition function, extract the above-mentioned voice command-related data and process it, and capture the located operation and maintenance personnel through multiple groups of camera devices set in the operation and maintenance environment, and collect and recognize the gestures made by the operation and maintenance personnel in real time. After successful recognition, perform corresponding operations based on the operation and maintenance commands corresponding to the current gesture command; specifically:

[0035] Use the sound source localization technology to determine the approximate position where the current voice command is issued, and based on this positioning result, lock the operation and maintenance personnel closest to this positioning result, control the multiple groups of camera devices in the operation and maintenance environment to face the located operation and maintenance personnel, and recognize and analyze the gestures made by this operation and maintenance personnel, specifically:

[0036] Collect the image information of the gestures made by this operation and maintenance personnel after the gesture recognition function is awakened through multiple groups of camera devices, and select the image information collected by the camera device corresponding to the body orientation of this operation and maintenance personnel as the analysis image;

[0037] After preprocessing the image, segment it, extract the hand area of the operation and maintenance personnel in the image information, and rotate and scale the extracted hand area by angle and ratio to make it consistent with the preset multiple groups of gesture images;

[0038] Compare the adjusted hand region image with multiple sets of preset gesture images, calculate the overlapping area between the two, and obtain the area overlap values for each group; select the preset gesture image corresponding to the maximum area overlap value as the probability execution instruction; and calculate the ratio between the maximum area overlap value and the total area of the hand region image to obtain the overlap ratio PC;

[0039] Based on the obtained overlapping area, determine the overlapping region in the two sets of images, remove this region from the adjusted hand region image, calculate the area of the remaining region after removal, compare the remaining region area with a preset remaining threshold, and obtain different weight coefficients based on the comparison result (weight coefficient one is 1.08, weight coefficient two is 1.15). If it is less than the preset remaining threshold, automatically match weight coefficient one; otherwise, match weight coefficient two; calculate the ratio between the remaining region area and the total area of the hand region image, and multiply it with the corresponding weight coefficient to obtain the remaining ratio PR;

[0040] Mark the hand region image and the corresponding preset gesture image as R and T respectively, calculate the means of the two sets of images, that is, the average value of all pixel values in each image and mark them as R1 and T1; use the variance formula to calculate the variance values of the two sets of images respectively, and mark them as R2 and T2;

[0041] Calculate the cross - correlation coefficient NCC between the two sets of images, that is, \(NCC=\frac{\sum_{x,y}(R(x,y)-R1)*(T(x,y)-T1)}{\sqrt{\sum_{x,y}(R(x,y)-RI)R2}*\sqrt{\sum_{x,y}(T(x,y)-T1)T2}}\);

[0042] The value range of NCC is between - 1 and 1. 1 means complete match, - 1 means complete non - match, and 0 means no linear correlation; the value of NCC can be used to measure the similarity between two images, and the closer it is to 1, the more similar the two images are;

[0043] Substitute the overlap ratio PC, remaining ratio PR, and cross - correlation coefficient NCC in the above parameters into the formula , and calculate to obtain the hand alignment value SBZ; where nva, nvb, and nvc respectively represent the minimum allowable overlap ratio, maximum allowable remaining ratio, and passing cross - correlation coefficient corresponding to the probability execution instruction; ty1, ty2, and ty3 are the influence weight factors of the overlap ratio PC, remaining ratio PR, and cross - correlation coefficient NCC respectively;

[0044] Compare the calculated hand ratio standard value SBZ with the corresponding hand ratio threshold. If it is greater than the corresponding hand ratio threshold, it is determined that the recognition is successful, and the instruction corresponding to the gesture is directly executed; if it is less than the corresponding hand ratio threshold, it is determined that the recognition is unsuccessful, and at the same time, the probability execution instruction selected is cancelled. Recalculate the hand ratio standard value of each group of preset gesture images and the image of this hand area, and compare it with the corresponding hand ratio threshold. If there is no value greater than the hand ratio threshold, it is determined that the gesture recognition of the operation and maintenance personnel this time fails; if there is a value greater than the hand ratio threshold and the number is multiple groups, select the largest value as the recognition success instruction. At this time, the system responds to the specific content of this recognition success instruction and asks whether to execute, that is, "The recognition instruction is XXX, do you want to execute?".

[0045] Compared with the prior art, the beneficial effects of the present invention are as follows:

[0046] The present invention receives the voice instruction issued by the operation and maintenance personnel and conducts preliminary comparison and matching. After successful matching, it further analyzes the reception error of the voice instruction, uses voice recognition technology to convert the currently received voice instruction into text format, and matches the converted text format with the corresponding texts of multiple related voice instructions to obtain the corresponding text with the highest matching degree corresponding to the currently converted text format. Based on the analysis results between the two texts and comprehensive consideration, the reception difference value of the current voice instruction is obtained, so as to reflect the difference degree between the currently received voice instruction and the matching corresponding text. Compare the calculated reception difference value with the corresponding reception threshold, and based on the comparison result, perform the next step of response, avoiding the wrong execution of the voice instruction and improving the accuracy of voice instruction reception and determination;

[0047] Based on the comparison and matching results of the voice instruction, the gesture recognition function set in the gesture trigger unit is correspondingly awakened, and the above-mentioned voice instruction-related data is extracted and processed. The operation and maintenance personnel located are captured by multiple sets of camera devices set in the operation and maintenance environment, and the gestures made by the operation and maintenance personnel are collected and recognized in real time, improving the accuracy of gesture image collection; and the coincidence ratio, remaining ratio and cross-correlation coefficient in the current operation and maintenance personnel gesture image are comprehensively analyzed to obtain the hand ratio standard value. Based on the comparison result of the hand ratio standard value, the next operation is performed, so as to realize the interconnection between the voice interaction system and the camera devices in the power grid operation and maintenance site, and improve the intelligence level of the system. BRIEF DESCRIPTION OF THE DRAWINGS

[0048] In the following description of the exemplary embodiments in conjunction with the drawings, more details, features and advantages of the present application are disclosed. In the drawings:

[0049] Figure 1 is the principle block diagram of the present invention. DETAILED DESCRIPTION OF THE INVENTION

[0050] Several embodiments of the present application will be described in more detail below with reference to the accompanying drawings so that those skilled in the art can implement the present application. The present application can be embodied in many different forms and for many different purposes and should not be limited to the embodiments set forth herein. These embodiments are provided so that this application will be thorough and complete, and will fully convey the scope of the application to those skilled in the art. The embodiments do not limit the present application.

[0051] Unless otherwise defined, all terms (including technical and scientific terms) used herein have the same meaning as commonly understood by one of ordinary skill in the art to which this application belongs. It will be further understood that terms such as those defined in commonly used dictionaries should be interpreted as having a meaning that is consistent with their meaning in the relevant art and / or the context of this specification, and will not be interpreted in an idealized or overly formal sense unless expressly so defined herein.

[0052] Please refer to Figure 1 as shown, an artificial intelligence-based voice interaction system for power grid operation and maintenance, including an interaction analysis module and a model optimization module;

[0053] A voice processing unit and a gesture trigger unit are provided in the interaction analysis module;

[0054] The voice processing unit receives the voice commands issued by the operation and maintenance personnel through a microphone array deployed in the operation and maintenance environment and performs preliminary comparison and matching. After successful matching, it further analyzes the reception error of the voice commands and issues the next system response based on the analysis results. The specific steps are as follows:

[0055] Step 1: The operation and maintenance personnel wake up the interaction system through a wake-up voice command and issue relevant voice commands; the relevant voice commands include but are not limited to "query the current power grid status", "handle power grid alarms", "turn on or off device No. 3", "query power grid historical data", "trigger substation maintenance tasks";

[0056] Step 2: The system enters the received voice commands and extracts the decibel level of the audio corresponding to the currently received voice command. It calculates the decibel level of the current audio using an acoustic processing algorithm and compares the calculated decibel level with a preset decibel threshold. If the decibel level is greater than the preset decibel threshold, it is determined that the current voice command reception environment is too noisy and further processing of the sound is required. When the sound decibel is less than the threshold, no further processing is performed, and the reception error of the current voice command content is directly analyzed;

[0057] Step 3: When the decibel level is greater than the preset decibel threshold, through anti-aliasing filtering, analog-to-digital conversion, framing, and pre-emphasis processing steps, extract three feature parameters, namely Mel-frequency cepstral coefficients, formants, and zero-crossing rates, corresponding to the current sound data, and combine and combine the extracted feature parameters to obtain a complete acoustic feature representation;

[0058] Step 4: Use the pre-trained and built acoustic models of multiple operation and maintenance personnel to match the extracted acoustic features with each model to obtain the matching probability of each model; extract the maximum probability value among the current matching probabilities of each model, and use this probability value as comparison data to compare with the preset matching threshold. If it is greater than the preset matching threshold, it is determined that the matching is successful, and further analyze the reception error of the current voice command content; if it is less than the preset matching threshold, it is determined that the matching fails, and the system issues a response "Voice command recognition failed, please say it again", and at the same time wakes up the gesture recognition function set in the gesture trigger unit;

[0059] Analyze the reception error of the current voice command content, specifically:

[0060] Use speech recognition technology to convert the currently received voice command into text format, and match the converted text format with the corresponding texts of multiple related voice commands to obtain the corresponding text with the highest matching degree for the current converted text format;

[0061] Count the number of different characters between the current converted text format and the matching corresponding text and mark it as A1; count the total number of characters in the current converted text and the corresponding text respectively, and mark them as A2 and A3. Based on the comparison result of the two, calculate the proportion JG of the number of unmatched characters; that is, through the formula Calculate and obtain;

[0062] Based on the accurate word segmentation of the thesaurus, perform word segmentation processing on the current converted text format and the matching corresponding text;

[0063] Perform word-level comparison on the segmented text in the following way:

[0064] Count the number of words that appear in the converted text format but do not appear in the matching corresponding text to obtain the new value XZ;

[0065] Count the number of words that appear in the matching corresponding text but do not appear in the converted text format to obtain the deletion value XY;

[0066] Count the number of words with corresponding positions but inconsistent contents between the converted text format and the matching corresponding text to obtain the replacement value XM;

[0067] Substitute the above obtained parameters into the formula , perform calculations to obtain the error condition value JM; where a1, a2, and a3 are the influence weight factors of the new value XZ, the deleted value XY, and the replacement value XM respectively;

[0068] Calculate the Levenshtein distance between the transformed text and the matched corresponding text to obtain the difference value JW between the two; the specific steps are as follows:

[0069] Initialization: Create a (K + 1) * (H + 1) matrix, where K represents the length of the matched corresponding text and H represents the length of the transformed text format; the rows of the matrix represent the characters of the matched corresponding text, and the columns represent the characters of the transformed text format;

[0070] Filling: Fill the first row with numbers from 0 to H; fill the first column with numbers from 0 to K;

[0071] Calculate the edit distance: Use the method of dynamic programming to fill the created matrix; for each position (t, b) in the created matrix, which represents the edit distance between the first t characters of the matched corresponding text and the first b characters of the transformed text, calculate based on the following different situations:

[0072] If the t-th character of the matched corresponding text is the same as the b-th character of the transformed text, the value of this position is equal to the value of the previous position in the matrix;

[0073] If the t-th character of the matched corresponding text is different from the b-th character of the transformed text, the value of this position is equal to the minimum of the value of the previous position in the matrix plus 1, the value of the adjacent position in the previous row plus 1, and the value of the adjacent position in the previous column plus 1;

[0074] Obtain the Levenshtein distance: The value in the lower right corner of the finally created matrix is the Levenshtein distance between the matched corresponding text and the transformed text, representing the difference value JW between them;

[0075] Substitute the proportion value JG, the error condition value JM, and the difference value JW in the above parameters into the formula , perform calculations to obtain the received difference JSC of the current voice command; where F1, F2, and F3 respectively represent the maximum allowable proportion value, the maximum error condition value, and the maximum difference value of the matched corresponding text; va1, va2, and va3 are the influence weight factors of the proportion value JG, the error condition value JM, and the difference value JW respectively;

[0076] Compare the calculated received difference JSC with the corresponding received threshold. If it is greater than the corresponding received threshold, it is determined that the reception accuracy of the current voice command is low, and the system issues a response "Voice command recognition failed, please say it again", and at the same time wakes up the gesture recognition function set in the gesture trigger unit;

[0077] It should be noted that if, after the matching fails, the operation and maintenance personnel simultaneously make a gesture instruction and a new voice instruction, the two groups of instructions will be analyzed separately, and the analysis results of the two groups of instructions will be matched. If the matching fails, the analysis result of the gesture instruction will be taken as the main one.

[0078] After the gesture trigger unit is awakened through the built-in gesture recognition function, it extracts and processes the above-mentioned voice instruction-related data, captures the located operation and maintenance personnel through multiple sets of camera devices set in the operation and maintenance environment, and real-time collects and recognizes the gestures made by the operation and maintenance personnel. After successful recognition, corresponding operations are performed based on the operation and maintenance instructions corresponding to the current gesture instruction; specifically:

[0079] Use the sound source localization technology to determine the approximate position where the current voice instruction is issued. Based on this positioning result, lock the operation and maintenance personnel closest to this positioning result, control the multiple sets of camera devices in the operation and maintenance environment to face the located operation and maintenance personnel, and recognize and analyze the gestures made by this operation and maintenance personnel. Specifically:

[0080] Collect the image information of the gestures made by this operation and maintenance personnel after the gesture recognition function is awakened through multiple sets of camera devices, and select the image information collected by the camera device corresponding to the body orientation of this operation and maintenance personnel as the analysis image;

[0081] After preprocessing the image, segment it, extract the hand area of the operation and maintenance personnel in the image information, and rotate and scale the angle of the extracted hand area to make it consistent with multiple sets of preset gesture images;

[0082] Compare the adjusted hand area image with multiple sets of preset gesture images, calculate the overlapping area between the two, and obtain the area overlap values of each group; select the preset gesture image corresponding to the maximum area overlap value as the probability execution instruction; and calculate the ratio between the maximum area overlap value and the total area of the hand area image to obtain the overlap ratio PC;

[0083] According to the obtained overlapping area, determine the overlapping area in the two groups of images, remove this area from the adjusted hand area image, calculate the area of the remaining area after removal, compare the remaining area with the preset remaining threshold, and obtain different weight coefficients based on the comparison result (weight coefficient one is 1.08, weight coefficient two is 1.15). If it is less than the preset remaining threshold, automatically match weight coefficient one, otherwise match weight coefficient two; calculate the ratio between the remaining area and the total area of the hand area image, and multiply it by the corresponding weight coefficient to obtain the remaining ratio PR;

[0084] Label the hand region image and the corresponding preset gesture image as R and T respectively, calculate the means of the two sets of images, that is, the average of all pixel values in their respective images, and label them as R1 and T1; use the variance formula to calculate the variance values of the two sets of images respectively, and label them as R2 and T2;

[0085] Calculate the cross-correlation coefficient NCC between the two sets of images, that is, through NCC = \frac{\sum_{x,y} (R(x,y)-R1)*(T(x,y)-T1)}{\sqrt{\sum_{x,y}(R(x,y)-RI)R2}*\sqrt{\sum_{x,y}(T(x,y)-T1)T2}};

[0086] It should be noted that the value range of NCC is between -1 and 1. 1 means a perfect match, -1 means a complete mismatch, and 0 means no linear correlation; the value of NCC can be used to measure the similarity between two images, and the closer it is to 1, the more similar the two images are;

[0087] Substitute the coincidence ratio PC, the remaining ratio PR, and the cross-correlation coefficient NCC in the above parameters into the formula , and calculate to obtain the hand alignment value SBZ; where nva, nvb, and nvc respectively represent the lowest allowable coincidence ratio, the maximum allowable remaining ratio, and the passing cross-correlation coefficient corresponding to the probability execution instruction; ty1, ty2, and ty3 are respectively the influence weight factors of the coincidence ratio PC, the remaining ratio PR, and the cross-correlation coefficient NCC;

[0088] Compare the calculated hand alignment value SBZ with the corresponding hand alignment threshold. If it is greater than the corresponding hand alignment threshold, it is determined that the recognition is successful, and directly execute the instruction corresponding to the gesture; if it is less than the corresponding hand alignment threshold, it is determined that the recognition is unsuccessful, and at the same time cancel the selected probability execution instruction, recalculate the hand alignment value of each preset gesture image and the hand region image, and compare it with the corresponding hand alignment threshold. If there is no value greater than the hand alignment threshold, it is determined that the gesture recognition of the operation and maintenance personnel this time fails; if there is a value greater than the hand alignment threshold and the number is multiple groups, select the largest value as the recognized successful instruction. At this time, the system responds to the specific content of this recognized successful instruction and asks whether to execute, that is, "The recognition instruction is XXX, do you want to execute?";

[0089] It should be noted that if the operation and maintenance personnel simultaneously show a gesture instruction and a new voice instruction, and the gesture recognition of the operation and maintenance personnel this time fails, and the voice instruction matches successfully or is determined to be successful, then the analysis result of the new voice instruction shall prevail.

[0090] The model optimization module is used to collect the recognition processes and results of each time of the grid operation and maintenance personnel, and use the data obtained to train and optimize the established model.

[0091] The preferred embodiments of the present invention disclosed above are only used to help illustrate the present invention. The preferred embodiments do not describe all the details in detail, nor do they limit the invention to only the specific embodiments. Obviously, according to the content of this specification, many modifications and variations can be made. These embodiments are selected and specifically described in this specification in order to better explain the principles and practical applications of the present invention, so that those skilled in the art can well understand and utilize the present invention. The present invention is only limited by the claims and their full scope and equivalents.

Claims

1. An artificial intelligence-based voice interaction system for power grid operation and maintenance, characterized in that, Including: Interaction analysis module: including a voice processing unit and a gesture trigger unit; The voice processing unit receives the voice commands issued by the operation and maintenance personnel through the microphone array deployed in the operation and maintenance environment, conducts preliminary comparison and matching, and further analyzes the reception error of the voice commands after successful matching. Based on the analysis results, it issues the next system response. The specific steps are as follows: Step 1: The operation and maintenance personnel wake up the interaction system through the wake-up voice command and issue relevant voice commands; Step 2: The system records the received voice commands, extracts the decibel level of the audio corresponding to the currently received voice command, calculates the decibel level of the current audio using the acoustic processing algorithm, and compares the calculated decibel level with the preset decibel threshold. If the decibel level is greater than the preset decibel threshold, it is determined that the current voice command reception environment is too noisy and further processing of the sound is required. When the sound decibel is less than the threshold, no further processing is performed, and the reception error of the current voice command content is directly analyzed; Step 3: When the decibel level is greater than the preset decibel threshold, through anti-aliasing filtering, analog-to-digital conversion, frame segmentation, and pre-emphasis processing steps, three feature parameters, namely Mel-frequency cepstral coefficients, formants, and zero-crossing rates, corresponding to the current sound data are extracted, and the extracted feature parameters are combined to obtain a complete acoustic feature representation; Step 4: Using the acoustic models of multiple operation and maintenance personnel that have been trained and built in advance, the extracted acoustic features are matched with each model to obtain the matching probability of each model; the maximum probability value among the current matching probabilities of each model is extracted, and this probability value is used as comparison data and compared with the preset matching threshold. If it is greater than the preset matching threshold, it is determined that the matching is successful, and the reception error of the current voice command content is further analyzed to obtain the reception difference. The calculated reception difference is compared with the corresponding reception threshold. If it is greater than the corresponding reception threshold, it is determined that the reception accuracy of the current voice command is low, and the system issues a response "Voice command recognition failed, please say it again", and at the same time wakes up the gesture recognition function set in the gesture trigger unit; If it is less than the preset matching threshold, it is determined that the matching fails, and the system issues a response "Voice command recognition failed, please say it again", and at the same time wakes up the gesture recognition function set in the gesture trigger unit; After the gesture trigger unit is woken up by the built-in gesture recognition function, it extracts and processes the above-mentioned voice command-related data, captures the located operation and maintenance personnel through multiple groups of camera devices set in the operation and maintenance environment, and real-time collects and recognizes the gestures made by the operation and maintenance personnel. After successful recognition, corresponding operations are performed based on the operation and maintenance commands corresponding to the current gesture commands.

2. The voice interaction system for power grid operation and maintenance based on artificial intelligence according to claim 1, characterized in that, The specific steps to obtain the occupancy ratio JG are as follows: Using voice recognition technology, the currently received voice command is converted into text format, and the converted text format is matched with the texts corresponding to multiple relevant voice commands to obtain the text with the highest matching degree corresponding to the current converted text format; Count the number of different characters between the current converted text format and the matching corresponding text and mark it as A1; Count the total number of words in the current converted text and the corresponding text respectively, and mark them as A2 and A3. Based on the comparison result of the two, calculate the ratio JG of the number of unmatched words; that is, through the formula Calculate and obtain.

3. The voice interaction system for power grid operation and maintenance based on artificial intelligence according to claim 2, characterized in that, The specific steps to obtain the error condition value JM are as follows: Based on the exact word segmentation of the thesaurus, perform word segmentation on the current converted text format and the matching corresponding text; compare the words at the word level after word segmentation in the following way: Count the number of words that appear in the converted text format but do not appear in the matching corresponding text to obtain the new value XZ; Count the number of words that appear in the matching corresponding text but do not appear in the converted text format to obtain the deletion value XY; Count the number of words with corresponding positions but inconsistent contents in the converted text format and the matching corresponding text to obtain the replacement value XM; Substitute the parameters obtained above into the formula , and calculate to obtain the error condition value JM; where a1, a2, and a3 are the influence weight factors of the new value XZ, the deleted value XY, and the replacement value XM, respectively.

4. The voice interaction system for power grid operation and maintenance based on artificial intelligence according to claim 3, wherein, The specific steps to obtain the difference value JW are as follows: S1: Create a matrix of (K + 1) * (H + 1), where K represents the length of the matching corresponding text and H represents the length of the converted text format; the rows of the matrix represent the characters of the matching corresponding text, and the columns represent the characters of the converted text format; S2: Fill the first row with numbers from 0 to H from 0 to H; fill the first column with numbers from 0 to K from 0 to K; S3: Use the method of dynamic programming to fill the created matrix; for each position (t, b) in the created matrix, which represents the edit distance between the first t characters of the matching corresponding text and the first b characters of the converted text, calculate based on the following different situations: If the t-th character of the matching corresponding text is the same as the b-th character of the converted text, the value of this position is equal to the value of the previous position in the matrix; If the t-th character of the matching corresponding text is different from the b-th character of the converted text, the value of this position is equal to the minimum of the value of the previous position in the matrix plus 1, the value of the adjacent position in the previous row plus 1, and the value of the adjacent position in the previous column plus 1; S4: Finally, the value in the lower right corner of the created matrix is the Levenshtein distance between the matching corresponding text and the converted text, representing the difference value JW between them.

5. The voice interaction system for power grid operation and maintenance based on artificial intelligence according to claim 4, characterized in that, Collect and recognize the gestures made by the operation and maintenance personnel in real time. The specific first step is: Use the sound source localization technology to determine the position where the current voice command is issued, and based on the positioning result, lock the operation and maintenance personnel closest to the positioning result, and control multiple groups of camera devices in the operation and maintenance environment to face the positioned operation and maintenance personnel; Collect the image information of the gestures made by the operation and maintenance personnel after the gesture recognition function is awakened through multiple groups of camera devices, and select the image information collected by the camera device corresponding to the body orientation of the operation and maintenance personnel as the analysis image; After preprocessing the image, segment it, extract the hand area of the operation and maintenance personnel in the image information, and rotate and scale the extracted hand area by angle and proportion to make it consistent with multiple groups of preset gesture images; Compare the adjusted hand area image with multiple groups of preset gesture images, calculate the overlapping area between the two, and obtain the area overlapping values of each group; select the preset gesture image corresponding to the maximum area overlapping value as the probability execution instruction; and calculate the ratio of the maximum area overlapping value to the total area of the hand area image to obtain the overlapping ratio; Based on the obtained overlapping area, determine the overlapping region in the two sets of images, and remove this region from the adjusted hand region image. Calculate the area of the remaining region after removal, compare the area of the remaining region with a preset remaining threshold, and obtain different weight coefficients based on the comparison result. If it is less than the preset remaining threshold, automatically match weight coefficient one; otherwise, match weight coefficient two. Calculate the ratio between the area of the remaining region and the total area of the hand region image, and multiply it by the corresponding weight coefficient to obtain the remaining ratio.

6. The voice interaction system for power grid operation and maintenance based on artificial intelligence according to claim 5, characterized in that, Collect and recognize the gestures made by the operation and maintenance personnel in real time. The specific step two is as follows: Mark the hand region image and the corresponding preset gesture image as R and T respectively. Calculate the means of the two sets of images, that is, the average value of all pixel values in each image, and mark them as R1 and T1. Use the variance formula to calculate the variance values of the two sets of images respectively, and mark them as R2 and T2. Calculate the cross-correlation coefficient NCC between the two sets of images, that is, through NCC = \frac{\sum_{x,y} (R(x,y)-R1)*(T(x,y)-T1)}{\sqrt{\sum_{x,y}(R(x,y)-RI)R2}*\sqrt{\sum_{x,y}(T(x,y)-T1)T2}}; Comprehensively analyze the coincidence ratio, the remaining ratio, and the cross-correlation coefficient to obtain the hand comparison accuracy value. Compare the calculated hand comparison accuracy value with the corresponding hand comparison threshold. If it is greater than the corresponding hand comparison threshold, it is determined that the recognition is successful, and directly execute the instruction corresponding to this gesture. If it is less than the corresponding hand comparison threshold, it is determined that the recognition is unsuccessful, and at the same time cancel the selected probability execution instruction. Recalculate the hand comparison accuracy value between each preset gesture image and this hand region image, and compare it with the corresponding hand comparison threshold. If there is no value greater than the hand comparison threshold, it is determined that the gesture recognition of the operation and maintenance personnel this time fails. If there is a value greater than the hand comparison threshold and the number of groups is multiple, select the largest value as the recognized successful instruction. At this time, the system responds to the specific content of this recognized successful instruction and asks whether to execute, that is, "The recognition instruction is XXX, do you want to execute?".

7. An artificial intelligence-based voice interaction system for power grid operation and maintenance according to claim 1, characterized in that, The specific process for analyzing the reception error of the current voice command content is as follows: Substitute the occupancy ratio JG, error condition value JM, and difference value JW corresponding to the current voice command content into the formula for calculation to obtain the reception difference JSC of the current voice command; Among them, F1, F2, and F3 respectively represent the maximum occupancy ratio, the maximum miscondition value, and the maximum difference value allowed for matching the corresponding text; va1, va2, and va3 are the influence weight factors of the occupancy ratio JG, the miscondition value JM, and the difference value JW respectively.

8. The voice interaction system for power grid operation and maintenance based on artificial intelligence according to claim 6, characterized in that, The specific process of comprehensive analysis of the coincidence ratio, remaining ratio, and cross-correlation coefficient is as follows: Substitute the coincidence ratio PC, remaining ratio PR, and cross-correlation coefficient NCC into the formula for calculation to obtain the hand ratio reference value SBZ; where nva, nvb, and nvc respectively represent the minimum allowable coincidence ratio, maximum allowable remaining ratio, and passing cross-correlation coefficient corresponding to the probability execution instruction; ty1, ty2, and ty3 are respectively the influence weight factors of the coincidence ratio PC, remaining ratio PR, and cross-correlation coefficient NCC.

Citation Information

Patent Citations

  • Smart pet collar with active noise reduction and voice interaction

    CN110754391A

  • Man-machine interaction management system based on intelligent set top box

    CN117409781A