Sound recognition processing method and device for smart power applications
Patent Information
- Application Number
- CN202510772303.3
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Filing Date
- 2025-06-10
- Publication Date
- 2026-09-01
- Estimated Expiration
- 2045-06-10
AI Technical Summary
[0003]然而,实践发现,在现有技术中,所需识别的语音信号在电力系统的应用中的采样率通常较低,而且监测到的语音信号可能受到环境噪声干扰以及双方同时说话等因素的影响,这些因素使得语音信号的质量较低,难以提取出明显的声纹特征,容易增加声纹识别的难度,影响声纹识别的效果和准确率
本发明实施例中,获取电力系统的当前操作人员的第一语音信号;基于预设语音信号处理算法,对第一语音信号进行处理,得到第一语音信号对应的第二语音信号,第二语音信号的质量高于第一语音信号的质量;对第二语音信号进行特征提取,得到第二语音信号对应的声纹特征信息;基于预设深度神经网络算法,对声纹特征信息进行识别,以完成对当前操作人员的身份验证。可见,实施本发明能够提高语音信号的质量从而能够提高声纹特征的提取准确性,有利于提高声纹识别的准确性从而有利于提高电力系统当前操作人员的身份验证准确性进而有利于提高电力系统操作的安全性。
Smart Images

Figure CN120600034B_ABST
Abstract
Description
Technical Field
[0001] This invention relates to the field of communication technology, and in particular to a voice recognition processing method and apparatus for smart power applications. Background Technology
[0002] With the global energy transition and the rapid development of new-generation information technology, cybersecurity issues in the power industry are becoming increasingly prominent. Therefore, identity verification is required to ensure the safe and stable operation of the system, such as through voiceprint recognition to verify the speaker's identity.
[0003] However, practical experience has shown that in existing technologies, the sampling rate of the voice signals to be identified in power system applications is typically low. Furthermore, the monitored voice signals may be affected by environmental noise interference and simultaneous speaking by both parties. These factors result in low voice signal quality, making it difficult to extract obvious voiceprint features, increasing the difficulty of voiceprint recognition, and affecting its effectiveness and accuracy. Therefore, proposing a new information processing scheme to improve the accuracy of identity verification and thus enhance the security of power system operations is particularly important. Summary of the Invention This invention provides a voice recognition processing method and apparatus for smart power applications, which can improve the quality of voice signals and the accuracy of voiceprint feature extraction, thereby improving the accuracy of voiceprint recognition, the accuracy of identity verification for power system operators, and the security of the power system.
[0004] To address the aforementioned technical problems, the first aspect of this invention discloses a sound recognition processing method for smart power applications, the method comprising: Acquire the first voice signal from the current operator of the power system; Based on a preset speech signal processing algorithm, the first speech signal is processed to obtain a second speech signal corresponding to the first speech signal; Feature extraction is performed on the second speech signal to obtain the voiceprint feature information corresponding to the second speech signal; Based on a preset deep neural network algorithm, the voiceprint feature information is identified to complete the identity verification of the current operator.
[0005] As an optional implementation, in the first aspect of the present invention, the step of extracting features from the second speech signal to obtain the voiceprint feature information corresponding to the second speech signal includes: Analyze the second speech signal to obtain the signal analysis results; Based on the signal analysis results, a target feature extraction method that meets the preset feature extraction conditions is selected from a plurality of preset feature extraction methods; Based on the target feature extraction method, feature extraction is performed on the second speech signal to obtain the voiceprint feature information corresponding to the second speech signal.
[0006] As an optional implementation, in the first aspect of the invention, the signal analysis result includes at least the number of users corresponding to the second voice signal; or, The signal analysis results include the number of users corresponding to the second speech signal, and also the speaking time period of each user; And, the step of selecting a target feature extraction method that meets the preset feature extraction conditions from a plurality of preset feature extraction methods based on the signal analysis results includes: Based on the number of users included in the signal analysis results, determine whether the number of users is less than or equal to a first preset number; When it is determined that the number of users is less than or equal to the first preset number, a first feature extraction method based on GMM-UBM is determined as the target feature extraction method that meets the preset feature extraction conditions. When it is determined that the number of users is greater than the first preset number, the speech time periods of all users included in the signal analysis result are compared to obtain the speech overlap time period of each user, and the duration of the speech overlap time period is greater than or equal to zero; and based on the number of users and the speech overlap time periods of all users, it is determined whether the current condition of the second voice signal meets the preset user scenario condition. When it is determined that the current conditions of the second voice signal meet the user scenario conditions, a second feature extraction method based on a deep neural network is determined as the target feature extraction method that meets the preset feature extraction conditions. When it is determined that the current conditions of the second voice signal meet the user scenario conditions, the third feature extraction method based on i-vector is determined as the target feature extraction method that meets the preset feature extraction conditions. Each of the aforementioned feature extraction methods includes one of the first feature extraction method, the second feature extraction method, and the third feature extraction method.
[0007] As an optional implementation, in the first aspect of the present invention, determining whether the current conditions of the second voice signal meet preset user scenario conditions based on the number of users and the overlapping time period of all the users' speech includes: Based on the number of users, determine whether the number of users is greater than or equal to a second preset number, and record this as the first condition; Based on the overlapping time periods of all the users' speech, determine whether there are at least a third preset number of users whose speech overlap time periods are greater than zero, and record this as the second condition; the first preset number is less than the third preset number, and the third preset number is less than the second preset number; Specifically, when the first condition and / or the second condition is met, it is determined that the current condition of the second voice signal meets the preset user scenario conditions; when neither the first condition nor the second condition is met, it is determined that the current condition of the second voice signal does not meet the preset user scenario conditions.
[0008] As an optional implementation, in the first aspect of the present invention, the step of identifying the voiceprint feature information based on a preset deep neural network algorithm to complete the identity verification of the current operator includes: Based on a voice recognition model trained by a preset deep neural network algorithm, the voiceprint feature information input to the voice recognition model is identified to obtain the probability score of the voiceprint feature information in different initial operators. The voice recognition model is obtained by training the voice sample information of all the initial operators using a deep neural network algorithm. The current operator is authenticated based on the probability score of the voiceprint feature information among all the initial operators.
[0009] As an optional implementation, in a first aspect of the invention, the step of verifying the identity of the current operator based on the probability score of the voiceprint feature information among all the initial operators includes: The probability scores of the voiceprint feature information among all the initial operators are compared to obtain the target operator with the highest probability score, and it is determined whether the probability score of the target operator is greater than or equal to a preset score. When it is determined that the probability score of the target operator is greater than or equal to the preset score, the identity verification of the current operator is confirmed to be successful. When it is determined that the probability score of the target operator is less than the preset score, it is determined that the identity verification of the current operator is unsuccessful.
[0010] As an optional implementation, in a first aspect of the invention, after verifying the identity of the current operator based on the probability score of the voiceprint feature information among all the initial operators, the method further includes: When verification fails, a verification failure message is output to the current operator. When the verification is successful, the identity tag of the target operator is determined as the identity tag of the current operator; Based on the identity tag of the current operator, the operation permission that matches the identity tag is determined from all the operation permissions of the initial operators pre-set in the power system, and is used as the target operation permission of the current operator. Each operation permission has a corresponding identity tag. Based on the target operation permission, the power system is controlled to unlock the functional modules corresponding to the target operation permission for the current operator, so that the current operator can perform corresponding control operations on the functional modules.
[0011] A second aspect of the present invention discloses a voice recognition processing device for smart power applications, the device comprising: The acquisition module is used to acquire the first voice signal of the current operator of the power system; The processing module is used to process the first speech signal based on a preset speech signal processing algorithm to obtain a second speech signal corresponding to the first speech signal. The extraction module is used to extract features from the second speech signal to obtain the voiceprint feature information corresponding to the second speech signal; The identification module is used to identify the voiceprint feature information based on a preset deep neural network algorithm in order to complete the identity verification of the current operator.
[0012] As an optional implementation, in the second aspect of the present invention, the extraction module performs feature extraction on the second speech signal to obtain the voiceprint feature information corresponding to the second speech signal in the following specific ways: Analyze the second speech signal to obtain the signal analysis results; Based on the signal analysis results, a target feature extraction method that meets the preset feature extraction conditions is selected from a plurality of preset feature extraction methods; Based on the target feature extraction method, feature extraction is performed on the second speech signal to obtain the voiceprint feature information corresponding to the second speech signal.
[0013] As an optional implementation, in the second aspect of the invention, the signal analysis result includes at least the number of users corresponding to the second voice signal; or, The signal analysis results include the number of users corresponding to the second speech signal, and also the speaking time period of each user; Furthermore, the extraction module selects the target feature extraction method that meets the preset feature extraction conditions from a plurality of preset feature extraction methods based on the signal analysis results, specifically including: Based on the number of users included in the signal analysis results, determine whether the number of users is less than or equal to a first preset number; When it is determined that the number of users is less than or equal to the first preset number, a first feature extraction method based on GMM-UBM is determined as the target feature extraction method that meets the preset feature extraction conditions. When it is determined that the number of users is greater than the first preset number, the speech time periods of all users included in the signal analysis result are compared to obtain the speech overlap time period of each user, and the duration of the speech overlap time period is greater than or equal to zero; and based on the number of users and the speech overlap time periods of all users, it is determined whether the current condition of the second voice signal meets the preset user scenario condition. When it is determined that the current conditions of the second voice signal meet the user scenario conditions, a second feature extraction method based on a deep neural network is determined as the target feature extraction method that meets the preset feature extraction conditions. When it is determined that the current conditions of the second voice signal meet the user scenario conditions, the third feature extraction method based on i-vector is determined as the target feature extraction method that meets the preset feature extraction conditions. Each of the aforementioned feature extraction methods includes one of the first feature extraction method, the second feature extraction method, and the third feature extraction method.
[0014] As an optional implementation, in a second aspect of the present invention, the extraction module determines whether the current conditions of the second voice signal meet preset user scenario conditions based on the number of users and the overlapping time period of all the users' speech, specifically including: Based on the number of users, determine whether the number of users is greater than or equal to a second preset number, and record this as the first condition; Based on the overlapping time periods of all the users' speech, determine whether there are at least a third preset number of users whose speech overlap time periods are greater than zero, and record this as the second condition; the first preset number is less than the third preset number, and the third preset number is less than the second preset number; Specifically, when the first condition and / or the second condition is met, it is determined that the current condition of the second voice signal meets the preset user scenario conditions; when neither the first condition nor the second condition is met, it is determined that the current condition of the second voice signal does not meet the preset user scenario conditions.
[0015] As an optional implementation, in a second aspect of the present invention, the identification module identifies the voiceprint feature information based on a preset deep neural network algorithm to complete the identity verification of the current operator. Specifically, this includes: Based on a voice recognition model trained by a preset deep neural network algorithm, the voiceprint feature information input to the voice recognition model is identified to obtain the probability score of the voiceprint feature information in different initial operators. The voice recognition model is obtained by training the voice sample information of all the initial operators using a deep neural network algorithm. The current operator is authenticated based on the probability score of the voiceprint feature information among all the initial operators.
[0016] As an optional implementation, in a second aspect of the present invention, the method by which the identification module verifies the identity of the current operator based on the probability score of the voiceprint feature information among all the initial operators specifically includes: The probability scores of the voiceprint feature information among all the initial operators are compared to obtain the target operator with the highest probability score, and it is determined whether the probability score of the target operator is greater than or equal to a preset score. When it is determined that the probability score of the target operator is greater than or equal to the preset score, the identity verification of the current operator is confirmed to be successful. When it is determined that the probability score of the target operator is less than the preset score, it is determined that the identity verification of the current operator is unsuccessful.
[0017] As an optional implementation, in a second aspect of the invention, the apparatus further includes: The output module is used to output a verification failure prompt message to the current operator when the verification fails after the recognition module verifies the identity of the current operator based on the probability score of the voiceprint feature information among all the initial operators. The determination module is used to determine the identity tag of the target operator as the identity tag of the current operator when the verification is successful; The determining module is further configured to determine, based on the identity tag of the current operator, the operation permission that matches the identity tag from all the operation permissions of the initial operators pre-set in the power system, and use it as the target operation permission of the current operator. Each operation permission has a corresponding identity tag. The control module is used to control the power system to unlock the functional modules corresponding to the target operation permissions for the current operator, so that the current operator can perform corresponding control operations on the functional modules.
[0018] A third aspect of the present invention discloses another voice recognition processing device for smart power applications, the device comprising: Memory containing executable program code; A processor coupled to the memory; The processor calls the executable program code stored in the memory to execute the voice recognition processing method for smart power applications disclosed in the first aspect of the present invention.
[0019] The fourth aspect of the present invention discloses a computer storage medium storing computer instructions, which, when invoked, are used to execute the voice recognition processing method for smart power applications disclosed in the first aspect of the present invention.
[0020] Compared with the prior art, the embodiments of the present invention have the following beneficial effects: In this embodiment of the invention, a first voice signal of the current operator of the power system is acquired; based on a preset voice signal processing algorithm, the first voice signal is processed to obtain a second voice signal corresponding to the first voice signal, the quality of the second voice signal being higher than that of the first voice signal; features are extracted from the second voice signal to obtain voiceprint feature information corresponding to the second voice signal; based on a preset deep neural network algorithm, the voiceprint feature information is identified to complete the identity verification of the current operator. It is evident that implementing this invention can improve the quality of the voice signal, thereby improving the accuracy of voiceprint feature extraction, which is beneficial to improving the accuracy of voiceprint recognition, thus improving the accuracy of identity verification of the current operator of the power system, and ultimately improving the security of power system operations. Attached Figure Description
[0021] To more clearly illustrate the technical solutions in the embodiments of the present invention, the accompanying drawings used in the description of the embodiments will be briefly introduced below. Obviously, the accompanying drawings described below are only some embodiments of the present invention. For those skilled in the art, other drawings can be obtained based on these drawings without creative effort.
[0022] Figure 1 This is a schematic flowchart of a sound recognition processing method for smart power applications disclosed in an embodiment of the present invention; Figure 2This is a flowchart illustrating another voice recognition processing method for smart power applications disclosed in an embodiment of the present invention; Figure 3 This is a schematic diagram of the structure of a voice recognition processing device for smart power applications disclosed in an embodiment of the present invention; Figure 4 This is a schematic diagram of another voice recognition processing device for smart power applications disclosed in an embodiment of the present invention; Figure 5 This is a schematic diagram of the structure of another voice recognition processing device for smart power applications disclosed in an embodiment of the present invention. Detailed Implementation
[0023] To enable those skilled in the art to better understand the present invention, the technical solutions of the present invention will be clearly and completely described below with reference to the accompanying drawings of the embodiments of the present invention. Obviously, the described embodiments are only some embodiments of the present invention, and not all embodiments. Based on the embodiments of the present invention, all other embodiments obtained by those skilled in the art without creative effort are within the scope of protection of the present invention.
[0024] The terms "first," "second," etc., used in the specification, claims, and accompanying drawings of this invention are used to distinguish different objects, not to describe a specific order. Furthermore, the terms "comprising" and "having," and any variations thereof, are intended to cover non-exclusive inclusion. For example, a process, method, apparatus, product, or end that includes a series of steps or units is not limited to the listed steps or units, but may optionally include steps or units not listed, or may optionally include other steps or units inherent to these processes, methods, products, or ends.
[0025] In this document, the term "embodiment" means that a particular feature, structure, or characteristic described in connection with an embodiment may be included in at least one embodiment of the invention. The appearance of this phrase in various places throughout the specification does not necessarily refer to the same embodiment, nor is it a separate or alternative embodiment mutually exclusive with other embodiments. It will be explicitly and implicitly understood by those skilled in the art that the embodiments described herein can be combined with other embodiments.
[0026] This invention discloses a voice recognition processing method and apparatus for smart power applications. It improves the quality of voice signals, thereby enhancing the accuracy of voiceprint feature extraction. This, in turn, improves the accuracy of voiceprint recognition, which in turn improves the accuracy of identity verification for current power system operators, ultimately enhancing the security of power system operations. Detailed descriptions follow.
[0027] Example 1 Please see Figure 1 , Figure 1 This is a flowchart illustrating a voice recognition processing method for smart power applications disclosed in an embodiment of the present invention. Figure 1 The described voice recognition processing method for smart power applications can be applied to voice recognition processing devices for smart power applications. These devices may include recognition processing equipment or recognition processing servers, where the recognition processing server may include a cloud server or a local server; this embodiment of the invention does not impose limitations. Figure 1 As shown, the voice recognition processing method applied to smart power applications may include the following operations: 101. Obtain the first voice signal from the current operator of the power system.
[0028] In this embodiment of the invention, specifically, when the current operator triggers an authentication command for the power system, the power system outputs a voiceprint authentication prompt to the current operator to trigger the current operator to perform voiceprint authentication, and the power system receives the voice signal issued by the current operator as the current operator's first voice signal.
[0029] 102. Based on a preset speech signal processing algorithm, the first speech signal is processed to obtain the second speech signal corresponding to the first speech signal.
[0030] In this embodiment of the invention, the quality of the second speech signal is higher than that of the first speech signal. Optionally, the preset speech signal processing algorithm can be an audio enhancement algorithm based on machine learning or digital signal processing (DSP) technology to enhance the low-quality audio signal (i.e., the first speech signal) and thus improve the recognition ability of the audio signal. Alternatively, it can be an algorithm of a deep learning model (such as GAN). Based on the low-quality audio signal, the algorithm of the deep learning model is used to synthesize a high-quality audio signal (i.e., the second speech signal) to facilitate subsequent voiceprint recognition training. This embodiment of the invention does not limit the scope of the invention.
[0031] Specifically, after acquiring the first voice signal of the current operator of the power system, the method may further include: Analyze the voice content of the first voice signal and determine whether the voice content matches the pre-set verification content in the power system; When it is determined that the voice content does not match the pre-set verification content in the power system, an error message is output to the current operator, and the voice content is marked as abnormal in the power system. When the voice content is determined to match the pre-set verification content in the power system, step 102 is triggered to process the first voice signal using a preset voice signal processing algorithm to obtain the corresponding second voice signal. Specifically, a match is indicated when the voice content is identical to the verification content, or when the content similarity between the voice content and the verification content is greater than or equal to a preset similarity (e.g., 90%). Conversely, a mismatch is indicated when the content similarity is less than the preset similarity. This dual verification process, involving both voice content verification and voiceprint feature verification, improves the accuracy and reliability of operator identity verification, thereby enhancing the security of the power system.
[0032] 103. Perform feature extraction on the second speech signal to obtain the voiceprint feature information corresponding to the second speech signal.
[0033] In this embodiment of the invention, specifically, based on the determined target feature extraction method, feature extraction is performed on the second speech signal to obtain the voiceprint feature information corresponding to the second speech signal. The target feature extraction method can be obtained by selecting a preset feature extraction method from a plurality of preset feature extraction methods based on the signal analysis results of the second speech signal, and selecting the method that meets preset feature extraction conditions.
[0034] 104. Based on a preset deep neural network algorithm, voiceprint feature information is identified to complete the identity verification of the current operator.
[0035] It is evident that implementation Figure 1The described voice recognition processing method for smart power applications can acquire the first voice signal of the current operator in the power system, and process the first voice signal to obtain the corresponding second voice signal based on a preset voice signal processing algorithm. Subsequently, feature extraction is performed on the second voice signal to obtain the corresponding voiceprint feature information. By processing the first voice signal, a higher quality second voice signal can be obtained, which can improve the quality and accuracy of the voice signal, thereby improving the accuracy of voiceprint feature information extraction. Then, based on a preset deep neural network algorithm, the voiceprint feature information is identified to complete the identity verification of the current operator. This is beneficial to improving the accuracy of voiceprint feature information recognition, thereby improving the accuracy and security of identity verification of the current operator in the power system, enabling accurate identification of user identities, and further improving the safety and effectiveness of power system operations. It also helps to improve the security protection level of the power system and ensure the stable and safe operation of the power system.
[0036] In an optional embodiment, step 103 above, which involves extracting features from the second speech signal to obtain the voiceprint feature information corresponding to the second speech signal, may include: Analyze the second speech signal to obtain the signal analysis results; Based on the signal analysis results, the target feature extraction method that meets the preset feature extraction conditions is selected from multiple preset feature extraction methods. Based on the target feature extraction method, feature extraction is performed on the second speech signal to obtain the voiceprint feature information corresponding to the second speech signal.
[0037] In this embodiment of the invention, optionally, the signal analysis result includes at least the number of users corresponding to the second voice signal; or, the signal analysis result includes the number of users corresponding to the second voice signal, and also includes the speaking time period of each user, wherein the users include at least the current operator. Specifically, when the power system receives the voice signal emitted by the current operator, the actual source of the received voice signal includes at least the current operator, and may also include other personnel in the scene where the current operator is located; this embodiment of the invention does not impose any limitations.
[0038] Optionally, each feature extraction method may include one of the following: a first feature extraction method based on GMM-UBM (where GMM represents Gaussian Mixture Model, UBM represents Universal Background Model, and GMM-UBM represents a model combining Gaussian Mixture Model and Universal Background Model), a second feature extraction method based on deep neural networks (such as DNN, CNN, RNN, etc.), and a third feature extraction method based on i-vector (where i-vector is a technique for speaker identification, mainly used to extract the speaker's feature vector). This embodiment of the invention does not limit the method.
[0039] As can be seen, this optional embodiment can analyze the second speech signal to obtain signal analysis results, and based on the signal analysis results, select the target feature extraction method that meets the preset feature extraction conditions from a plurality of preset feature extraction methods. This can improve the accuracy and efficiency of the selection of target feature extraction methods for the second speech signal. Subsequently, based on the target feature extraction method, feature extraction is performed on the second speech signal to obtain the voiceprint feature information corresponding to the second speech signal. This can improve the accuracy and efficiency of voiceprint feature information extraction based on the accurately selected target feature extraction method, and by providing a variety of feature extraction methods, it is beneficial to improve the diversity and flexibility of voiceprint feature extraction for the second speech signal.
[0040] In this optional embodiment, as an optional implementation method, selecting a target feature extraction method that meets the preset feature extraction conditions from a plurality of preset feature extraction methods based on the signal analysis results may include: Based on the number of users included in the signal analysis results, determine whether the number of users is less than or equal to the first preset number; When it is determined that the number of users is less than or equal to the first preset number, the first feature extraction method based on GMM-UBM is determined as the target feature extraction method that meets the preset feature extraction conditions. When it is determined that the number of users is greater than the first preset number, the speech time periods of all users included in the signal analysis results are compared to obtain the speech overlap time period of each user. The length of the speech overlap time period is greater than or equal to zero. Based on the number of users and the speech overlap time periods of all users, it is determined whether the current conditions of the second voice signal meet the preset user scenario conditions. When it is determined that the current conditions of the second speech signal meet the user scenario conditions, the second feature extraction method based on the deep neural network is determined as the target feature extraction method that meets the preset feature extraction conditions. When it is determined that the current conditions of the second speech signal meet the user scenario conditions, the third feature extraction method based on i-vector is selected as the target feature extraction method that meets the preset feature extraction conditions.
[0041] In this embodiment of the invention, optionally, preset feature extraction conditions are used to indicate that the environment in which the speaker corresponding to the corresponding voice signal is located is a corresponding speaking environment (e.g., few people, multi-person dialogue, etc.). User scenario conditions are used to indicate that the scenario complexity of the environment in which the speaker corresponding to the corresponding voice signal is located is within a preset scenario complexity range. The speaking overlap time period of all users is used as the basis for determining the scenario complexity of the environment in which the speaker corresponding to the corresponding voice signal is located. Specifically, when the duration of a user's speaking overlap time period is zero, it indicates that the user is not speaking simultaneously with other users. When the duration of a user's speaking overlap time period is greater than zero, it indicates that the user is speaking simultaneously with other users. This embodiment of the invention does not impose any limitations.
[0042] For example, when the number of users corresponding to the second speech signal is small, the first feature extraction method based on GMM-UBM can be directly selected as the target feature extraction method for the second speech signal. In this way, when the data is small, selecting the GMM-UBM-based method can ensure that the extracted voiceprint features perform well and have good interpretability. When the number of users corresponding to the second speech signal is large, the complexity of the scene in which the speaker is located can be further analyzed to determine whether to select the second feature extraction method based on deep neural networks or the third feature extraction method based on i-vectors as the target feature extraction method for the second speech signal. In this way, when the data is large, selecting the second or third feature extraction method can help improve the accuracy of speaker identification, thereby improving the accuracy and reliability of subsequent identification of the current operator.
[0043] As can be seen, this optional implementation can determine whether the number of users is less than or equal to a first preset number based on the number of users included in the signal analysis results. When the number of users is determined to be less than or equal to the first preset number, a first feature extraction method based on GMM-UBM is determined as the target feature extraction method that meets the preset feature extraction conditions. This improves the accuracy and efficiency of determining the first feature extraction method based on GMM-UBM as the target feature extraction method, and is beneficial to improving the extraction efficiency and speed of the voiceprint features of the second speech signal through the first feature extraction method. When the number of users is determined to be greater than the first preset number, the speech overlap time period of each user is obtained by comparing the speech time periods of all users included in the signal analysis results. The duration of the speech overlap time period is greater than or equal to zero. Based on the number of users and the speech overlap time periods of all users, it is determined whether the current conditions of the second speech signal meet the preset user scenario conditions. This can be achieved when the number of users is relatively large. In many cases, the accuracy and efficiency of determining whether the second speech signal meets the user scenario conditions are improved by considering the overlapping time period of all users' speech. When it is determined that the current conditions of the second speech signal meet the user scenario conditions, the second feature extraction method based on deep neural networks is determined as the target feature extraction method that meets the preset feature extraction conditions. This improves the accuracy and reliability of determining the second feature extraction method based on deep neural networks as the target feature extraction method, which is conducive to improving the recognition accuracy of the speaker corresponding to the second speech signal by the deep neural network model. When it is determined that the current conditions of the second speech signal meet the user scenario conditions, the third feature extraction method based on i-vector is determined as the target feature extraction method that meets the preset feature extraction conditions. This improves the accuracy and reliability of determining the third feature extraction method based on i-vector as the target feature extraction method, which is conducive to achieving accurate speaker identification.
[0044] In this optional implementation, determining whether the current conditions of the second voice signal meet the preset user scenario conditions based on the number of users and the overlapping time period of all users' speech may include: Based on the number of users, determine whether the number of users is greater than or equal to the second preset number, and record it as the first condition; Based on the overlapping time periods of all users' speech, determine whether there are at least three preset numbers of users whose overlapping time periods are greater than zero, which is recorded as the second condition; the first preset number is less than the third preset number, and the third preset number is less than the second preset number. Specifically, when the first condition and / or the second condition are met, it is determined that the current condition of the second voice signal meets the preset user scenario conditions; when neither the first condition nor the second condition is met, it is determined that the current condition of the second voice signal does not meet the preset user scenario conditions.
[0045] In this embodiment of the invention, specifically, when it is determined that the number of users is greater than or equal to a second preset number, and / or when it is determined that the duration of the overlapping speaking time of at least a third preset number of users is greater than zero, it indicates that the scene complexity of the environment in which the speaker corresponding to the second voice signal is located is within a preset complexity range. At this time, a second feature extraction method that can effectively improve recognition accuracy on a large dataset can be selected for voiceprint extraction. When it is determined that the number of users is less than the second preset number and less than the third preset number of users, the duration of the overlapping speaking time of the second voice signal is greater than zero, it indicates that the scene complexity of the environment in which the speaker corresponding to the second voice signal is located is not within the preset complexity range. At this time, a third feature extraction method with relatively low computational complexity and suitable for multi-speaker environments can be selected for voiceprint extraction.
[0046] For example, assuming the first preset quantity is 1, the second preset quantity is 5, and the third preset quantity is 3, when it is determined that the number of users corresponding to the second voice signal is greater than or equal to 5, and / or at least 3 users are speaking simultaneously with other users, it can be determined that the current condition of the second voice signal meets the preset user scenario condition; when it is determined that the number of users corresponding to the second voice signal is less than 5, and less than 3 users are speaking simultaneously with other users, it can be determined that the current condition of the second voice signal does not meet the preset user scenario condition.
[0047] As can be seen, this optional implementation can also determine whether the number of users is greater than or equal to a second preset number, denoted as the first condition, and determine whether there are at least a third preset number of users whose speaking overlap time is greater than zero, denoted as the second condition, based on the speaking overlap time of all users. When the first condition and / or the second condition are met, it is determined that the current condition of the second voice signal meets the preset user scenario condition. This can improve the accuracy, diversity, and flexibility of determining whether the current condition of the second voice signal meets the preset user scenario condition by providing a variety of scenario judgment methods. When neither the first nor the second condition is met, it is determined that the current condition of the second voice signal does not meet the preset user scenario condition. This can improve the accuracy and reliability of determining whether the current condition of the second voice signal does not meet the preset user scenario condition by using multiple judgment methods of the user's scenario. This is beneficial for improving the accuracy and reliability of subsequent feature extraction methods based on accurate judgment results.
[0048] Example 2 Please see Figure 2 , Figure 2 This is a flowchart illustrating a voice recognition processing method for smart power applications disclosed in an embodiment of the present invention. Figure 2The described voice recognition processing method for smart power applications can be applied to voice recognition processing devices for smart power applications. These devices may include recognition processing equipment or recognition processing servers, where the recognition processing server may include a cloud server or a local server; this embodiment of the invention does not impose limitations. Figure 2 As shown, the voice recognition processing method applied to smart power applications may include the following operations: 201. Obtain the first voice signal from the current operator of the power system.
[0049] 202. Based on a preset speech signal processing algorithm, the first speech signal is processed to obtain the second speech signal corresponding to the first speech signal.
[0050] 203. Perform feature extraction on the second speech signal to obtain the voiceprint feature information corresponding to the second speech signal.
[0051] In this embodiment of the invention, for other descriptions of steps 201-203, please refer to the detailed description of steps 101-103 in Embodiment 1. These descriptions will not be repeated in this embodiment of the invention.
[0052] 204. Based on the voice recognition model trained by the preset deep neural network algorithm, the voiceprint feature information input to the voice recognition model is identified to obtain the probability score of the voiceprint feature information in different initial operators.
[0053] In this embodiment of the invention, the voice recognition model is trained using a deep neural network algorithm on the voice sample information of all initial operators. Specifically, features are extracted from the voice sample information of each initial operator to obtain the voiceprint feature sample information of each initial operator; the voiceprint feature sample information of all initial operators is then input into a pre-constructed model based on a preset deep neural network algorithm for training, and the training results are obtained; based on the training results, the recognition accuracy of the model is evaluated; and it is determined whether the recognition accuracy of the model reaches a preset recognition accuracy. When it is determined that the preset recognition accuracy has been reached, the model is determined as the voice recognition model trained based on the preset deep neural network algorithm; when it is determined that the preset recognition accuracy has not been reached, the model training operation is re-executed until the recognition accuracy of the trained model reaches the preset recognition accuracy.
[0054] 205. Verify the identity of the current operator based on the probability score of voiceprint feature information among all initial operators.
[0055] In this embodiment of the invention, optionally, the initial operator with the highest probability score among all initial operators is identified as the current operator, thereby completing the identity verification of the current operator and improving the efficiency and speed of the identity verification of the current operator; or, the initial operator with the highest probability score is selected from all initial operators as the target operator suspected of being the current operator, and then the probability score of the target operator is used to further determine whether the current operator is the target operator, thereby completing the identity verification of the current operator and improving the accuracy and reliability of the identity verification of the current operator.
[0056] It is evident that implementation Figure 2 The described voice recognition processing method for smart power applications can acquire the first voice signal of the current operator in the power system, and process the first voice signal to obtain the corresponding second voice signal based on a preset voice signal processing algorithm. Subsequently, feature extraction is performed on the second voice signal to obtain the corresponding voiceprint feature information. By processing the first voice signal, a higher quality second voice signal can be obtained, which can improve the quality and accuracy of the voice signal, thereby improving the accuracy of voiceprint feature information extraction. Then, based on a preset deep neural network algorithm, the voiceprint feature information is identified to complete the identity verification of the current operator. This is beneficial to improving the accuracy of voiceprint feature information recognition, thereby improving the accuracy and security of identity verification of the current operator in the power system, enabling accurate identification of user identities, and further improving the safety and effectiveness of power system operations. It also helps to improve the security protection level of the power system and ensure the stable and safe operation of the power system. Furthermore, based on a voice recognition model trained by a pre-defined deep neural network algorithm, the system can identify the voiceprint feature information input to the voice recognition model and obtain the probability score of the voiceprint feature information among different initial operators. This enables accurate identification of voiceprint feature information based on the voice recognition model, which helps improve the accuracy and efficiency of obtaining the probability score of voiceprint feature information among different initial operators. Subsequently, based on the probability score of the voiceprint feature information among all initial operators, the system can verify the identity of the current operator. This allows for accurate identity verification of the current operator based on the accurately identified probability score, which helps improve the accuracy and efficiency of identity verification for the current operator.
[0057] In an optional embodiment, step 205 above, which involves verifying the identity of the current operator based on the probability score of the voiceprint feature information among all initial operators, may include: Compare the probability scores of voiceprint feature information among all initial operators to obtain the target operator with the highest probability score, and determine whether the probability score of the target operator is greater than or equal to the preset score. When the probability score of the target operator is determined to be greater than or equal to the preset score, the identity verification of the current operator is confirmed to be successful. When the probability score of the target operator is determined to be less than the preset score, the identity verification of the current operator is deemed unsuccessful.
[0058] In this embodiment of the invention, optionally, the probability score of voiceprint feature information in each initial operator can be in the form of a percentage. The preset score can be 70%, 80%, or any other pre-set value. This embodiment of the invention does not limit the score.
[0059] As can be seen, this optional embodiment can compare the probability scores of voiceprint feature information among all initial operators to obtain the target operator with the highest probability score, and determine whether the probability score of the target operator is greater than or equal to a preset score. When it is determined that the probability score of the target operator is greater than or equal to the preset score, it is determined that the current operator's identity verification is successful, which improves the accuracy and reliability of determining whether the current operator's identity verification is successful. When it is determined that the probability score of the target operator is less than the preset score, it is determined that the current operator's identity verification is unsuccessful, which improves the accuracy and reliability of determining whether the current operator's identity verification is unsuccessful.
[0060] In this optional embodiment, as an optional implementation, after verifying the identity of the current operator based on the probability score of voiceprint feature information among all initial operators, the method may further include: If the verification fails, output a verification failure message to the current operator; When the verification is successful, the identity tag of the target operator will be set as the identity tag of the current operator. Based on the current operator's identity tag, the operation permission that matches the identity tag is determined from all the operation permissions of the initial operators pre-set in the power system, and is used as the target operation permission of the current operator. Each operation permission has a corresponding identity tag. Based on the target operating permissions, the control power system unlocks the corresponding functional modules for the current operator, enabling the current operator to perform corresponding control operations on the functional modules.
[0061] In this embodiment of the invention, multiple initial operator permissions are pre-set in the power system, and each initial operator corresponds to an identity tag, which is associated with the corresponding initial operator's operating permissions. This allows the corresponding operating permissions of the initial operator to be directly determined through the initial operator's identity tag. The power system may be equipped with multiple functional modules that can be managed by operators, such as at least one of the following: a monitoring and data acquisition module, a scheduling and optimization module, a fault detection and protection module, an equipment management and maintenance module, a load management and demand response module, and a communication and data management module.
[0062] As can be seen, this optional implementation can directly output a verification failure message to the current operator when the current operator's identity verification fails, which is beneficial for prompting the current operator to re-verify and facilitates the re-verification process. If the current operator's identity verification is successful, the target operator's identity tag is determined as the current operator's identity tag. Based on the current operator's identity tag, the operation permissions matching the identity tag are determined from all the pre-set operation permissions of initial operators in the power system, serving as the target operation permissions for the current operator. This improves the accuracy of determining the current operator's identity tag, thereby improving the accuracy and efficiency of determining the target operation permissions. Subsequently, based on the target operation permissions, the power system unlocks the corresponding functional modules for the current operator, enabling the current operator to perform corresponding control operations on the functional modules. This achieves accurate granting of corresponding operation permissions to the current operator and unlocking of corresponding functional modules, facilitating the current operator's performance of corresponding control operations on the power system's functional modules and improving the security of power system operations.
[0063] Example 3 Please see Figure 3 , Figure 3 This is a schematic diagram of the structure of a voice recognition processing device for smart power applications disclosed in an embodiment of the present invention. Figure 3 The described voice recognition processing device for smart power applications may include a recognition processing device or a recognition processing server, wherein the recognition processing server may include a cloud server or a local server, and the embodiments of the present invention are not limited thereto. Figure 3 As shown, the voice recognition processing device for smart power applications may include: The acquisition module 301 is used to acquire the first voice signal of the current operator of the power system.
[0064] The processing module 302 is used to process the first speech signal based on a preset speech signal processing algorithm to obtain the second speech signal corresponding to the first speech signal.
[0065] The extraction module 303 is used to extract features from the second speech signal to obtain the voiceprint feature information corresponding to the second speech signal.
[0066] The identification module 304 is used to identify voiceprint feature information based on a preset deep neural network algorithm in order to complete the identity verification of the current operator.
[0067] It is evident that implementation Figure 3 The described voice recognition processing device for smart power applications can acquire the first voice signal of the current operator in the power system, and process the first voice signal to obtain a second voice signal corresponding to the first voice signal based on a preset voice signal processing algorithm. Subsequently, feature extraction is performed on the second voice signal to obtain the corresponding voiceprint feature information. By processing the first voice signal to obtain a higher quality second voice signal, the quality and accuracy of the voice signal can be improved, thereby improving the accuracy of voiceprint feature information extraction. Then, based on a preset deep neural network algorithm, the voiceprint feature information is identified to complete the identity verification of the current operator. This is beneficial to improving the accuracy of voiceprint feature information recognition, thereby improving the accuracy and security of identity verification of the current operator in the power system, enabling accurate identification of user identities, and further improving the safety and effectiveness of power system operations. It also helps to improve the security protection level of the power system and ensure the stable and safe operation of the power system.
[0068] In an optional embodiment, the extraction module 303 performs feature extraction on the second speech signal to obtain the voiceprint feature information corresponding to the second speech signal. Specifically, this may include: Analyze the second speech signal to obtain the signal analysis results; Based on the signal analysis results, the target feature extraction method that meets the preset feature extraction conditions is selected from multiple preset feature extraction methods. Based on the target feature extraction method, feature extraction is performed on the second speech signal to obtain the voiceprint feature information corresponding to the second speech signal.
[0069] As can be seen, this optional embodiment can analyze the second speech signal to obtain signal analysis results, and based on the signal analysis results, select the target feature extraction method that meets the preset feature extraction conditions from a plurality of preset feature extraction methods. This can improve the accuracy and efficiency of the selection of target feature extraction methods for the second speech signal. Subsequently, based on the target feature extraction method, feature extraction is performed on the second speech signal to obtain the voiceprint feature information corresponding to the second speech signal. This can improve the accuracy and efficiency of voiceprint feature information extraction based on the accurately selected target feature extraction method, and by providing a variety of feature extraction methods, it is beneficial to improve the diversity and flexibility of voiceprint feature extraction for the second speech signal.
[0070] In this optional embodiment, as an optional implementation, the signal analysis result includes at least the number of users corresponding to the second speech signal; or, the signal analysis result includes the number of users corresponding to the second speech signal, and also includes the speaking time period of each user. Furthermore, the extraction module 303, based on the signal analysis result, selects a target feature extraction method that meets the preset feature extraction conditions from a plurality of preset feature extraction methods. Specifically, this may include: Based on the number of users included in the signal analysis results, determine whether the number of users is less than or equal to the first preset number; When it is determined that the number of users is less than or equal to the first preset number, the first feature extraction method based on GMM-UBM is determined as the target feature extraction method that meets the preset feature extraction conditions. When it is determined that the number of users is greater than the first preset number, the speech time periods of all users included in the signal analysis results are compared to obtain the speech overlap time period of each user. The length of the speech overlap time period is greater than or equal to zero. Based on the number of users and the speech overlap time periods of all users, it is determined whether the current conditions of the second voice signal meet the preset user scenario conditions. When it is determined that the current conditions of the second speech signal meet the user scenario conditions, the second feature extraction method based on the deep neural network is determined as the target feature extraction method that meets the preset feature extraction conditions. When it is determined that the current conditions of the second speech signal meet the user scenario conditions, the third feature extraction method based on i-vector is determined as the target feature extraction method that meets the preset feature extraction conditions. Each feature extraction method includes one of the first feature extraction method, the second feature extraction method, and the third feature extraction method.
[0071] As can be seen, this optional implementation can determine whether the number of users is less than or equal to a first preset number based on the number of users included in the signal analysis results. When the number of users is determined to be less than or equal to the first preset number, a first feature extraction method based on GMM-UBM is determined as the target feature extraction method that meets the preset feature extraction conditions. This improves the accuracy and efficiency of determining the first feature extraction method based on GMM-UBM as the target feature extraction method, and is beneficial to improving the extraction efficiency and speed of the voiceprint features of the second speech signal through the first feature extraction method. When the number of users is determined to be greater than the first preset number, the speech overlap time period of each user is obtained by comparing the speech time periods of all users included in the signal analysis results. The duration of the speech overlap time period is greater than or equal to zero. Based on the number of users and the speech overlap time periods of all users, it is determined whether the current conditions of the second speech signal meet the preset user scenario conditions. This can be achieved when the number of users is relatively large. In many cases, the accuracy and efficiency of determining whether the second speech signal meets the user scenario conditions are improved by considering the overlapping time period of all users' speech. When it is determined that the current conditions of the second speech signal meet the user scenario conditions, the second feature extraction method based on deep neural networks is determined as the target feature extraction method that meets the preset feature extraction conditions. This improves the accuracy and reliability of determining the second feature extraction method based on deep neural networks as the target feature extraction method, which is conducive to improving the recognition accuracy of the speaker corresponding to the second speech signal by the deep neural network model. When it is determined that the current conditions of the second speech signal meet the user scenario conditions, the third feature extraction method based on i-vector is determined as the target feature extraction method that meets the preset feature extraction conditions. This improves the accuracy and reliability of determining the third feature extraction method based on i-vector as the target feature extraction method, which is conducive to achieving accurate speaker identification.
[0072] In this optional implementation, the extraction module 303 may determine whether the current conditions of the second voice signal meet the preset user scenario conditions based on the number of users and the overlapping time period of all users' speech. Based on the number of users, determine whether the number of users is greater than or equal to the second preset number, and record it as the first condition; Based on the overlapping time periods of all users' speech, determine whether there are at least three preset numbers of users whose overlapping time periods are greater than zero, which is recorded as the second condition; the first preset number is less than the third preset number, and the third preset number is less than the second preset number. Specifically, when the first condition and / or the second condition are met, it is determined that the current condition of the second voice signal meets the preset user scenario conditions; when neither the first condition nor the second condition is met, it is determined that the current condition of the second voice signal does not meet the preset user scenario conditions.
[0073] As can be seen, this optional implementation can also determine whether the number of users is greater than or equal to a second preset number, denoted as the first condition, and determine whether there are at least a third preset number of users whose speaking overlap time is greater than zero, denoted as the second condition, based on the speaking overlap time of all users. When the first condition and / or the second condition are met, it is determined that the current condition of the second voice signal meets the preset user scenario condition. This can improve the accuracy, diversity, and flexibility of determining whether the current condition of the second voice signal meets the preset user scenario condition by providing a variety of scenario judgment methods. When neither the first nor the second condition is met, it is determined that the current condition of the second voice signal does not meet the preset user scenario condition. This can improve the accuracy and reliability of determining whether the current condition of the second voice signal does not meet the preset user scenario condition by using multiple judgment methods of the user's scenario. This is beneficial for improving the accuracy and reliability of subsequent feature extraction methods based on accurate judgment results.
[0074] In another optional embodiment, the identification module 304 identifies voiceprint feature information based on a preset deep neural network algorithm to complete the identity verification of the current operator. Specifically, this method may include: Based on the voice recognition model trained by the preset deep neural network algorithm, the voiceprint feature information input to the voice recognition model is identified to obtain the probability score of the voiceprint feature information in different initial operators. The voice recognition model is obtained by training the voice sample information of all initial operators through the deep neural network algorithm. The current operator is identified based on the probability score of voiceprint feature information among all initial operators.
[0075] As can be seen, this optional embodiment can identify the voiceprint feature information input to the voice recognition model based on a preset deep neural network algorithm, and obtain the probability score of the voiceprint feature information in different initial operators. It can achieve accurate identification of voiceprint feature information based on the voice recognition model, which is beneficial to improving the accuracy and efficiency of identifying the probability score of voiceprint feature information in different initial operators. Subsequently, based on the probability score of voiceprint feature information in all initial operators, the current operator is identified. It can accurately identify the current operator based on the accurately identified probability score, which is beneficial to improving the accuracy and efficiency of identifying the current operator.
[0076] In this optional embodiment, as an optional implementation, the identification module 304 may specifically verify the identity of the current operator based on the probability score of voiceprint feature information among all initial operators in the following ways: Compare the probability scores of voiceprint feature information among all initial operators to obtain the target operator with the highest probability score, and determine whether the probability score of the target operator is greater than or equal to the preset score. When the probability score of the target operator is determined to be greater than or equal to the preset score, the identity verification of the current operator is confirmed to be successful. When the probability score of the target operator is determined to be less than the preset score, the identity verification of the current operator is deemed unsuccessful.
[0077] As can be seen, this optional implementation can compare the probability scores of voiceprint feature information among all initial operators to obtain the target operator with the highest probability score, and determine whether the probability score of the target operator is greater than or equal to a preset score. When it is determined that the probability score of the target operator is greater than or equal to the preset score, it is determined that the current operator's identity verification is successful, which improves the accuracy and reliability of determining whether the current operator's identity verification is successful. When it is determined that the probability score of the target operator is less than the preset score, it is determined that the current operator's identity verification is unsuccessful, which improves the accuracy and reliability of determining whether the current operator's identity verification is unsuccessful.
[0078] In this optional implementation, alternatively, such as Figure 4 As shown, Figure 4 This is a schematic diagram of another voice recognition processing device for smart power applications disclosed in an embodiment of the present invention, as shown below. Figure 4 As shown, the device may further include: The output module 305 is used to output a verification failure prompt message to the current operator after the recognition module 304 verifies the identity of the current operator based on the probability score of the voiceprint feature information among all initial operators.
[0079] The determination module 306 is used to determine the identity tag of the target operator as the identity tag of the current operator when the verification is successful.
[0080] The determination module 306 is also used to determine the operation permission that matches the identity tag from all the operation permissions of the initial operators preset in the power system, based on the identity tag of the current operator, and use it as the target operation permission of the current operator. Each operation permission has a corresponding identity tag.
[0081] The control module 307 is used to control the power system to unlock the functional modules corresponding to the target operation permissions for the current operator, so that the current operator can perform corresponding control operations on the functional modules.
[0082] It is evident that implementation Figure 4The described device can also directly output a verification failure message to the current operator if the current operator's identity verification fails, which helps prompt the current operator to re-verify and facilitates the re-verification process. If the current operator's identity verification is successful, the target operator's identity tag is determined as the current operator's identity tag. Based on the current operator's identity tag, the device determines the operation permission matching the identity tag from all the pre-set operation permissions of all initial operators in the power system, which serves as the target operation permission for the current operator. This improves the accuracy of determining the current operator's identity tag, thereby improving the accuracy and efficiency of determining the target operation permission. Subsequently, based on the target operation permission, the device controls the power system to unlock the functional module corresponding to the target operation permission for the current operator, enabling the current operator to perform corresponding control operations on the functional module. This device can accurately grant the current operator the corresponding operation permission and unlock the corresponding functional module, which is beneficial for the current operator to perform corresponding control operations on the power system's functional modules and improves the security of power system operation.
[0083] Example 4 Please see Figure 5 , Figure 5 This is a schematic diagram of the structure of another voice recognition processing device for smart power applications disclosed in an embodiment of the present invention. Figure 5 As shown, the voice recognition processing device for smart power applications may include: Memory 401 storing executable program code; Processor 402 coupled to memory 401; The processor 402 calls the executable program code stored in the memory 401 to execute the steps in the voice recognition processing method for smart power applications described in Embodiment 1 or Embodiment 2 of the present invention.
[0084] Example 5 This invention discloses a computer storage medium storing computer instructions. When these computer instructions are invoked, they are used to execute the steps in the voice recognition processing method for smart power applications described in Embodiment 1 or Embodiment 2 of this invention.
[0085] Example 6 This invention discloses a computer program product, which includes a non-transitory computer-readable storage medium storing a computer program, and the computer program is operable to cause a computer to perform the steps in the voice recognition processing method for smart power applications described in Embodiment 1 or Embodiment 2.
[0086] The device embodiments described above are merely illustrative. The modules described as separate components may or may not be physically separate. The components shown as modules may or may not be physical modules; that is, they may be located in one place or distributed across multiple network modules. Some or all of the modules can be selected to achieve the purpose of this embodiment according to actual needs. Those skilled in the art can understand and implement this without any creative effort.
[0087] Through the detailed description of the above embodiments, those skilled in the art can clearly understand that each implementation method can be implemented by means of software plus necessary general-purpose hardware platforms, and of course, it can also be implemented by hardware. Based on this understanding, the above technical solutions, in essence or the part that contributes to the prior art, can be embodied in the form of a software product. This computer software product can be stored in a computer-readable storage medium, including read-only memory (ROM), random access memory (RAM), programmable read-only memory (PROM), erasable programmable read-only memory (EPROM), one-time programmable read-only memory (OTPROM), electrically-Erasable Programmable Read-Only Memory (EEPROM), compact disc read-only memory (CD-ROM) or other optical disc storage, disk storage, magnetic tape storage, or any other computer-readable medium that can be used to carry or store data.
[0088] Finally, it should be noted that the voice recognition processing method and apparatus for smart power applications disclosed in the embodiments of the present invention are merely preferred embodiments of the present invention and are only used to illustrate the technical solutions of the present invention, not to limit it. Although the present invention has been described in detail with reference to the foregoing embodiments, those skilled in the art should understand that modifications can still be made to the technical solutions described in the foregoing embodiments, or equivalent substitutions can be made to some of the technical features. Such modifications or substitutions do not cause the essence of the corresponding technical solutions to deviate from the spirit and scope of the technical solutions of the embodiments of the present invention.
Claims
1. A sound recognition processing method for smart power applications, characterized in that, The method includes: Acquire the first voice signal from the current operator of the power system; Based on a preset speech signal processing algorithm, the first speech signal is processed to obtain a second speech signal corresponding to the first speech signal; Analyze the second speech signal to obtain a signal analysis result, the signal analysis result including at least the number of users corresponding to the second speech signal; or, the signal analysis result including the number of users corresponding to the second speech signal, and also including the speaking time period of each user; Based on the number of users included in the signal analysis results, determine whether the number of users is less than or equal to a first preset number; When it is determined that the number of users is less than or equal to the first preset number, a first feature extraction method based on GMM-UBM is determined as the target feature extraction method that meets the preset feature extraction conditions. When it is determined that the number of users is greater than the first preset number, the speech time periods of all users included in the signal analysis result are compared to obtain the speech overlap time period of each user, and the duration of the speech overlap time period is greater than or equal to zero; and based on the number of users and the speech overlap time periods of all users, it is determined whether the current condition of the second voice signal meets the preset user scenario condition. When it is determined that the current conditions of the second voice signal meet the user scenario conditions, a second feature extraction method based on a deep neural network is determined as the target feature extraction method that meets the preset feature extraction conditions. When it is determined that the current conditions of the second voice signal do not meet the user scenario conditions, a third feature extraction method based on i-vector is determined as the target feature extraction method that meets the preset feature extraction conditions. Based on the target feature extraction method, feature extraction is performed on the second speech signal to obtain the voiceprint feature information corresponding to the second speech signal; Based on a preset deep neural network algorithm, the voiceprint feature information is identified to complete the identity verification of the current operator.
2. The voice recognition processing method for smart power applications according to claim 1, characterized in that, The step of determining whether the current conditions of the second voice signal meet the preset user scenario conditions based on the number of users and the overlapping time period of all the users' speech includes: Based on the number of users, determine whether the number of users is greater than or equal to a second preset number, and record this as the first condition; Based on the overlapping time periods of all the users' speech, determine whether there are at least a third preset number of users whose speech overlap time periods are greater than zero, and record this as the second condition; the first preset number is less than the third preset number, and the third preset number is less than the second preset number; Specifically, when the first condition and / or the second condition is met, it is determined that the current condition of the second voice signal meets the preset user scenario conditions; when neither the first condition nor the second condition is met, it is determined that the current condition of the second voice signal does not meet the preset user scenario conditions.
3. The voice recognition processing method for smart power applications according to claim 1 or 2, characterized in that, The process of identifying the voiceprint feature information based on a preset deep neural network algorithm to complete the identity verification of the current operator includes: Based on a voice recognition model trained by a preset deep neural network algorithm, the voiceprint feature information input to the voice recognition model is identified to obtain the probability score of the voiceprint feature information in different initial operators. The voice recognition model is obtained by training the voice sample information of all the initial operators using a deep neural network algorithm. The current operator is authenticated based on the probability score of the voiceprint feature information among all the initial operators.
4. The voice recognition processing method for smart power applications according to claim 3, characterized in that, The step of verifying the identity of the current operator based on the probability score of the voiceprint feature information among all the initial operators includes: The probability scores of the voiceprint feature information among all the initial operators are compared to obtain the target operator with the highest probability score, and it is determined whether the probability score of the target operator is greater than or equal to a preset score. When it is determined that the probability score of the target operator is greater than or equal to the preset score, the identity verification of the current operator is confirmed to be successful. When it is determined that the probability score of the target operator is less than the preset score, it is determined that the identity verification of the current operator is unsuccessful.
5. The voice recognition processing method for smart power applications according to claim 4, characterized in that, After verifying the identity of the current operator based on the probability score of the voiceprint feature information among all the initial operators, the method further includes: When verification fails, a verification failure message is output to the current operator. When the verification is successful, the identity tag of the target operator is determined as the identity tag of the current operator; Based on the identity tag of the current operator, the operation permission that matches the identity tag is determined from all the operation permissions of the initial operators pre-set in the power system, and is used as the target operation permission of the current operator. Each operation permission has a corresponding identity tag. Based on the target operation permission, the power system is controlled to unlock the functional modules corresponding to the target operation permission for the current operator, so that the current operator can perform corresponding control operations on the functional modules.
6. A voice recognition processing device for smart power applications, characterized in that, The apparatus is used to perform the voice recognition processing method for smart power applications as described in any one of claims 1-5, the apparatus comprising: The acquisition module is used to acquire the first voice signal of the current operator of the power system; The processing module is used to process the first speech signal based on a preset speech signal processing algorithm to obtain a second speech signal corresponding to the first speech signal. The extraction module is used to extract features from the second speech signal to obtain the voiceprint feature information corresponding to the second speech signal; The identification module is used to identify the voiceprint feature information based on a preset deep neural network algorithm in order to complete the identity verification of the current operator.
7. A voice recognition processing device for smart power applications, characterized in that, The device includes: Memory containing executable program code; A processor coupled to the memory; The processor calls the executable program code stored in the memory to execute the voice recognition processing method for smart power applications as described in any one of claims 1-5.
8. A computer storage medium, characterized in that, The computer storage medium stores computer instructions, which, when invoked, are used to execute the voice recognition processing method for smart power applications as described in any one of claims 1-5.
Citation Information
Patent Citations
Self-adapting method of DNN acoustic model based on personal identity characteristics
CN109637526A
Model generation method, voiceprint recognition method and corresponding device
CN110838295A