Voiceprint recognition method and device, medium and program product
By extracting acoustic features and clustering data from the acoustic fingerprint data of cold source station equipment, and combining this with reconstruction error judgment using an autoencoder model, the problems of low accuracy and poor robustness in acoustic fingerprint recognition of cold source station equipment were solved, thus achieving efficient equipment anomaly detection.
Patent Information
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-12-29
- Publication Date
- 2026-03-27
AI Technical Summary
Existing voiceprint recognition technologies have low accuracy and poor robustness in cold source station equipment, and are particularly difficult to effectively identify under complex and variable voiceprint characteristics and environmental noise.
By acquiring the original voiceprint data of the target device, acoustic features are extracted and clustered. A pre-trained autoencoder model is used to obtain the reconstruction error, and a device anomaly alarm is generated when the error exceeds the threshold.
It improves the accuracy and robustness of voiceprint recognition, enabling rapid response to changes in device status and reducing maintenance costs and workload.
Smart Images

Figure CN121747584A_ABST
Abstract
Description
Technical Field
[0001] This invention relates to the field of audio processing technology, and in particular to a voiceprint recognition method, device, medium, and program product. Background Technology
[0002] As an important component of large buildings and industrial facilities, the operation of the cooling station directly affects the energy efficiency and safety of the entire system.
[0003] Currently, existing voiceprint recognition methods typically employ identity verification vector methods. These methods first generate an identity vector corresponding to the voice signal to be identified, and then evaluate the similarity between this identity vector and the identity vector of a registered voice signal to achieve identity recognition. However, the voiceprint characteristics of cold source station equipment are complex and variable, and easily affected by environmental noise. Existing technologies, facing these complex and variable voiceprint characteristics and strong environmental noise, are prone to low recognition accuracy and poor robustness. Summary of the Invention
[0004] This invention provides a voiceprint recognition method, device, medium, and program product, which can improve the accuracy and robustness of voiceprint recognition.
[0005] According to one aspect of the present invention, a voiceprint recognition method is provided, comprising:
[0006] Obtain the original voiceprint data corresponding to the target device, and extract acoustic features from the original voiceprint data to obtain the voiceprint features corresponding to the target device;
[0007] Clustering is performed on the voiceprint features corresponding to the target device to obtain at least one set of voiceprint features and the target voiceprint features corresponding to each set of voiceprint features;
[0008] Using a pre-trained autoencoder model, a reconstruction error is obtained based on each set of voiceprint features and the corresponding target voiceprint features. When the reconstruction error is greater than or equal to a preset error threshold, a device anomaly alarm is generated.
[0009] According to another aspect of the present invention, a voiceprint recognition device is provided, comprising:
[0010] The voiceprint feature acquisition module is used to acquire the original voiceprint data corresponding to the target device, and to extract acoustic features from the original voiceprint data to acquire the voiceprint features corresponding to the target device.
[0011] The voiceprint feature clustering module is used to perform clustering processing on the voiceprint features corresponding to the target device to obtain at least one voiceprint feature set and the target voiceprint features corresponding to each voiceprint feature set.
[0012] The reconstruction error acquisition module is used to acquire the reconstruction error based on each set of voiceprint features and the corresponding target voiceprint features through a pre-trained autoencoder model, and generate a device abnormality alarm when the reconstruction error is greater than or equal to a preset error threshold.
[0013] According to another aspect of the present invention, an electronic device is provided, the electronic device comprising:
[0014] At least one processor; and
[0015] A memory communicatively connected to the at least one processor; wherein,
[0016] The memory stores a computer program that can be executed by the at least one processor, which enables the at least one processor to perform the voiceprint recognition method according to any embodiment of the present invention.
[0017] According to another aspect of the present invention, a computer-readable storage medium is provided, the computer-readable storage medium storing a computer program configured to cause a processor to execute and implement the voiceprint recognition method according to any embodiment of the present invention.
[0018] According to another aspect of the present invention, a computer program product is provided, comprising a computer program that, when executed by a processor, implements the voiceprint recognition method according to any embodiment of the present invention.
[0019] The technical solution of this invention involves acquiring the original voiceprint data corresponding to the target device, extracting acoustic features from the original voiceprint data to obtain the voiceprint features corresponding to the target device, performing clustering processing on the voiceprint features corresponding to the target device to obtain at least one set of voiceprint features and target voiceprint features corresponding to each set of voiceprint features, obtaining the reconstruction error based on each set of voiceprint features and the corresponding target voiceprint features using a pre-trained autoencoder model, and generating a device anomaly alarm when the reconstruction error is greater than or equal to a preset error threshold. By performing clustering processing on the extracted voiceprint features and using the autoencoder model to determine whether the device is abnormal based on the clustering results, the accuracy and robustness of voiceprint recognition can be improved.
[0020] It should be understood that the description in this section is not intended to identify key or essential features of the embodiments of the present invention, nor is it intended to limit the scope of the invention. Other features of the invention will become readily apparent from the following description. Attached Figure Description
[0021] To more clearly illustrate the technical solutions in the embodiments of the present invention, the accompanying drawings used in the description of the embodiments will be briefly introduced below. Obviously, the accompanying drawings described below are only some embodiments of the present invention. For those skilled in the art, other drawings can be obtained based on these drawings without creative effort.
[0022] Figure 1 This is a flowchart of a voiceprint recognition method provided in Embodiment 1 of the present invention;
[0023] Figure 2 This is a schematic diagram of the training and testing process of an autoencoder model according to Embodiment 1 of the present invention;
[0024] Figure 3 This is a flowchart of a voiceprint recognition method provided in Embodiment 2 of the present invention;
[0025] Figure 4 This is a schematic diagram of a process for obtaining a log-Mel spectrum according to Embodiment 2 of the present invention;
[0026] Figure 5 This is a schematic diagram of the structure of a voiceprint recognition device according to Embodiment 3 of the present invention;
[0027] Figure 6 This is a schematic diagram of the structure of an electronic device that implements the voiceprint recognition method of this invention. Detailed Implementation
[0028] To enable those skilled in the art to better understand the present invention, the technical solutions of the present invention will be clearly and completely described below with reference to the accompanying drawings of the embodiments of the present invention. Obviously, the described embodiments are only some embodiments of the present invention, and not all embodiments. Based on the embodiments of the present invention, all other embodiments obtained by those skilled in the art without creative effort should fall within the scope of protection of the present invention.
[0029] It should be noted that the terms "first," "second," "target," etc., used in the specification, claims, and accompanying drawings of this invention are used to distinguish similar objects and are not necessarily used to describe a specific order or sequence. It should be understood that such terms can be used interchangeably where appropriate so that embodiments of the invention described herein can be implemented in orders other than those illustrated or described herein. Furthermore, the terms "comprising" and "having," and any variations thereof, are intended to cover non-exclusive inclusion; for example, a process, method, system, product, or apparatus that comprises a series of steps or units is not necessarily limited to those steps or units explicitly listed, but may include other steps or units not explicitly listed or inherent to such processes, methods, products, or apparatus.
[0030] It is worth noting that the information collected by this invention is information and data authorized by the user or fully authorized by all parties, and the collection, storage, use, processing, transmission, provision, disclosure and application of related data all comply with the relevant laws, regulations and standards of relevant countries and regions, take necessary confidentiality measures, do not violate public order and good morals, and provide corresponding operation entry points for users to choose to authorize or refuse.
[0031] Example 1
[0032] Figure 1 This is a flowchart of a voiceprint recognition method provided in Embodiment 1 of the present invention. This embodiment is applicable to situations involving voiceprint recognition and status monitoring of cold source station equipment. The method can be executed by a voiceprint recognition device, which can be implemented in hardware and / or software. Typically, the voiceprint recognition device can be configured in an electronic device, such as a computer or server. Figure 1 As shown, the method includes:
[0033] S110. Obtain the original voiceprint data corresponding to the target device, and extract acoustic features from the original voiceprint data to obtain the voiceprint features corresponding to the target device.
[0034] The target device can be a cold source station device, such as a refrigeration unit or water pump. In this embodiment, audio data of the target device can be collected using pre-deployed audio acquisition equipment to serve as raw voiceprint data. The raw voiceprint data can then be preprocessed, such as through sampling and quantization, noise reduction, and filtering. Acoustic features that effectively characterize the device's operating status, such as time-domain and frequency-domain features, can then be extracted from the preprocessed raw voiceprint data to serve as voiceprint features.
[0035] Optionally, obtaining the raw voiceprint data corresponding to the target device may include:
[0036] The sound signals of the target device during operation are collected by a microphone array pre-deployed around the target device, which are used as the original voiceprint data corresponding to the target device.
[0037] In one optional example, a number of microphones can be deployed at equal intervals around the target device to form a microphone array. While the target device is running, the pre-deployed microphone array can be used to collect the sound signals generated by the target device in real time, forming raw voiceprint data.
[0038] In this embodiment, by collecting raw voiceprint data through a microphone array pre-deployed around the target device, effective and comprehensive collection of device voiceprint data can be achieved.
[0039] S120. Cluster the voiceprint features corresponding to the target device to obtain at least one set of voiceprint features and the target voiceprint features corresponding to each set of voiceprint features.
[0040] In this embodiment, a preset clustering algorithm can be used to cluster the voiceprint features, grouping similar voiceprint features into one category to obtain multiple voiceprint feature sets. Simultaneously, commonalities among the voiceprint features in each set can be extracted as target voiceprint features. These target voiceprint features provide important prior knowledge for the autoencoder model, helping it to better construct the encoder and decoder structures. The autoencoder model can then optimize the dimensionality reduction and reconstruction of the voiceprint features around the target voiceprint features.
[0041] Optionally, clustering the voiceprint features corresponding to the target device to obtain at least one set of voiceprint features and target voiceprint features corresponding to each set of voiceprint features may include:
[0042] The voiceprint features corresponding to the target device are clustered using the K-means clustering algorithm to obtain at least one set of voiceprint features.
[0043] Obtain the cluster center corresponding to each of the aforementioned voiceprint feature sets, and obtain the target voiceprint feature corresponding to each of the aforementioned voiceprint feature sets based on the voiceprint features corresponding to the cluster centers.
[0044] In one optional example, the K-means clustering algorithm can be used to cluster the extracted voiceprint features to initially distinguish different operating modes of the device, thereby obtaining multiple voiceprint feature sets. Then, based on the clustering results, the cluster center corresponding to each voiceprint feature set can be determined, and the voiceprint features corresponding to that cluster center, such as typical values of voiceprint frequency and amplitude, can be identified as the corresponding target voiceprint features.
[0045] In this embodiment, by using the K-means clustering algorithm to cluster voiceprint features and selecting the cluster center features as target voiceprint features, it is possible to distinguish voiceprint features under different working conditions, thereby improving the efficiency and accuracy of acquiring target voiceprint features.
[0046] S130. Using a pre-trained autoencoder model, based on each set of voiceprint features and the corresponding target voiceprint features, a reconstruction error is obtained, and when the reconstruction error is greater than or equal to a preset error threshold, a device abnormality alarm is generated.
[0047] An autoencoder model can include an encoder and a decoder. Specifically, the autoencoder model first maps the input data to a low-dimensional vector space using the encoder to extract key features, and then reconstructs the data based on the low-dimensional representation using the decoder to obtain the output data. The goal of the autoencoder model is to make the output as similar to the input as possible.
[0048] In this embodiment, the autoencoder model can learn from the divided voiceprint feature sets separately to avoid mutual interference between data from different device modes, thereby improving the accuracy of voiceprint recognition. Specifically, each voiceprint feature in the voiceprint feature set, along with the corresponding target voiceprint feature, can be used as input data and fed into the autoencoder model to obtain its output data. Then, based on loss functions such as mean squared error, the current loss function value can be calculated as the reconstruction error according to the input and output data. Finally, the reconstruction error can be compared with a preset error threshold. If the reconstruction error is greater than or equal to the preset error threshold, the target device is considered to be operating abnormally, and a corresponding device abnormality alarm is generated. If the reconstruction error is less than the preset error threshold, the target device is considered to be operating normally, and monitoring continues. This embodiment does not specifically limit the form of the device abnormality alarm.
[0049] Optionally, before obtaining the reconstruction error based on each set of voiceprint features and the corresponding target voiceprint features using a pre-trained autoencoder model, the method may further include: collecting voiceprint data during normal operation of the target device and preprocessing the voiceprint data to obtain preprocessed voiceprint data; then, extracting acoustic features from the preprocessed voiceprint data to obtain voiceprint features; further, inputting the voiceprint features into the encoder to obtain low-dimensional encoding, and inputting the low-dimensional encoding into the decoder to reconstruct the output features; then, calculating the error between the output features and the voiceprint features, and performing backpropagation and optimization based on the error, using an optimizer to adjust the network weights to minimize the loss function; finally, repeating the above steps until convergence, for example, when the loss function value is less than a preset error threshold.
[0050] In a specific example, the training and testing process of an autoencoder model can be as follows: Figure 2As shown. The encoder's input is the time-frequency domain features extracted from the audio data. During testing, if the reconstruction error is greater than or equal to a preset error threshold, it is determined to be an abnormal state; if the reconstruction error is less than the preset error threshold, it is determined to be a normal state.
[0051] Optionally, in this embodiment, the encoding dimension of the autoencoder model can be determined based on the number of voiceprint feature sets during normal device operation. The number of voiceprint feature sets represents the number of device modes. In this embodiment, the encoding dimension of the autoencoder model can be set to a suitable value that can effectively represent all device modes, thereby improving the model's learning efficiency and accuracy.
[0052] In this embodiment, by acquiring the target voiceprint features as prior knowledge and reference for the autoencoder model, the operating efficiency of the autoencoder model can be improved, thereby improving the real-time performance of voiceprint recognition.
[0053] The technical solution of this invention involves acquiring the original voiceprint data corresponding to the target device, extracting acoustic features from the original voiceprint data to obtain the voiceprint features corresponding to the target device, performing clustering processing on the voiceprint features corresponding to the target device to obtain at least one set of voiceprint features and target voiceprint features corresponding to each set of voiceprint features, obtaining the reconstruction error based on each set of voiceprint features and the corresponding target voiceprint features using a pre-trained autoencoder model, and generating a device anomaly alarm when the reconstruction error is greater than or equal to a preset error threshold. By performing clustering processing on the extracted voiceprint features and using the autoencoder model to determine whether the device is abnormal based on the clustering results, the accuracy and robustness of voiceprint recognition can be improved.
[0054] Example 2
[0055] Figure 3 This is a flowchart of a voiceprint recognition method provided in Embodiment 2 of the present invention. This embodiment is a further refinement of the above technical solution, and the technical solution in this embodiment can be combined with one or more of the above implementation methods. Figure 3 As shown, the method includes:
[0056] S210. By pre-deploying a microphone array around the target device, the sound signal of the target device during operation is collected as the original voiceprint data corresponding to the target device.
[0057] S220. Preprocess the original voiceprint data to obtain standard voiceprint data, and perform frequency domain conversion on the standard voiceprint data to obtain the frequency domain energy spectrum.
[0058] Preprocessing can include sampling, noise removal, filtering, and silence removal. Specifically, the preprocessed raw voiceprint data can be identified as standard voiceprint data, and Fourier transforms can be performed on the standard voiceprint data, such as Fast Fourier Transform or Short-Time Fourier Transform, to convert the time-domain signal into a frequency-domain energy spectrum.
[0059] Optionally, preprocessing the original voiceprint data to obtain standard voiceprint data may include: removing high-frequency noise from the original voiceprint data using a pre-emphasis filter to obtain the standard voiceprint data.
[0060] Among them, the pre-emphasis filter can be a first-order high-pass filter, which can increase the amplitude of the signal transition edge (corresponding to high frequency) to offset the high-frequency loss in subsequent transmission, thereby enhancing the high-frequency components and making the overall spectrum flatter.
[0061] In this embodiment, by using a pre-emphasis filter to process the original voiceprint data to obtain standard voiceprint data, the spectral energy can be balanced, thereby improving the accuracy of feature extraction.
[0062] Optionally, performing frequency domain transformation on the standard voiceprint data to obtain the frequency domain energy spectrum may include: performing frame segmentation and windowing processing on the standard voiceprint data to obtain multiple time-domain voiceprint signals, and performing fast Fourier transform on each of the time-domain voiceprint signals to obtain the frequency domain energy spectrum.
[0063] In an optional example, when performing frequency domain transformation on standard speaker data, the standard speaker data can first be divided into several short frames, each containing a certain number of sample points. Typically, a window size of 25 milliseconds per frame is used, with adjacent frames overlapping by 10 milliseconds. Then, a window function, such as a Hanning window, can be added to each frame to reduce spectral leakage, thereby obtaining the time-domain speaker signal. Finally, a Fast Fourier Transform can be performed on each time-domain speaker signal to convert the time-domain signal into a frequency-domain energy spectrum.
[0064] In this embodiment, the frequency domain energy spectrum is obtained by performing frame segmentation, windowing, and fast Fourier transform on the standard voiceprint signal, which can achieve efficient and accurate acquisition of the frequency domain energy spectrum.
[0065] S230. The frequency domain energy spectrum is filtered by a Mel filter to obtain an initial Mel spectrum, and the energy in the initial Mel spectrum is logarithmically processed to obtain a logarithmic Mel spectrum. Based on the logarithmic Mel spectrum, the voiceprint features corresponding to the target device are obtained.
[0066] It should be noted that sound signals are one-dimensional time-domain signals, lacking a clear pattern of frequency variation. While Fourier transform can transform sound signals into the frequency domain, allowing analysis of the signal's frequency distribution, it loses time-domain information and cannot analyze how the frequency distribution changes over time. To address this issue, this embodiment introduces Mel spectrograms, which combine the characteristics of both the time and frequency domains.
[0067] The Mel filter can be a set of triangular Mel filters, for example, 64-128 in number. This set of filters is denser in the low-frequency region and sparser in the high-frequency region to simulate the sensitivity of the human ear to different frequencies. The output energy of each filter represents the energy distribution at its corresponding Mel frequency.
[0068] In this embodiment, after obtaining the initial Mel spectrum through a Mel filter, the energy can be logarithmically calculated to simulate the nonlinear perception of intensity by the human ear, thereby obtaining the logarithmic Mel spectrum. Furthermore, time-frequency domain features can be extracted from the logarithmic Mel spectrum to obtain time-domain and frequency-domain features as the final voiceprint features.
[0069] Optionally, obtaining the voiceprint features corresponding to the target device based on the log-Mel spectrum may include:
[0070] Based on the log-Mel spectrum, the mapping relationship between time point, frequency value and energy value is obtained, and based on the mapping relationship between time point, frequency value and energy value, the voiceprint feature corresponding to the target device is obtained.
[0071] In the log-Mel spectrum, the horizontal axis represents time, the vertical axis represents Mel frequency, and the color depth represents log energy intensity. Therefore, the mapping relationship between time points, frequency values, and energy values can be extracted from the log-Mel spectrum. Subsequently, the extracted time points can be used as time-domain features, the frequency values as frequency-domain features, and the energy values as energy features. These time-domain features, frequency-domain features, and energy features are then combined to form the voiceprint feature.
[0072] In this embodiment, by extracting time-frequency domain features from the log-Mel spectrum, efficient acquisition of time-domain and frequency-domain features can be achieved, providing a data foundation for abnormal voiceprint recognition.
[0073] In one specific implementation of this embodiment, the process for obtaining the log-Melbourne spectrum can be as follows: Figure 4 As shown, firstly, the audio data is divided into frames to obtain several short frames. Then, the frames are processed frame by frame through windowing and fast Fourier transform to obtain the frequency domain energy spectrum. Next, the frequency domain energy spectrum is filtered through a set of Mel filters to obtain the initial Mel spectrum. Finally, the logarithm of the output energy of each Mel filter is taken to convert the initial Mel spectrum into a logarithmic Mel spectrum.
[0074] S240. Using the K-means clustering algorithm, cluster the voiceprint features corresponding to the target device to obtain at least one set of voiceprint features.
[0075] S250. Obtain the cluster center corresponding to each of the voiceprint feature sets, and obtain the target voiceprint feature corresponding to each of the voiceprint feature sets based on the voiceprint features corresponding to the cluster center.
[0076] S260. Using a pre-trained autoencoder model, based on each set of voiceprint features and the corresponding target voiceprint features, a reconstruction error is obtained, and when the reconstruction error is greater than or equal to a preset error threshold, a device abnormality alarm is generated.
[0077] In this embodiment, the combination of cluster analysis and autoencoder technology improves the robustness of the voiceprint recognition system to environmental noise and individual differences. Secondly, the use of an autoencoder model for anomaly detection enables rapid response to changes in the operating status of the cold source station equipment, improving the real-time performance of equipment anomaly detection. Finally, it eliminates the need for additional expensive sensors or complex manual inspections, reducing maintenance costs and workload.
[0078] The technical solution of this invention, after acquiring the original voiceprint data, preprocesses the original voiceprint data to obtain standard voiceprint data, and performs frequency domain conversion on the standard voiceprint data to obtain a frequency domain energy spectrum; filters the frequency domain energy spectrum using a Mel filter to obtain an initial Mel spectrum, and takes the logarithm of the energy in the initial Mel spectrum to obtain a logarithmic Mel spectrum; and obtains the voiceprint features corresponding to the target device based on the logarithmic Mel spectrum; by converting the frequency domain energy spectrum into a logarithmic Mel spectrum and extracting the voiceprint features from the logarithmic Mel spectrum, efficient and accurate acquisition of time-frequency domain features can be achieved, thereby improving the accuracy of voiceprint recognition.
[0079] Example 3
[0080] Figure 5 This is a schematic diagram of the structure of a voiceprint recognition device provided in Embodiment 3 of the present invention. Figure 5 As shown, the device includes: a voiceprint feature acquisition module 310, a voiceprint feature clustering module 320, and a reconstruction error acquisition module 330; wherein,
[0081] The voiceprint feature acquisition module 310 is used to acquire the original voiceprint data corresponding to the target device, and to extract acoustic features from the original voiceprint data to acquire the voiceprint features corresponding to the target device.
[0082] The voiceprint feature clustering module 320 is used to perform clustering processing on the voiceprint features corresponding to the target device to obtain at least one voiceprint feature set and the target voiceprint features corresponding to each voiceprint feature set.
[0083] The reconstruction error acquisition module 330 is used to acquire the reconstruction error based on each set of voiceprint features and the corresponding target voiceprint features through a pre-trained autoencoder model, and generate a device abnormality alarm when the reconstruction error is greater than or equal to a preset error threshold.
[0084] The technical solution of this invention involves acquiring the original voiceprint data corresponding to the target device, extracting acoustic features from the original voiceprint data to obtain the voiceprint features corresponding to the target device, performing clustering processing on the voiceprint features corresponding to the target device to obtain at least one set of voiceprint features and target voiceprint features corresponding to each set of voiceprint features, obtaining the reconstruction error based on each set of voiceprint features and the corresponding target voiceprint features using a pre-trained autoencoder model, and generating a device anomaly alarm when the reconstruction error is greater than or equal to a preset error threshold. By performing clustering processing on the extracted voiceprint features and using the autoencoder model to determine whether the device is abnormal based on the clustering results, the accuracy and robustness of voiceprint recognition can be improved.
[0085] Optionally, the voiceprint feature acquisition module 310 includes:
[0086] The data preprocessing unit is used to preprocess the original voiceprint data to obtain standard voiceprint data, and to perform frequency domain transformation on the standard voiceprint data to obtain the frequency domain energy spectrum.
[0087] The Mel spectrum acquisition unit is used to filter the frequency domain energy spectrum through a Mel filter to obtain an initial Mel spectrum, and to take the logarithm of the energy in the initial Mel spectrum to obtain a logarithmic Mel spectrum, and to obtain the voiceprint features corresponding to the target device based on the logarithmic Mel spectrum.
[0088] Optionally, the data preprocessing unit is specifically used to remove high-frequency noise from the original voiceprint data through a pre-emphasis filter to obtain the standard voiceprint data.
[0089] Optionally, the data preprocessing unit is further configured to perform framing and windowing processing on the standard voiceprint data to obtain multiple time-domain voiceprint signals, and to perform fast Fourier transform on each of the time-domain voiceprint signals to obtain the frequency-domain energy spectrum.
[0090] Optionally, the Mel spectrum acquisition unit is specifically used to acquire the mapping relationship between time point, frequency value and energy value according to the logarithmic Mel spectrum, and to acquire the voiceprint feature corresponding to the target device according to the mapping relationship between the time point, frequency value and energy value.
[0091] Optionally, the voiceprint feature clustering module 320 is specifically used to perform clustering processing on the voiceprint features corresponding to the target device using the K-means clustering algorithm to obtain at least one set of voiceprint features.
[0092] Obtain the cluster center corresponding to each of the aforementioned voiceprint feature sets, and obtain the target voiceprint feature corresponding to each of the aforementioned voiceprint feature sets based on the voiceprint features corresponding to the cluster centers.
[0093] Optionally, the voiceprint feature acquisition module 310 is specifically used to acquire the sound signal of the target device during operation by using a microphone array pre-deployed around the target device, so as to use the original voiceprint data corresponding to the target device.
[0094] The voiceprint recognition device provided in the embodiments of the present invention can execute the voiceprint recognition method provided in any embodiment of the present invention, and has the corresponding functional modules and beneficial effects of the method.
[0095] The collection, storage, use, processing, transmission, provision, and disclosure of user personal information involved in the technical solution disclosed herein comply with the provisions of relevant laws and regulations and do not violate public order and good morals.
[0096] Example 4
[0097] Figure 6 A schematic diagram of an electronic device 40 that can be used to implement embodiments of the present invention is shown. The electronic device 40 is intended to represent various forms of digital computers, such as laptop computers, desktop computers, workstations, personal digital assistants, servers, blade servers, mainframe computers, and other suitable computers. The electronic device 40 can also represent various forms of mobile devices, such as personal digital processors, cellular phones, smartphones, wearable devices (such as helmets, glasses, watches, etc.), and other similar computing devices. The components shown herein, their connections and relationships, and their functions are merely illustrative and are not intended to limit the implementation of the invention described and / or claimed herein.
[0098] like Figure 6As shown, the electronic device 40 includes at least one processor 41 and a memory, such as a read-only memory (ROM) 42 or a random access memory (RAM) 43, communicatively connected to the at least one processor 41. The memory stores computer programs executable by the at least one processor. The processor 41 can perform various appropriate actions and processes based on the computer program stored in the read-only memory 42 or loaded from the storage unit 48 into the random access memory 43. The RAM 43 can also store various programs and data required for the operation of the electronic device 40. The processor 41, ROM 42, and RAM 43 are interconnected via a bus 44. An input / output (I / O) interface 45 is also connected to the bus 44.
[0099] Multiple components in electronic device 40 are connected to I / O interface 45, including: input unit 46, such as keyboard, mouse, etc.; output unit 47, such as various types of monitors, speakers, etc.; storage unit 48, such as disk, optical disk, etc.; and communication unit 49, such as network card, modem, wireless transceiver, etc. Communication unit 49 allows electronic device 40 to exchange information / data with other devices through computer networks such as the Internet and / or various telecommunications networks.
[0100] Processor 41 can be a variety of general-purpose and / or special-purpose processing components with processing and computing capabilities. Some examples of processor 41 include, but are not limited to, central processing units, graphics processing units, various special-purpose artificial intelligence computing chips, various processors running machine learning model algorithms, digital signal processors, and any suitable processor, controller, microcontroller, etc. Processor 41 performs the various methods and processes described above, such as voiceprint recognition methods.
[0101] In some embodiments, the voiceprint recognition method may be implemented as a computer program tangibly contained in a computer-readable storage medium, such as storage unit 48. In some embodiments, part or all of the computer program may be loaded and / or installed on electronic device 40 via ROM 42 and / or communication unit 49. When the computer program is loaded into RAM 43 and executed by processor 41, one or more steps of the voiceprint recognition method described above may be performed. Alternatively, in other embodiments, processor 41 may be configured to perform the voiceprint recognition method by any other suitable means (e.g., by means of firmware).
[0102] Various embodiments of the systems and techniques described above herein can be implemented in digital electronic circuit systems, integrated circuit systems, field-programmable gate arrays, application-specific integrated circuits (ASICs), application-specific standard products (ASICs), system-on-a-chip (SoCs), complex programmable logic devices, computer hardware, firmware, software, and / or combinations thereof. These various embodiments may include implementations in one or more computer programs that can be executed and / or interpreted on a programmable system including at least one programmable processor, which may be a dedicated or general-purpose programmable processor, capable of receiving data and instructions from a storage system, at least one input device, and at least one output device, and transmitting data and instructions to the storage system, the at least one input device, and the at least one output device.
[0103] Computer programs used to implement the methods of the present invention may be written in any combination of one or more programming languages. These computer programs may be provided to a processor of a general-purpose computer, a special-purpose computer, or other programmable data processing device, such that when executed by the processor, the computer programs cause the functions / operations specified in the flowcharts and / or block diagrams to be performed. The computer programs may be executed entirely on a machine, partially on a machine, or as a standalone software package, partially on a machine and partially on a remote machine, or entirely on a remote machine or server.
[0104] In the context of this invention, a computer-readable storage medium can be a tangible medium that may contain or store a computer program for use by or in conjunction with an instruction execution system, apparatus, or device. A computer-readable storage medium may include, but is not limited to, electronic, magnetic, optical, electromagnetic, infrared, or semiconductor systems, apparatus, or devices, or any suitable combination thereof. Alternatively, a computer-readable storage medium may be a machine-readable signal medium. More specific examples of machine-readable storage media include electrical connections based on one or more wires, portable computer disks, hard disks, random access memory, read-only memory, erasable programmable read-only memory, optical fibers, portable compact disk read-only memory, optical storage devices, magnetic storage devices, or any suitable combination thereof.
[0105] To provide interaction with a user, the systems and techniques described herein can be implemented on an electronic device 40, which includes: a display device (e.g., a cathode ray tube or liquid crystal display) for displaying information to the user; and a keyboard and pointing device (e.g., a mouse or trackball) through which the user provides input to the electronic device 40. Other types of devices can also be used to provide interaction with the user; for example, feedback provided to the user can be any form of sensory feedback (e.g., visual feedback, auditory feedback, or tactile feedback); and input from the user can be received in any form (including sound input, voice input, or tactile input).
[0106] The systems and technologies described herein can be implemented in computing systems that include backend components (e.g., as data servers), or computing systems that include middleware components (e.g., application servers), or computing systems that include frontend components (e.g., user computers with graphical user interfaces or web browsers through which users can interact with implementations of the systems and technologies described herein), or any combination of such backend, middleware, or frontend components. The components of the system can be interconnected via digital data communication of any form or medium (e.g., communication networks). Examples of communication networks include local area networks (LANs), wide area networks (WANs), blockchain networks, and the Internet.
[0107] A computing system can include clients and servers. Clients and servers are generally located far apart and typically interact via a communication network. The client-server relationship is created by computer programs running on the respective computers and having a client-server relationship with each other. The server can be a cloud server.
[0108] This embodiment may also include a computer program product, which includes a computer program that, when executed by a processor, implements the voiceprint recognition method provided in any embodiment of the present invention.
[0109] It should be understood that the various forms of processes shown above can be used, with steps reordered, added, or deleted. For example, the steps described in this invention can be executed in parallel, sequentially, or in different orders, as long as the desired result of the technical solution of this invention can be achieved, and this is not limited herein.
[0110] The specific embodiments described above do not constitute a limitation on the scope of protection of this invention. Those skilled in the art should understand that various modifications, combinations, sub-combinations, and substitutions can be made according to design requirements and other factors. Any modifications, equivalent substitutions, and improvements made within the spirit and principles of this invention should be included within the scope of protection of this invention.
Claims
1. A voiceprint recognition method, characterized in that, The method comprises the following steps: obtaining original voiceprint data corresponding to a target device, and performing acoustic feature extraction on the original voiceprint data to obtain voiceprint features corresponding to the target device; performing clustering processing on the voiceprint features corresponding to the target device to obtain at least one voiceprint feature set and target voiceprint features corresponding to each voiceprint feature set; obtaining reconstruction errors according to each voiceprint feature set and the corresponding target voiceprint features through a pre-trained autoencoder model, and generating a device anomaly alarm when the reconstruction error is greater than or equal to a preset error threshold.
2. The method of claim 1, wherein, The method for performing acoustic feature extraction on the original voiceprint data to obtain voiceprint features corresponding to the target device comprises the following steps: performing preprocessing on the original voiceprint data to obtain standard voiceprint data, and performing frequency domain conversion on the standard voiceprint data to obtain a frequency domain energy spectrum; performing filtering processing on the frequency domain energy spectrum through a Mel filter to obtain an initial Mel spectrum, performing logarithmic processing on the energy in the initial Mel spectrum to obtain a log Mel spectrum, and obtaining voiceprint features corresponding to the target device according to the log Mel spectrum.
3. The method of claim 2, wherein, The method for performing preprocessing on the original voiceprint data to obtain standard voiceprint data comprises the following steps: performing high-frequency noise removal on the original voiceprint data through a pre-emphasis filter to obtain the standard voiceprint data.
4. The method of claim 2, wherein, The method for performing frequency domain conversion on the standard voiceprint data to obtain a frequency domain energy spectrum comprises the following steps: performing framing and windowing processing on the standard voiceprint data to obtain a plurality of time domain voiceprint signals, and performing fast Fourier transform on each time domain voiceprint signal to obtain a frequency domain energy spectrum.
5. The method of claim 2, wherein, The method for obtaining voiceprint features corresponding to the target device according to the log Mel spectrum comprises the following steps: obtaining a mapping relationship between time points, frequency values and energy values according to the log Mel spectrum, and obtaining voiceprint features corresponding to the target device according to the mapping relationship between the time points, the frequency values and the energy values.
6. The method of claim 1, wherein, The method for performing clustering processing on the voiceprint features corresponding to the target device to obtain at least one voiceprint feature set and target voiceprint features corresponding to each voiceprint feature set comprises the following steps: performing clustering processing on the voiceprint features corresponding to the target device through a K-means clustering algorithm to obtain at least one voiceprint feature set; obtaining a clustering center corresponding to each voiceprint feature set, and obtaining target voiceprint features corresponding to each voiceprint feature set according to the voiceprint features corresponding to the clustering center.
7. The method of claim 1, wherein, The method for obtaining original voiceprint data corresponding to a target device comprises the following steps: collecting sound signals generated by the target device through a microphone array pre-deployed around the target device as the original voiceprint data corresponding to the target device.
8. An electronic device, comprising: The electronic device comprises: at least one processor, and a memory connected in communication with the at least one processor; wherein the memory stores a computer program that can be executed by the at least one processor, and the computer program is executed by the at least one processor to enable the at least one processor to execute the voiceprint recognition method in any one of claims 1-7.
9. A computer-readable storage medium, characterized in that, The computer readable storage medium stores a computer program, which is used to make the processor execute the voiceprint recognition method in any one of claims 1-7.
10. A computer program product, characterised in that, The computer readable storage medium stores a computer program, which is used to make the processor execute the voiceprint recognition method in any one of claims 1-7.