Voice matching method and system for router, gateway and camera

By creating and fine-tuning the voice matching model on edge computing devices, combining knowledge distillation and dynamic pruning technology, the resource requirements and accuracy problems of edge computing devices in voice matching are solved, and efficient and accurate voice matching is achieved, improving user experience.

CN120148487APending Publication Date: 2025-06-13FUJIAN NEWLAND COMM SCI TECH
View PDF 0 Cites 1 Cited by

Patent Information

Application Number
CN202510269128.6
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-03-07
Publication Date
2025-06-13

AI Technical Summary

Technical Problem

Edge computing devices such as routers, gateways, and cameras face technical bottlenecks such as imbalance in computing power and accuracy, defects in noise robustness and inefficient dynamic matching, affecting the timeliness and accuracy of voice matching.

Method used

By creating a speech matching model, including speech data input, transformation, feature extraction, similarity calculation and output modules, combined with knowledge distillation and dynamic pruning techniques, model compression is performed, and fine-tuned on edge computing devices, to reduce resource requirements and improve matching accuracy and timeliness.

Benefits of technology

It effectively reduces the resource requirements for voice matching, improves the accuracy and timeliness of voice matching, improves user experience, and supports real-time matching of voice data up to 30 seconds.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120148487A_ABST
    Figure CN120148487A_ABST
Patent Text Reader

Abstract

The invention provides a voice matching method and system for a router, a gateway and a camera in the technical field of edge computing, and the method comprises the steps: S1, creating a voice matching model, and setting a loss function of the voice matching model; s2, acquiring a large amount of historical voice data to construct a data set; s3, training the voice matching model based on the data set and the loss function, and compressing the voice matching model through a knowledge distillation technology and a dynamic pruning technology in the training process; s4, deploying the trained voice matching model on an edge computing device, and performing fine adjustment on the deployed voice matching model by the edge computing device; and S5, acquiring real-time voice data by the edge computing device, preprocessing the real-time voice data, and inputting the preprocessed real-time voice data into the fine-tuned voice matching model to obtain a voice matching result. The method has the advantages that the resource demand of voice matching is greatly reduced, and the accuracy and timeliness of voice matching are greatly improved.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention relates to the technical field of edge computing, and particularly to a voice matching method and system for routers, gateways, and cameras. Background Art

[0002] With the rise of edge computing devices (such as routers, gateways, and cameras), many innovative application scenarios have emerged. Among them, voice matching is an important application scenario, and functions such as voice command recognition and voice search can be realized through voice matching.

[0003] For voice matching, traditionally, it is necessary to obtain the voice data input by the user and compare it with the known voice data on the cloud server. However, traditional methods usually require a large amount of computing resources and network bandwidth, which are often limited on edge computing devices, thereby affecting the timeliness and accuracy of voice matching. In order to ensure the timeliness of voice matching and thus ensure the user experience, there is a need to perform voice matching on edge computing devices.

[0004] However, performing voice matching on edge computing devices faces the following technical bottlenecks: 1. The contradiction between computing power and accuracy: Traditional voice matching algorithms (such as HMM-GMM) are difficult to meet the real-time requirements on low-computing-power edge computing devices, while matching methods based on deep learning (such as end-to-end ASR models) often require more than 500MB of memory occupancy and are difficult to be deployed on resource-constrained edge computing devices. 2. Defects in noise robustness: In voice front-end processing, spectral subtraction with a fixed threshold is mostly used. In sudden noise scenarios (such as the noise of a range hood when a home camera is installed in the kitchen), the signal-to-noise ratio deteriorates severely, resulting in inaccurate acoustic features. 3. Low efficiency of dynamic matching: The time complexity of the traditional DTW algorithm is high. When processing a voice sequence with a length exceeding 5 seconds (such as the continuous conversation scenario of an intelligent gateway), the single matching time exceeds 800ms.

[0005] Therefore, how to provide a voice matching method and system for routers, gateways, and cameras to reduce the resource requirements of voice matching, improve the accuracy and timeliness of voice matching, and enhance the user experience has become an urgent technical problem to be solved. Summary of the Invention

[0006] The technical problem to be solved by the present invention is to provide a voice matching method and system for routers, gateways, and cameras to reduce the resource requirements of voice matching, improve the accuracy and timeliness of voice matching, and enhance the user experience.

[0007] In a first aspect, the present invention provides a voice matching method for routers, gateways, and cameras, including the following steps:

[0008] Step S1: Create a voice matching model based on a voice data input module, a voice data conversion module, an acoustic feature extraction module, a similarity calculation module, and an output module, and set the loss function of the voice matching model;

[0009] Step S2: Obtain a large amount of historical voice data, preprocess and label each piece of historical voice data, and then construct a data set;

[0010] Step S3: Train the voice matching model based on the data set and the loss function, and compress the voice matching model through knowledge distillation technology and dynamic pruning technology during the training process;

[0011] Step S4: Deploy the trained voice matching model on an edge computing device of which the device type is a router, a gateway or a camera, and the edge computing device fine-tunes the deployed voice matching model;

[0012] Step S5: The edge computing device obtains real-time voice data, preprocesses the real-time voice data, and then inputs it into the fine-tuned voice matching model to obtain a voice matching result, so as to complete voice matching;

[0013] Step S6: The edge computing device continuously optimizes the voice matching model based on the voice matching result and the real-time voice data.

[0014] Further, the specific content of Step S1 is as follows:

[0015] Create a voice matching model based on a voice data input module, a voice data conversion module, an acoustic feature extraction module, a similarity calculation module, and an output module, and set the loss function of the voice matching model; the voice data input module, the voice data conversion module, the acoustic feature extraction module, the similarity calculation module, and the output module are connected in sequence; the loss function is constructed based on a contrastive loss function and a binary cross-entropy loss function;

[0016] The speech data input module is used to perform first-stage noise reduction on the input speech data through spectral subtraction, and then perform second-stage noise reduction through a three-layer one-dimensional convolutional network to obtain noise-reduced data; the speech data conversion module is used to convert the noise-reduced data into a feature representation through 12 low-frequency Mel filters and 16 medium-high-frequency Mel filters; the frame length of the low-frequency Mel filter and the medium-high-frequency Mel filter is 25 ms, and the frame shift is 10 ms; the acoustic feature extraction module is used to extract acoustic features from the feature representation, and is constructed based on a time-domain feature extraction unit and a frequency-domain feature extraction unit; the time-domain feature extraction unit is constructed based on a CNN network, and the frequency-domain feature extraction unit is constructed based on a BiLSTM network; the similarity calculation module is used to segment the speech data into several segments of speech sub-data according to phoneme boundaries, and calculate the similarity of the acoustic features corresponding to the speech sub-data according to the DTW algorithm improved by slope constraint conditions; the output module is used to output a speech matching result according to the similarity.

[0017] Further, the specific steps of step S2 are as follows:

[0018] Obtain a large amount of historical speech data, perform preprocessing on each of the historical speech data including at least format conversion, noise reduction, sampling rate unification, and removal of silent segments, and construct a data set after annotating similar speech data for each of the preprocessed historical speech data.

[0019] Further, the specific steps of step S3 are as follows:

[0020] Divide the data set into a training set, a validation set, and a test set based on a preset segmentation ratio, train the speech matching model through the training set until the loss value of the loss function is less than a preset loss threshold, and compress the speech matching model through knowledge distillation technology and dynamic pruning technology during the training process; verify the trained speech matching model through the validation set, and determine whether the matching accuracy rate is greater than a preset accuracy threshold. If not, the verification fails, and the training set is expanded and training continues. If so, the verification is successful; test the speech matching model with successful verification through the test set, and determine whether the confidence level is greater than a preset confidence level threshold. If not, the test fails, and the training set is expanded and training continues. If so, the test is successful, and the training ends.

[0021] Further, the specific steps of step S4 are as follows:

[0022] Deploy the trained speech matching model on an edge computing device with a device type of router, gateway, or camera. The edge computing device obtains a preset number of real-time speech data to construct a fine-tuning data set, and based on a preset learning rate, batch size, and optimizer, perform adversarial training on the deployed speech matching model through the fine-tuning data set, and then fine-tune the deployed speech matching model.

[0023] In a second aspect, the present invention provides a voice matching system for routers, gateways, and cameras, including the following modules:

[0024] A voice matching model creation module, configured to create a voice matching model based on a voice data input module, a voice data conversion module, an acoustic feature extraction module, a similarity calculation module, and an output module, and set a loss function of the voice matching model;

[0025] A dataset construction module, configured to obtain a large amount of historical voice data, preprocess and annotate each piece of the historical voice data, and then construct a dataset;

[0026] A voice matching model training module, configured to train the voice matching model based on the dataset and the loss function, and compress the voice matching model through knowledge distillation technology and dynamic pruning technology during the training process;

[0027] A voice matching model deployment module, configured to deploy the trained voice matching model on an edge computing device of a device type of a router, a gateway, or a camera, and the edge computing device fine-tunes the deployed voice matching model;

[0028] A voice matching module, configured to enable the edge computing device to obtain real-time voice data, preprocess the real-time voice data, and then input the preprocessed real-time voice data into the fine-tuned voice matching model to obtain a voice matching result, so as to complete voice matching;

[0029] A voice matching model optimization module, configured to enable the edge computing device to continuously optimize the voice matching model based on the voice matching result and the real-time voice data.

[0030] Further, the voice matching model creation module is specifically configured to:

[0031] Create a voice matching model based on a voice data input module, a voice data conversion module, an acoustic feature extraction module, a similarity calculation module, and an output module, and set a loss function of the voice matching model; the voice data input module, the voice data conversion module, the acoustic feature extraction module, the similarity calculation module, and the output module are connected in sequence; the loss function is constructed based on a contrastive loss function and a binary cross-entropy loss function;

[0032] The voice data input module is used to perform first-stage noise reduction on the input voice data through spectral subtraction, and then perform second-stage noise reduction through a three-layer one-dimensional convolutional network to obtain noise-reduced data; the voice data conversion module is used to convert the noise-reduced data into a feature representation through 12 low-frequency Mel filters and 16 medium-high frequency Mel filters; the frame length of the low-frequency Mel filter and the medium-high frequency Mel filter is 25 ms, and the frame shift is 10 ms; the acoustic feature extraction module is used to extract acoustic features from the feature representation and is constructed based on a time-domain feature extraction unit and a frequency-domain feature extraction unit; the time-domain feature extraction unit is constructed based on a CNN network, and the frequency-domain feature extraction unit is constructed based on a BiLSTM network; the similarity calculation module is used to segment the voice data into several segments of voice sub-data according to phoneme boundaries, and calculate the similarity of the acoustic features corresponding to the voice sub-data according to the DTW algorithm improved by slope constraint conditions; the output module is used to output a voice matching result according to the similarity.

[0033] Further, the dataset construction module is specifically used for:

[0034] Obtain a large amount of historical voice data, perform preprocessing on each of the historical voice data including at least format conversion, noise reduction, sampling rate unification, and removal of silent segments, and construct a dataset after annotating similar voice data for each of the preprocessed historical voice data.

[0035] Further, the voice matching model training module is specifically used for:

[0036] Divide the dataset into a training set, a validation set, and a test set based on a preset segmentation ratio, train the voice matching model through the training set until the loss value of the loss function is less than a preset loss threshold, and compress the voice matching model through knowledge distillation technology and dynamic pruning technology during the training process; verify the trained voice matching model through the validation set, and judge whether the matching accuracy rate is greater than a preset accuracy threshold. If not, the verification fails, and the training set is expanded and training continues. If so, the verification is successful; test the voice matching model that has passed the verification through the test set, and judge whether the confidence level is greater than a preset confidence threshold. If not, the test fails, and the training set is expanded and training continues. If so, the test is successful, and training ends.

[0037] Further, the voice matching model deployment module is specifically used for:

[0038] Deploy the trained voice matching model on an edge computing device with a device type of router, gateway, or camera. The edge computing device obtains a preset number of real-time voice data to construct a fine-tuning data set, and based on a preset learning rate, batch size, and optimizer, performs adversarial training on the deployed voice matching model through the fine-tuning data set, and then fine-tunes the deployed voice matching model.

[0039] The advantages of the present invention are as follows:

[0040] 1. Create a voice matching model through a voice data input module, a voice data conversion module, an acoustic feature extraction module, a similarity calculation module, and an output module, and set the loss function of the voice matching model; then obtain a large amount of historical voice data for preprocessing and annotation to construct a data set, and train the voice matching model based on the data set and the loss function. During the training process, compress the voice matching model through knowledge distillation technology and dynamic pruning technology, and then deploy the trained voice matching model on an edge computing device with a device type of router, gateway, or camera. The edge computing device fine-tunes the deployed voice matching model; finally, the edge computing device obtains real-time voice data, preprocesses the real-time voice data and inputs it into the fine-tuned voice matching model to obtain a voice matching result, and continuously optimizes the voice matching model based on the voice matching result and the real-time voice data; that is, perform voice matching based on the voice matching model created by the voice data input module, the voice data conversion module, the acoustic feature extraction module, the similarity calculation module, and the output module. The Mel filter in the voice data conversion module is optimized from 40 conventional ones to 28 (12 low-frequency Mel filters and 16 medium-high-frequency Mel filters) to simplify the network structure; the similarity calculation module segments the voice data according to phoneme boundaries and calculates the similarity using the DTW algorithm improved based on slope constraint conditions to improve the calculation efficiency; compress the voice matching model by combining knowledge distillation technology and dynamic pruning technology, effectively reducing the volume of the voice matching model, facilitating the deployment of the voice matching model on resource-constrained edge computing devices; and the voice data input module performs two-stage noise reduction (spectral subtraction, one-dimensional convolutional network) on the input voice, effectively overcoming the influence of noise and improving robustness. Combining the fine-tuning and optimization of the voice matching model, ultimately greatly reduces the resource requirements for voice matching, greatly improves the accuracy and timeliness of voice matching, and thus greatly improves the user experience.

[0041] 2. The loss function is constructed based on the contrastive loss function and the binary cross-entropy loss function, that is, the contrastive loss function and the binary cross-entropy loss function are weighted; since the contrastive loss function is used to learn the feature embedding space, making the distance between similar samples closer and the distance between dissimilar samples farther, and the binary cross-entropy loss function is used to output binary classification problems, it can effectively combine the advantages of the contrastive loss function and the binary cross-entropy loss function, thus greatly improving the training effect of the speech matching model.

[0042] 3. The acoustic feature extraction module is constructed based on the time-domain feature extraction unit and the frequency-domain feature extraction unit. The time-domain feature extraction unit is based on the CNN network, and the frequency-domain feature extraction unit is based on the Bi LSTM network, so that the acoustic features extracted by the acoustic feature extraction module combine time-domain features and frequency-domain features, effectively improving the feature extraction ability, and thus greatly improving the accuracy of speech matching.

[0043] 4. By dividing the data set into a training set, a validation set, and a test set, the speech matching model is trained by the training set until the loss value of the loss function is less than the preset loss threshold. During the training process, the speech matching model is compressed by the knowledge distillation technology and the dynamic pruning technology; the matching accuracy is calculated by the validation set to verify the trained speech matching model, and the confidence is calculated by the test set to test the speech matching model that has passed the verification; that is, during the training process of the speech matching model, compression, verification, and testing are continuously performed to effectively balance the model volume and matching accuracy of the speech matching model, so as to be better deployed on edge computing devices. BRIEF DESCRIPTION OF THE DRAWINGS

[0044] The present invention will be further described below with reference to the accompanying drawings in conjunction with embodiments.

[0045] Figure 1 It is a flowchart of a speech matching method for routers, gateways, and cameras according to the present invention.

[0046] Figure 2 It is a schematic structural diagram of a speech matching system for routers, gateways, and cameras according to the present invention. DETAILED DESCRIPTION OF THE EMBODIMENTS

[0047] The overall idea of the technical solution in the embodiments of this application is as follows: Perform voice matching based on a voice matching model created by a voice data input module, a voice data conversion module, an acoustic feature extraction module, a similarity calculation module, and an output module. The Mel filter in the voice data conversion module is optimized from 40 conventional ones to 28 to simplify the network structure; the similarity calculation module segments the voice data according to phoneme boundaries and calculates the similarity using the DTW algorithm improved based on slope constraint conditions to improve the calculation efficiency; compress the voice matching model by combining knowledge distillation technology and dynamic pruning technology to effectively reduce the volume of the voice matching model; and the voice data input module performs two-stage noise reduction on the input voice to effectively overcome the influence of noise, combined with the fine-tuning and optimization of the voice matching model to reduce the resource requirements for voice matching and improve the accuracy and timeliness of voice matching.

[0048] Please refer to Figures 1 to 2 as shown, a preferred embodiment of a voice matching method for a router, a gateway, and a camera according to the present invention includes the following steps:

[0049] Step S1: Create a voice matching model based on a voice data input module, a voice data conversion module, an acoustic feature extraction module, a similarity calculation module, and an output module, and set the loss function of the voice matching model;

[0050] Step S2: Obtain a large amount of historical voice data, preprocess and annotate each piece of historical voice data, and then construct a data set;

[0051] Step S3: Train the voice matching model based on the data set and the loss function, and compress the voice matching model through knowledge distillation technology and dynamic pruning technology during the training process;

[0052] Step S4: Deploy the trained voice matching model on an edge computing device of a device type of a router, a gateway, or a camera, and the edge computing device fine-tunes the deployed voice matching model; through fine-tuning, the model performance can be quickly improved on specific tasks;

[0053] Step S5: The edge computing device obtains real-time voice data, preprocesses the real-time voice data, and then inputs it into the fine-tuned voice matching model to obtain a voice matching result, so as to complete voice matching;

[0054] Step S6: The edge computing device continuously optimizes the voice matching model based on the voice matching result and the real-time voice data.

[0055] The present invention can achieve an end-to-end delay of <150 ms on an ARM Cortex-A53 level processor, maintain an instruction recognition accuracy rate of ≥99% in an environment with a signal-to-noise ratio ≥5 dB, and support real-time matching of up to 30 seconds of voice data.

[0056] Specifically, step S1 is as follows:

[0057] Create a voice matching model based on a voice data input module, a voice data conversion module, an acoustic feature extraction module, a similarity calculation module, and an output module, and set the loss function of the voice matching model; the voice data input module, the voice data conversion module, the acoustic feature extraction module, the similarity calculation module, and the output module are connected in sequence; the loss function is constructed based on a contrastive loss function and a binary cross-entropy loss function;

[0058] By setting the loss function to be constructed based on a contrastive loss function and a binary cross-entropy loss function, that is, weighting the contrastive loss function and the binary cross-entropy loss function; since the contrastive loss function is used to learn the feature embedding space, making the distance between similar samples closer and the distance between dissimilar samples farther, and the binary cross-entropy loss function is used to output a binary classification problem, it can effectively combine the advantages of the contrastive loss function and the binary cross-entropy loss function, thereby greatly improving the training effect of the voice matching model.

[0059] The voice data input module is used to perform first-stage noise reduction on the input voice data through spectral subtraction, and then perform second-stage noise reduction through a three-layer one-dimensional convolutional network (with a parameter quantity <50 KB) to obtain noise-reduced data; the voice data conversion module is used to convert the noise-reduced data into a feature representation through 12 low-frequency Mel filters and 16 medium-high-frequency Mel filters; the frame length of the low-frequency Mel filters and the medium-high-frequency Mel filters is 25 ms, and the frame shift is 10 ms; the frame length determines the duration and the number of sampling points of each frame, affecting the frequency resolution; the frame shift determines the overlapping degree between adjacent frames, affecting the time resolution and the computational amount; the acoustic feature extraction module is used to extract acoustic features from the feature representation and is constructed based on a time-domain feature extraction unit and a frequency-domain feature extraction unit; the time-domain feature extraction unit is constructed based on a CNN network, and the frequency-domain feature extraction unit is constructed based on a BiLSTM network; the similarity calculation module is used to segment the voice data into several segments of voice sub-data according to phoneme boundaries, and calculate the similarity of the acoustic features corresponding to the voice sub-data according to the DTW algorithm improved by the slope constraint condition; the output module is used to output the voice matching result according to the similarity.

[0060] By setting that the acoustic feature extraction module is constructed based on a time-domain feature extraction unit and a frequency-domain feature extraction unit, the time-domain feature extraction unit is constructed based on a CNN network, and the frequency-domain feature extraction unit is constructed based on a Bi LSTM network, the acoustic features extracted by the acoustic feature extraction module combine time-domain features and frequency-domain features, effectively improving the feature extraction ability, and thus greatly improving the accuracy of speech matching.

[0061] A phoneme boundary refers to the demarcation point between different phonemes in a speech signal, that is, the end of one phoneme and the start position of another phoneme; phoneme boundary detection is an important topic in speech signal processing and is widely used in fields such as speech recognition, speech synthesis, speech coding, and language learning.

[0062] Introducing a slope constraint condition in the DTW algorithm (Dynamic Time Warping algorithm) aims to limit the slope of the bending path, thereby reducing the computational complexity and improving the efficiency of the algorithm. The traditional DTW algorithm allows the path to search freely within the entire matrix of the time series, which may lead to a relatively high computational complexity, while the slope constraint can significantly reduce the number of paths that need to be calculated, thus reducing the computational complexity.

[0063] The specific steps of step S2 are as follows:

[0064] Obtain a large amount of historical speech data, perform preprocessing on each of the historical speech data, including at least format conversion, noise reduction, sampling rate unification, and removal of silent segments, and construct a dataset after annotating similar speech data for each of the preprocessed historical speech data.

[0065] The specific steps of step S3 are as follows:

[0066] Based on a preset splitting ratio, divide the dataset into a training set, a validation set, and a test set. Train the speech matching model using the training set until the loss value of the loss function is less than a preset loss threshold. During the training process, compress the speech matching model using knowledge distillation technology and dynamic pruning technology; verify the trained speech matching model using the validation set and determine whether the matching accuracy is greater than a preset accuracy threshold. If not, the verification fails, expand the training set and continue training. If so, the verification is successful; test the speech matching model with successful verification using the test set and determine whether the confidence level is greater than a preset confidence level threshold. If not, the test fails, expand the training set and continue training. If so, the test is successful and the training ends.

[0067] By dividing the dataset into a training set, a validation set, and a test set, training the speech matching model with the training set until the loss value of the loss function is less than a preset loss threshold, and compressing the speech matching model through knowledge distillation technology and dynamic pruning technology during the training process; calculating the matching accuracy through the validation set to verify the trained speech matching model, and calculating the confidence through the test set to test the speech matching model that has passed the verification; that is, continuously compressing, validating, and testing during the training process of the speech matching model to effectively balance the model volume and matching accuracy of the speech matching model, so as to be better deployed on edge computing devices.

[0068] Knowledge distillation technology is a machine learning model compression method aimed at transferring the knowledge of a large model to a small model to improve the model performance and generalization ability; the core idea of knowledge distillation is to transform the knowledge of a complex model into a more concise and effective representation, which can reduce the computational complexity and resource requirements while maintaining high performance. Dynamic pruning technology aims to remove parts of the neural network that have little impact on the model performance (such as accuracy), such as neurons, connections (weights), etc., thereby reducing the model complexity and computational resource requirements. When applying dynamic pruning technology, channels with a contribution degree <15% in the CNN layer can be removed based on the importance score.

[0069] The specific steps of step S4 are as follows:

[0070] Deploy the trained speech matching model on edge computing devices of device types such as routers, gateways, or cameras. The edge computing device obtains a preset number of real-time speech data to construct a fine-tuning dataset, and based on a preset learning rate, batch size, and optimizer, performs adversarial training on the deployed speech matching model through the fine-tuning dataset, and then fine-tunes the deployed speech matching model.

[0071] A preferred embodiment of a speech matching system for routers, gateways, and cameras according to the present invention includes the following modules:

[0072] A speech matching model creation module for creating a speech matching model based on a speech data input module, a speech data conversion module, an acoustic feature extraction module, a similarity calculation module, and an output module, and setting the loss function of the speech matching model;

[0073] A dataset construction module for obtaining a large amount of historical speech data, preprocessing and annotating each piece of historical speech data, and constructing a dataset;

[0074] A speech matching model training module for training the speech matching model based on the dataset and the loss function, and compressing the speech matching model through knowledge distillation technology and dynamic pruning technology during the training process;

[0075] A voice matching model deployment module, which is used to deploy the trained voice matching model on edge computing devices of device types such as routers, gateways, or cameras. The edge computing devices fine-tune the deployed voice matching model; through fine-tuning, the model performance can be quickly improved on specific tasks;

[0076] A voice matching module, which is used for edge computing devices to obtain real-time voice data, preprocess the real-time voice data and then input it into the fine-tuned voice matching model to obtain a voice matching result, so as to complete voice matching;

[0077] A voice matching model optimization module, which is used for edge computing devices to continuously optimize the voice matching model based on the voice matching result and real-time voice data.

[0078] The present invention can achieve an end-to-end delay of <150ms on an ARM Cortex-A53 level processor, maintain an instruction recognition accuracy of ≥99% in an environment with a signal-to-noise ratio ≥5dB, and support real-time matching of voice data up to 30 seconds long.

[0079] The voice matching model creation module is specifically used for:

[0080] Create a voice matching model based on a voice data input module, a voice data conversion module, an acoustic feature extraction module, a similarity calculation module, and an output module, and set the loss function of the voice matching model; the voice data input module, the voice data conversion module, the acoustic feature extraction module, the similarity calculation module, and the output module are connected in sequence; the loss function is constructed based on a contrastive loss function and a binary cross-entropy loss function;

[0081] By setting the loss function to be constructed based on a contrastive loss function and a binary cross-entropy loss function, that is, weighting the contrastive loss function and the binary cross-entropy loss function; since the contrastive loss function is used to learn the feature embedding space, making the distance between similar samples closer and the distance between dissimilar samples farther, and the binary cross-entropy loss function is used to output binary classification problems, it can effectively combine the advantages of the contrastive loss function and the binary cross-entropy loss function, thereby greatly improving the training effect of the voice matching model.

[0082] The voice data input module is used to perform first-stage noise reduction on the input voice data through spectral subtraction, and then perform second-stage noise reduction on the voice data through a three-layer one-dimensional convolutional network (parameter quantity <50KB) to obtain noise-reduced data; the voice data conversion module is used to convert the noise-reduced data into a feature representation through 12 low-frequency Mel filters and 16 medium-high frequency Mel filters; the frame length of the low-frequency Mel filters and the medium-high frequency Mel filters is 25ms, and the frame shift is 10ms; the frame length determines the duration and the number of sampling points of each frame, affecting the frequency resolution; the frame shift determines the overlapping degree between adjacent frames, affecting the time resolution and the calculation amount; the acoustic feature extraction module is used to extract acoustic features from the feature representation and is constructed based on a time-domain feature extraction unit and a frequency-domain feature extraction unit; the time-domain feature extraction unit is constructed based on a CNN network, and the frequency-domain feature extraction unit is constructed based on a Bi LSTM network; the similarity calculation module is used to segment the voice data into several segments of voice sub-data according to the phoneme boundary, and calculate the similarity of the acoustic features corresponding to the voice sub-data according to the DTW algorithm improved by the slope constraint condition; the output module is used to output the voice matching result according to the similarity.

[0083] By setting that the acoustic feature extraction module is constructed based on a time-domain feature extraction unit and a frequency-domain feature extraction unit, the time-domain feature extraction unit is constructed based on a CNN network, and the frequency-domain feature extraction unit is constructed based on a BiLSTM network, the acoustic features extracted by the acoustic feature extraction module combine time-domain features and frequency-domain features, effectively improving the feature extraction ability, and thus greatly improving the accuracy of voice matching.

[0084] The phoneme boundary refers to the demarcation point between different phonemes in the voice signal, that is, the end position of one phoneme and the start position of another phoneme; phoneme boundary detection is an important topic in voice signal processing and is widely used in fields such as speech recognition, speech synthesis, speech coding, and language learning.

[0085] Introducing a slope constraint condition into the DTW algorithm (Dynamic Time Warping algorithm) aims to limit the slope of the bending path, thereby reducing the computational complexity and improving the efficiency of the algorithm. The traditional DTW algorithm allows the path to freely search within the entire matrix of the time series, which may lead to a relatively high computational complexity, while the slope constraint can significantly reduce the number of paths that need to be calculated, thereby reducing the computational complexity.

[0086] The dataset construction module is specifically used for:

[0087] Obtain a large amount of historical voice data, perform preprocessing on each of the historical voice data including at least format conversion, noise reduction, sampling rate unification, and removal of silent segments, and construct a dataset after annotating similar voice data for each of the preprocessed historical voice data.

[0088] The voice matching model training module is specifically used for:

[0089] Dividing the data set into a training set, a validation set, and a test set based on a preset segmentation ratio, training the voice matching model through the training set until the loss value of the loss function is less than a preset loss threshold, and compressing the voice matching model through knowledge distillation technology and dynamic pruning technology during the training process; validating the trained voice matching model through the validation set to determine whether the matching accuracy rate is greater than a preset accuracy threshold. If not, the validation fails, and the training set is expanded and training continues. If so, the validation succeeds; testing the voice matching model that has passed the validation through the test set to determine whether the confidence level is greater than a preset confidence threshold. If not, the test fails, and the training set is expanded and training continues. If so, the test succeeds, and the training ends.

[0090] By dividing the data set into a training set, a validation set, and a test set, training the voice matching model through the training set until the loss value of the loss function is less than a preset loss threshold, and compressing the voice matching model through knowledge distillation technology and dynamic pruning technology during the training process; calculating the matching accuracy rate through the validation set to validate the trained voice matching model, and calculating the confidence level through the test set to test the voice matching model that has passed the validation; that is, continuously compressing, validating, and testing during the training process of the voice matching model to effectively balance the model volume and matching accuracy of the voice matching model, so as to be better deployed on edge computing devices.

[0091] Knowledge distillation technology is a machine learning model compression method aimed at transferring the knowledge of a large model to a small model to improve the model performance and generalization ability; the core idea of knowledge distillation is to transform the knowledge of a complex model into a more concise and effective representation, so as to reduce the computational complexity and resource requirements while maintaining high performance. Dynamic pruning technology aims to remove parts of the neural network that have less impact on the model performance (such as accuracy), such as neurons, connections (weights), etc., thereby reducing the complexity and computational resource requirements of the model. When applying dynamic pruning technology, channels with a contribution degree <15% in the CNN layer can be removed based on the importance score.

[0092] The voice matching model deployment module is specifically used for:

[0093] Deploying the trained voice matching model on edge computing devices of device types such as routers, gateways, or cameras. The edge computing device obtains a preset number of real-time voice data to construct a fine-tuning data set, and based on a preset learning rate, batch size, and optimizer, performs adversarial training on the deployed voice matching model through the fine-tuning data set, and then fine-tunes the deployed voice matching model.

[0094] In summary, the advantages of the present invention are as follows:

[0095] 1. A voice matching model is created through a voice data input module, a voice data conversion module, an acoustic feature extraction module, a similarity calculation module, and an output module, and the loss function of the voice matching model is set; then a large amount of historical voice data is obtained, preprocessed and labeled to construct a data set, and the voice matching model is trained based on the data set and the loss function. During the training process, the voice matching model is compressed through knowledge distillation technology and dynamic pruning technology, and then the trained voice matching model is deployed on an edge computing device with a device type of router, gateway, or camera. The edge computing device fine-tunes the deployed voice matching model; finally, the edge computing device obtains real-time voice data, preprocesses the real-time voice data and inputs it into the fine-tuned voice matching model to obtain a voice matching result, and continuously optimizes the voice matching model based on the voice matching result and the real-time voice data; that is, voice matching is performed based on the voice matching model created through the voice data input module, the voice data conversion module, the acoustic feature extraction module, the similarity calculation module, and the output module. The Mel filter in the voice data conversion module is optimized from the conventional 40 to 28 (12 low-frequency Mel filters and 16 medium-high-frequency Mel filters) to simplify the network structure; the similarity calculation module segments the voice data according to phoneme boundaries and calculates the similarity using the DTW algorithm improved based on slope constraint conditions to improve the calculation efficiency; the voice matching model is compressed by combining knowledge distillation technology and dynamic pruning technology, effectively reducing the volume of the voice matching model, facilitating the deployment of the voice matching model on resource-constrained edge computing devices; and the voice data input module performs two-stage noise reduction (spectral subtraction, one-dimensional convolutional network) on the input voice, effectively overcoming the influence of noise and improving the robustness. Combining the fine-tuning and optimization of the voice matching model, ultimately, the resource requirements for voice matching are greatly reduced, and the accuracy and timeliness of voice matching are greatly improved, thereby greatly enhancing the user experience.

[0096] 2. By setting the loss function based on a combination of a contrastive loss function and a binary cross-entropy loss function, that is, weighting the contrastive loss function and the binary cross-entropy loss function; since the contrastive loss function is used to learn the feature embedding space, making the distance between similar samples closer and the distance between dissimilar samples farther, and the binary cross-entropy loss function is used to output binary classification problems, the advantages of the contrastive loss function and the binary cross-entropy loss function can be effectively combined, thereby greatly improving the training effect of the voice matching model.

[0097] 3. By setting that the acoustic feature extraction module is constructed based on the time-domain feature extraction unit and the frequency-domain feature extraction unit, the time-domain feature extraction unit is constructed based on the CNN network, and the frequency-domain feature extraction unit is constructed based on the Bi LSTM network, the acoustic features extracted by the acoustic feature extraction module combine the time-domain features and the frequency-domain features, effectively improving the feature extraction ability, and further greatly improving the accuracy of speech matching.

[0098] 4. By dividing the data set into a training set, a validation set and a test set, training the speech matching model with the training set until the loss value of the loss function is less than the preset loss threshold, and compressing the speech matching model through the knowledge distillation technology and the dynamic pruning technology during the training process; calculating the matching accuracy through the validation set to verify the trained speech matching model, and calculating the confidence through the test set to test the speech matching model that has passed the verification; that is, continuously compressing, verifying and testing during the training process of the speech matching model to effectively balance the model volume and the matching accuracy of the speech matching model, so as to be better deployed on edge computing devices.

[0099] Although the specific implementation manners of the present invention have been described above, those skilled in the art of this technology should understand that the specific embodiments we described are illustrative rather than used to limit the scope of the present invention. Equivalent modifications and variations made by those skilled in the art in accordance with the spirit of the present invention should be covered by the scope protected by the claims of the present invention.

Claims

1. A voice matching method for routers, gateways, and cameras, characterized in that: The steps include: Step S1, creating a voice matching model based on the voice data input module, the voice data conversion module, the acoustic feature extraction module, the similarity calculation module and the output module, and setting the loss function of the voice matching model; Step S2, obtaining a large amount of historical voice data, and constructing a data set after preprocessing and annotating each of the historical voice data; Step S3, training the speech matching model based on the data set and the loss function, and compressing the speech matching model through knowledge distillation technology and dynamic pruning technology during the training process; Step S4: deploy the trained voice matching model on an edge computing device whose device type is a router, a gateway or a camera, and the edge computing device fine-tunes the deployed voice matching model; Step S5: The edge computing device obtains real-time voice data, pre-processes the real-time voice data, and inputs the fine-tuned voice matching model to obtain a voice matching result to complete the voice matching; Step S6: The edge computing device continuously optimizes the voice matching model based on the voice matching results and real-time voice data.

2. A voice matching method for a router, a gateway, and a camera as claimed in claim 1, characterized in that: The step S1 is specifically as follows: A speech matching model is created based on a speech data input module, a speech data conversion module, an acoustic feature extraction module, a similarity calculation module and an output module, and a loss function of the speech matching model is set; the speech data input module, the speech data conversion module, the acoustic feature extraction module, the similarity calculation module and the output module are connected in sequence; the loss function is constructed based on a contrast loss function and a binary cross entropy loss function; The speech data input module is used to perform a first-stage noise reduction on the input speech data through spectral subtraction, and then perform a second-stage noise reduction through a three-layer one-dimensional convolutional network to obtain noise reduction data; the speech data conversion module is used to convert the noise reduction data into feature representation through 12 low-frequency Mel filters and 16 medium- and high-frequency Mel filters; the frame length of the low-frequency Mel filter and the medium- and high-frequency Mel filter is 25ms, and the frame shift is 10ms; the acoustic feature extraction module is used to extract acoustic features from the feature representation, and is constructed based on a time domain feature extraction unit and a frequency domain feature extraction unit; the time domain feature extraction unit is constructed based on a CNN network, and the frequency domain feature extraction unit is constructed based on a BiLSTM network; the similarity calculation module is used to segment the speech data into several segments of speech sub-data according to the phoneme boundary, and calculate the similarity of the acoustic features corresponding to the speech sub-data according to the DTW algorithm improved by the slope constraint condition; the output module is used to output the speech matching result according to the similarity.

3. A voice matching method for a router, a gateway, and a camera as claimed in claim 1, characterized in that: The step S2 is specifically as follows: A large amount of historical voice data is obtained, and each of the historical voice data is preprocessed including at least format conversion, noise reduction, sampling rate unification, and removal of silent segments. After the preprocessed historical voice data are annotated with similar voice data, a data set is constructed.

4. A voice matching method for a router, a gateway, and a camera as claimed in claim 1, characterized in that: The step S3 is specifically as follows: The data set is divided into a training set, a validation set and a test set based on a preset segmentation ratio; the speech matching model is trained with the training set until the loss value of the loss function is less than a preset loss threshold; the speech matching model is compressed with knowledge distillation technology and dynamic pruning technology during the training process; the trained speech matching model is verified with the validation set to determine whether the matching accuracy is greater than a preset accuracy threshold; if not, the verification fails, and the training set is expanded to continue training; if yes, the verification succeeds; the successfully verified speech matching model is tested with the test set to determine whether the confidence is greater than a preset confidence threshold; if not, the test fails, and the training set is expanded to continue training; if yes, the test succeeds and the training ends.

5. A voice matching method for a router, a gateway, and a camera as claimed in claim 1, characterized in that: The step S4 is specifically as follows: The trained voice matching model is deployed on an edge computing device whose device type is a router, gateway or camera. The edge computing device obtains a preset amount of real-time voice data to build a fine-tuning data set. Based on a preset learning rate, batch size and optimizer, adversarial training is performed on the deployed voice matching model through the fine-tuning data set, and then the deployed voice matching model is fine-tuned.

6. A voice matching system for routers, gateways, and cameras, characterized in that: Includes the following modules: A voice matching model creation module, used to create a voice matching model based on the voice data input module, the voice data conversion module, the acoustic feature extraction module, the similarity calculation module and the output module, and set the loss function of the voice matching model; A data set construction module is used to obtain a large amount of historical speech data, and to construct a data set after preprocessing and annotating each of the historical speech data; A speech matching model training module, used to train the speech matching model based on the data set and the loss function, and compress the speech matching model through knowledge distillation technology and dynamic pruning technology during the training process; A voice matching model deployment module, used to deploy the trained voice matching model on an edge computing device whose device type is a router, a gateway or a camera, and the edge computing device fine-tunes the deployed voice matching model; The voice matching module is used for the edge computing device to obtain real-time voice data, pre-process the real-time voice data and input the fine-tuned voice matching model to obtain the voice matching result to complete the voice matching; The voice matching model optimization module is used for the edge computing device to continuously optimize the voice matching model based on the voice matching results and real-time voice data.

7. A voice matching system for a router, a gateway, and a camera as claimed in claim 6, characterized in that: The speech matching model creation module is specifically used for: A speech matching model is created based on a speech data input module, a speech data conversion module, an acoustic feature extraction module, a similarity calculation module and an output module, and a loss function of the speech matching model is set; the speech data input module, the speech data conversion module, the acoustic feature extraction module, the similarity calculation module and the output module are connected in sequence; the loss function is constructed based on a contrast loss function and a binary cross entropy loss function; The speech data input module is used to perform a first-stage noise reduction on the input speech data through spectral subtraction, and then perform a second-stage noise reduction through a three-layer one-dimensional convolutional network to obtain noise reduction data; the speech data conversion module is used to convert the noise reduction data into feature representation through 12 low-frequency Mel filters and 16 medium- and high-frequency Mel filters; the frame length of the low-frequency Mel filter and the medium- and high-frequency Mel filter is 25ms, and the frame shift is 10ms; the acoustic feature extraction module is used to extract acoustic features from the feature representation, and is constructed based on a time domain feature extraction unit and a frequency domain feature extraction unit; the time domain feature extraction unit is constructed based on a CNN network, and the frequency domain feature extraction unit is constructed based on a BiLSTM network; the similarity calculation module is used to segment the speech data into several segments of speech sub-data according to the phoneme boundary, and calculate the similarity of the acoustic features corresponding to the speech sub-data according to the DTW algorithm improved by the slope constraint condition; the output module is used to output the speech matching result according to the similarity.

8. A voice matching system for a router, a gateway, and a camera as claimed in claim 6, characterized in that: The data set construction module is specifically used for: A large amount of historical voice data is obtained, and each of the historical voice data is preprocessed including at least format conversion, noise reduction, sampling rate unification, and removal of silent segments. After the preprocessed historical voice data are annotated with similar voice data, a data set is constructed.

9. A voice matching system for a router, a gateway, and a camera as claimed in claim 6, characterized in that: The speech matching model training module is specifically used for: The data set is divided into a training set, a validation set and a test set based on a preset segmentation ratio; the speech matching model is trained with the training set until the loss value of the loss function is less than a preset loss threshold; the speech matching model is compressed with knowledge distillation technology and dynamic pruning technology during the training process; the trained speech matching model is verified with the validation set to determine whether the matching accuracy is greater than a preset accuracy threshold; if not, the verification fails, and the training set is expanded to continue training; if yes, the verification succeeds; the successfully verified speech matching model is tested with the test set to determine whether the confidence is greater than a preset confidence threshold; if not, the test fails, and the training set is expanded to continue training; if yes, the test succeeds and the training ends.

10. A voice matching system for a router, a gateway, and a camera as claimed in claim 6, characterized in that: The speech matching model deployment module is specifically used for: The trained voice matching model is deployed on an edge computing device whose device type is a router, gateway or camera. The edge computing device obtains a preset amount of real-time voice data to build a fine-tuning data set. Based on a preset learning rate, batch size and optimizer, adversarial training is performed on the deployed voice matching model through the fine-tuning data set, and then the deployed voice matching model is fine-tuned.

Citation Information

Cited By

  • Target detection method and system combining router, gateway and camera

    CN121030429A