Edge computing device and method for voice copying
Through edge computing devices and deep learning models, the misjudgment and poor real-time real-time real-time real-time problem of speech recognition in complex environments is solved, and high accuracy and real-time voice imprinting is achieved, which is applied to conference records, legal documents and medical records.
Patent Information
- Application Number
- CN202510262542.4
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-03-06
- Publication Date
- 2025-08-08
AI Technical Summary
Existing voice recognition technology is prone to misjudgment and poor real-time performance in complex environments, and is difficult to meet the high real-time requirements of voice copying application scenarios.
Edge computing devices are adopted, including speech acquisition unit, feature extraction unit, edge computing unit and storage unit. Through deep learning speech synthesis model and attention transfer model, voice signals similar to the speech signal to be copied are generated, and real-time processing and storage are performed.
It improves the accuracy and real-timeness of voice copying, solves the problem of misjudgment in complex environments, meets application scenarios with high requirements for real-timeness, and improves the work efficiency in the fields of meeting minutes, legal documents and medical records.
Smart Images

Figure CN120452420A_ABST
Abstract
Description
Technical Field
[0001] The present invention relates to the field of edge computing and speech recognition technology, and more particularly to an edge computing device and method for speech imprinting. Background Art
[0002] In today's digital age, speech recognition technology, as a key branch of artificial intelligence, has experienced rapid development. As a key application of speech recognition technology, voice transcription has gradually penetrated numerous fields. For example, in meeting record-keeping scenarios, it can convert participants' speeches into text in real time, greatly improving the efficiency and accuracy of meeting records. In the field of legal documents, voice transcription helps quickly record witness testimony, lawyer debates, and other content, facilitating judicial work. In medical records, doctors can use voice transcription to more conveniently record patient conditions, diagnoses, and treatment plans, thereby optimizing medical service processes.
[0003] However, existing speech processing technologies still face numerous challenges in achieving accurate speech capture. For one thing, speech recognition accuracy needs to be improved, especially in complex environments or when dealing with users with similar speech characteristics, where misjudgment is prone to occur. Furthermore, traditional speech processing methods are mostly based on centralized computing architectures, which suffer from high data transmission latency and poor real-time performance, making them difficult to meet the requirements of speech capture applications requiring high real-time performance.
[0004] Therefore, how to improve the accuracy and real-time performance of voice imitation is a problem that those skilled in the art need to solve urgently. Summary of the Invention
[0005] In view of this, the present invention provides an edge computing device and method for voice imprinting to solve the problems existing in the background technology.
[0006] In order to achieve the above object, the present invention adopts the following technical solutions:
[0007] An edge computing device for voice imprinting, comprising: a voice acquisition unit, a feature extraction unit, an edge computing unit, and a storage unit;
[0008] The voice acquisition unit acquires the voice signal to be transcribed in real time and pre-processes it;
[0009] The feature extraction unit extracts features from the pre-processed speech signal to be transcribed to obtain key speech feature parameters;
[0010] The edge computing unit uses a deep learning-based speech synthesis model to take the extracted key speech feature parameters as input to generate a speech signal similar to the speech signal to be emulated;
[0011] The storage unit uses a memory to store the speech signal to be emulated and the generated similar speech signal.
[0012] Optionally, after collecting the voice signal to be emulated, the voice acquisition unit is further used to amplify, gain control, filter and sample the voice signal to be emulated in sequence, and then perform format conversion and encoding on the voice signal to be emulated, so that the voice signal to be emulated is divided into short-time signals composed of multiple frames.
[0013] Optionally, the feature extraction unit obtains the key speech feature parameters by extracting frequency cepstral coefficient MFCC features from the preprocessed speech signal to be transcribed.
[0014] The optional deep learning-based speech synthesis model is:
[0015] The acquired short-time signal is used as the first-layer convolution input, and the input short-time signal is convolved layer by layer according to the key speech feature parameters. The output layer is the feature representation layer of the short-time signal. The features output by the feature representation layer are used as the input features of the SVM classifier, and the recognition task of the speech imprint is completed by the SVM classifier. In this embodiment, an attention transfer model is introduced into the convLSTM network of the speech synthesis model; the attention transfer model includes a channel attention module and a spatial attention module;
[0016] The channel attention module is used to convert each two-dimensional feature channel into a real number and generate an intermediate map representing the dependency relationship between channels;
[0017] The spatial attention module is used to compress and generate a feature matrix in the tensor space of the channel dimension, and then obtain a two-dimensional spatial attention map through the softmax activation function.
[0018] Optionally, the attention function used by the channel attention module is:
[0019]
[0020] Among them, C i is the output result of the previous convolution layer of channel i, is the convolution output after the channel attention module conversion, σ represents the i The standard deviation on C i ζ represents the mean of the attention weights obtained using the Gaussian function.
[0021] Optionally, the spatial attention module is used to compress and generate three feature matrices using three pooling methods in the tensor space of the channel dimension, then generate a unified pooling layer through a set fusion rule, and then obtain a two-dimensional spatial attention map through a softmax activation function; wherein the three pooling methods include maximum pooling, local saliency value pooling, and transfer attention value pooling;
[0022] The pooling layer function corresponding to the transfer attention value pooling method is
[0023]
[0024] Among them, w i is the superposition output of the saliency weight matrix of the previous layer or the previous layers, σ w Indicates that W i The standard deviation on w Indicates that i The mean value on w Represents the pooling weight value obtained using the Gaussian function, Represents the weight matrix of the pooling layer after transferring attention and saliency.
[0025] Optionally, the set fusion rule is specifically:
[0026] P=λ1P m +λ2P t +λ3P1
[0027] stλ1+λ2+λ3=1
[0028] Among them, P m is the pooling layer function corresponding to the maximum pooling method, P1 is the pooling layer function corresponding to the local saliency value pooling method, P is a unified pooling layer function, and λ1, λ2, and λ3 all represent feature weighted fusion normalization constraint parameters.
[0029] Optionally, the workflow of the edge computing unit is:
[0030] Establish a speech synthesis model based on deep learning;
[0031] Receive a short-term signal composed of multiple frames pre-processed by a voice acquisition unit;
[0032] The short-term signal composed of multiple frames is sequentially input into the speech synthesis model based on deep learning;
[0033] The key speech feature parameters are used to perform a convolution operation on the input short-term signal composed of multiple frames. The attention transfer model is added to generate new convolution parameter weights. The attention transfer model is applied to the subsequent convolution and pooling layer operations of the input speech signal, transferring the previous learning experience to the current speech recognition process to generate a fusion target feature representation.
[0034] The output result of the last layer is used as the input feature of the SVM classifier to realize the imprinting of the speech signal to be imprinted.
[0035] Optionally, the memory of the storage unit adopts 2G memory + 16G EMMC.
[0036] An edge computing method for voice imprinting, comprising:
[0037] Collect the voice signal to be transcribed in real time and pre-process it;
[0038] Extract features from the pre-processed speech signal to be transcribed to obtain key speech feature parameters;
[0039] A speech synthesis model based on deep learning is used to take the extracted key speech feature parameters as input to generate a speech signal similar to the speech signal to be imitated;
[0040] The memory is used to store the speech signal to be imitated and the generated similar speech signal.
[0041] It can be seen from the above technical solution that compared with the existing technology, the present invention discloses an edge computing device and method for voice imprinting, including a voice collection unit, a feature extraction unit, an edge computing unit and a storage unit; the present invention improves the accuracy of voice imprinting as a whole, solves the problem that the existing technology is prone to misjudgment in complex environments or with similar voice features, and reduces data transmission delay based on edge computing, improves real-time performance, and meets voice imprinting application scenarios with high real-time requirements. It can also be applied to conference records, legal documents, medical records and other fields to improve work efficiency and service quality, and has broad application prospects and practical value. BRIEF DESCRIPTION OF THE DRAWINGS
[0042] In order to more clearly illustrate the embodiments of the present invention or the technical solutions in the prior art, the following briefly introduces the drawings required for use in the embodiments or the description of the prior art. Obviously, the drawings described below are merely embodiments of the present invention. For ordinary technicians in this field, other drawings can be obtained based on the provided drawings without paying any creative work.
[0043] Figure 1 A schematic diagram of the device structure provided by the present invention;
[0044] Figure 2 The present invention provides a flow chart of the method. DETAILED DESCRIPTION
[0045] The following will clearly and completely describe the technical solutions in the embodiments of the present invention in conjunction with the accompanying drawings. Obviously, the described embodiments are only part of the embodiments of the present invention, not all of the embodiments. Based on the embodiments of the present invention, all other embodiments obtained by ordinary technicians in this field without making creative efforts are within the scope of protection of the present invention.
[0046] The embodiment of the present invention discloses an edge computing device for voice imprinting, such as Figure 1 As shown, it includes: a voice acquisition unit, a feature extraction unit, an edge computing unit, and a storage unit;
[0047] The voice acquisition unit collects the voice signal to be transcribed in real time and pre-processes it;
[0048] The feature extraction unit extracts features from the pre-processed speech signal to be transcribed to obtain key speech feature parameters;
[0049] The edge computing unit uses a deep learning-based speech synthesis model, takes the extracted key speech feature parameters as input, and generates a speech signal similar to the speech signal to be emulated;
[0050] The storage unit uses a memory to store the speech signal to be imitated and the generated similar speech signal.
[0051] In a specific embodiment, after collecting the voice signal to be emulated, the voice acquisition unit is further used to sequentially amplify, gain control, filter and sample the voice signal to be emulated, and then perform format conversion and encoding on the voice signal to be emulated, so that the voice signal to be emulated is divided into short-time signals composed of multiple frames.
[0052] Specifically, after collecting the voice signal to be transcribed, the voice collection unit is used to amplify, gain control, filter and sample the voice signal to be transcribed in sequence, and then perform the following steps on the voice signal to be transcribed:
[0053] The system performs format conversion and coding so that the voice signal to be copied is divided into short-time signals composed of multiple frames; and is also used to perform pre-emphasis processing on the voice signal to be copied after format conversion and coding using a window function.
[0054] In speaker recognition technology, speech acquisition is actually the process of digitizing speech signals. Through amplification and gain control, anti-aliasing filtering, sampling, A / D (analog / digital) conversion and encoding (generally PCM (pulse code modulation) code), the simulated speech signal is filtered and amplified, and the filtered and amplified analog speech signal is converted into a digital speech signal.
[0055] In the above process, filtering is performed to suppress all components of the frequency domain components of the input signal whose frequencies exceed fs / 2 (fs is the sampling frequency) to prevent aliasing interference and at the same time suppress the 50Hz power supply frequency interference.
[0056] The voice acquisition unit also performs the reverse process of digitizing the encoded voice signal to be digitized, reconstructing the voice waveform from the digitized voice. This is known as D / A (digital-to-analog) conversion. Furthermore, smoothing filtering is required after D / A conversion to smooth the higher harmonics of the reconstructed voice waveform and remove harmonic distortion.
[0057] Through the processing described above, the speech signal is segmented into short-duration frames. Each frame is then treated as a stationary random signal, and digital signal processing techniques are used to extract speech feature parameters. During processing, data is extracted from the data area frame by frame. After processing is complete, the next frame is extracted, and so on. Ultimately, a time series of speech feature parameters is obtained, consisting of the parameters of each frame.
[0058] In addition, the voice collection unit is also used to perform pre-emphasis processing on the voice signal to be transcribed after format conversion and encoding using a window function.
[0059] In a specific embodiment, the feature extraction unit obtains key speech feature parameters by extracting frequency cepstral coefficient MFCC features from the preprocessed speech signal to be transcribed.
[0060] MFCC parameters have the following advantages (compared to LPCC parameters):
[0061] Speech information is mostly concentrated in the low-frequency part, while the high-frequency part is easily interfered by environmental noise. The MFCC parameter converts the linear frequency scale into the Mel frequency scale, emphasizing the low-frequency information of speech. In addition to the advantages of LPCC, it also highlights the information that is conducive to recognition and shields the interference of noise. LPCC parameters are based on the linear frequency scale and therefore do not have this feature.
[0062] MFCC parameters have no assumptions and can be used in various situations. However, LPCC parameters assume that the signal they process is an AR signal. For consonants with strong dynamic characteristics, this assumption is not strictly true. Therefore, MFCC parameters are better than LPCC parameters in speaker recognition.
[0063] FFT transformation is required in the MFCC parameter extraction process to obtain all the information in the frequency domain of the speech signal.
[0064] The height of the sound heard by the human ear is not linearly proportional to the frequency of the sound, and the Mel frequency scale is more in line with the auditory characteristics of the human ear. The so-called Mel frequency scale, its value roughly corresponds to the logarithmic distribution of the actual frequency. The specific relationship between Mel frequency and actual frequency can be expressed as: Mel(f) = 2595lg(1+f / 700), where the unit of actual frequency f is Hz. The critical frequency bandwidth changes with the change of frequency and is consistent with the growth of Mel frequency. Below 1000Hz, it is roughly linearly distributed with a bandwidth of about 100Hz; above 1000Hz, it grows logarithmically. Similar to the division of critical bands, the speech frequency can be divided into a series of triangular filter sequences, namely the Mel filter group
[0065] The output of the triangular filter is:
[0066]
[0067] Where n=1,2,…,P;Y n is the output of the nth filter.
[0068] The filter output is transformed into the cepstral domain using the discrete cosine transform (DCT):
[0069]
[0070] Where P is the order of MFCC parameters, and P=12 is selected in the actual software algorithm. k} k =1,2,…,12 is the desired MFCC parameters.
[0071] In a specific embodiment, the speech synthesis model based on deep learning is specifically:
[0072] The acquired short-time signal is used as the input of the first convolution layer. The input short-time signal is convolved layer by layer according to the key speech feature parameters. The output layer is the feature representation layer of the short-time signal. The features output by the feature representation layer are used as the input features of the SVM classifier, and the SVM classifier completes the speech imprint recognition task. Among them, the attention transfer model is introduced into the convLSTM network of the speech synthesis model; the attention transfer model includes a channel attention module and a spatial attention module.
[0073] The channel attention module is used to convert each two-dimensional feature channel into a real number and generate an intermediate map representing the dependency relationship between channels;
[0074] The spatial attention module is used to compress and generate a feature matrix in the tensor space of the channel dimension, and then obtain a two-dimensional spatial attention map through the softmax activation function.
[0075] The attention function used by the channel attention module is:
[0076]
[0077] Among them, C i is the output result of the previous convolutional layer of channel i, is the convolution output after the channel attention module conversion, σ represents the i The standard deviation on C i ζ represents the mean of the attention weights obtained using the Gaussian function.
[0078] The spatial attention module is used to compress and generate three feature matrices using three pooling methods in the channel-dimensional tensor space. It then generates a unified pooling layer through the set fusion rules, and finally obtains a two-dimensional spatial attention map through the softmax activation function. The three pooling methods include maximum pooling, local saliency pooling, and transfer attention pooling.
[0079] The pooling layer function corresponding to the transfer attention value pooling method is
[0080]
[0081] Among them, w i is the superposition output of the saliency weight matrix of the previous layer or the previous layers, σ w Indicates that W i The standard deviation on w Indicates that i The mean value on w Represents the pooling weight value obtained using the Gaussian function, Represents the weight matrix of the pooling layer after transferring attention and saliency.
[0082] The specific fusion rules are as follows:
[0083] P=λ1P m +λ2P t +λ3P1
[0084] stλ1+λ2+λ3=1
[0085] Among them, P m is the pooling layer function corresponding to the maximum pooling method, P1 is the pooling layer function corresponding to the local saliency value pooling method, P is a unified pooling layer function, and λ1, λ2, and λ3 all represent feature weighted fusion normalization constraint parameters.
[0086] Specifically, the maximum pooling P m Use the conventional maximum pooling operation to retain the local maximum of the upper layer output and retain its most influential element output to generate attention Figure 1 ; Local saliency pooling P1 generates attention by convolution operation based on the output result generated by the upper channel attention model and the weight matrix of this layer Figure 2 ; Migrate attention pooling layer P t , where W i The superposition output of the saliency weight matrix of the previous layer or the previous layers is used to generate a new attention mapping matrix through the attention model, and then a convolution operation is performed with the upper layer output result to generate an attention map 3.
[0087] Finally, the target saliency distribution features are converted into an attention matrix, and after a certain affine transformation, the inner product is calculated with the original weight parameters. The result is then embedded into the convolution layer of each channel to complete the recalibration in the channel dimension.
[0088] In this embodiment of the present invention, the attention map generated by the attention transfer model can implement pruning operations in deep convolutional networks, discarding redundant parameters. In other words, it can enhance useful information while suppressing useless information. The attention transfer mechanism can transfer this advantage to the training of new unlabeled datasets, quickly obtaining target locations and feature representations.
[0089] In a specific embodiment, the workflow of the edge computing unit is as follows:
[0090] Establish a speech synthesis model based on deep learning;
[0091] Receive a short-term signal composed of multiple frames pre-processed by a voice acquisition unit;
[0092] The short-term signal composed of multiple frames is sequentially input into the speech synthesis model based on deep learning;
[0093] The key speech feature parameters are used to perform a convolution operation on the input short-term signal composed of multiple frames. The attention transfer model is added to generate new convolution parameter weights. The attention transfer model is applied to the subsequent convolution and pooling layer operations of the input speech signal, transferring the previous learning experience to the current speech recognition process to generate a fusion target feature representation.
[0094] The output result of the last layer is used as the input feature of the SVM classifier to realize the imprinting of the speech signal to be imprinted.
[0095] In a specific embodiment, the memory of the storage unit uses 2G memory + 16G EMMC to store data, and uses memory to store and quickly process real-time data. Data with low real-time requirements can be stored through EMMC. This data storage method has great advantages in improving the processing rate of edge computing and improving hardware resource utilization.
[0096] An edge computing method for voice imprinting, such as Figure 2 Shown, including:
[0097] Collect the voice signal to be transcribed in real time and pre-process it;
[0098] Extract features from the pre-processed speech signal to be transcribed to obtain key speech feature parameters;
[0099] A speech synthesis model based on deep learning is used to take the extracted key speech feature parameters as input to generate a speech signal similar to the speech signal to be imitated;
[0100] The memory is used to store the speech signal to be imitated and the generated similar speech signal.
[0101] The various embodiments in this specification are described in a progressive manner, with each embodiment focusing on the differences from other embodiments. Reference can be made to the common and similar parts between the various embodiments. For the devices disclosed in the embodiments, since they correspond to the methods disclosed in the embodiments, the description is relatively simple, and the relevant parts can be referred to the method description.
[0102] The above description of the disclosed embodiments is intended to enable one skilled in the art to implement or use the present invention. Various modifications to these embodiments will be readily apparent to one skilled in the art, and the general principles defined herein may be implemented in other embodiments without departing from the spirit or scope of the present invention. Therefore, the present invention is not limited to the embodiments shown herein but is intended to conform to the widest scope consistent with the principles and novel features disclosed herein.
Claims
1. An edge computing device for voice imprinting, characterized in that: include: Voice collection unit, feature extraction unit, edge computing unit, storage unit; The voice acquisition unit acquires the voice signal to be transcribed in real time and pre-processes it; The feature extraction unit extracts features from the pre-processed speech signal to be transcribed to obtain key speech feature parameters; The edge computing unit uses a deep learning-based speech synthesis model to take the extracted key speech feature parameters as input to generate a speech signal similar to the speech signal to be emulated; The storage unit uses a memory to store the speech signal to be emulated and the generated similar speech signal.
2. The edge computing device for voice imprinting according to claim 1, characterized in that: After collecting the voice signal to be transcribed, the voice collection unit is further used to sequentially amplify, gain control, filter and sample the voice signal to be transcribed, and then perform format conversion and encoding on the voice signal to be transcribed, so that the voice signal to be transcribed is divided into short-time signals composed of multiple frames.
3. The edge computing device for voice imprinting according to claim 1, characterized in that: The feature extraction unit obtains the key speech feature parameters by extracting frequency cepstral coefficient MFCC features from the preprocessed speech signal to be transcribed.
4. The edge computing device for voice imprinting according to claim 2, characterized in that: The speech synthesis model based on deep learning is specifically: The acquired short-time signal is used as the first-layer convolution input, and the input short-time signal is convolved layer by layer according to the key speech feature parameters. The output layer is the feature representation layer of the short-time signal. The features output by the feature representation layer are used as the input features of the SVM classifier, and the recognition task of the speech imprint is completed by the SVM classifier. In this embodiment, an attention transfer model is introduced into the convLSTM network of the speech synthesis model; the attention transfer model includes a channel attention module and a spatial attention module; The channel attention module is used to convert each two-dimensional feature channel into a real number and generate an intermediate map representing the dependency relationship between channels; The spatial attention module is used to compress and generate a feature matrix in the tensor space of the channel dimension, and then obtain a two-dimensional spatial attention map through the softmax activation function.
5. The edge computing device for voice imprinting according to claim 4, characterized in that: The attention function used by the channel attention module is: Among them, C i is the output result of the previous convolutional layer of channel i, is the convolution output after the channel attention module conversion, σ represents the i The standard deviation on C i ζ represents the mean of the attention weights obtained using the Gaussian function.
6. The edge computing device for voice imprinting according to claim 4, characterized in that: The spatial attention module is used to compress and generate three feature matrices using three pooling methods in the tensor space of the channel dimension, then generate a unified pooling layer through the set fusion rule, and then obtain a two-dimensional spatial attention map through the softmax activation function; wherein the three pooling methods include maximum pooling, local saliency value pooling, and transfer attention value pooling; The pooling layer function corresponding to the transfer attention value pooling method is Among them, w i is the superposition output of the saliency weight matrix of the previous layer or the previous layers, σ w Indicates that W i The standard deviation on w Indicates that i The mean value on w Represents the pooling weight value obtained using the Gaussian function, Represents the weight matrix of the pooling layer after transferring attention and saliency.
7. The edge computing device for voice imprinting according to claim 6, characterized in that: The fusion rules set are specifically as follows: P=λ1P m +λ2P t +λ3P1 stλ1+λ2+λ3=1 Among them, P m is the pooling layer function corresponding to the maximum pooling method, P1 is the pooling layer function corresponding to the local saliency value pooling method, P is a unified pooling layer function, and λ1, α2, and α3 all represent feature weighted fusion normalization constraint parameters.
8. The edge computing device for voice imprinting according to claim 1, characterized in that: The workflow of the edge computing unit is as follows: Establish a speech synthesis model based on deep learning; Receive a short-term signal composed of multiple frames pre-processed by a voice acquisition unit; The short-term signal composed of multiple frames is sequentially input into the speech synthesis model based on deep learning; The key speech feature parameters are used to perform a convolution operation on the input short-term signal composed of multiple frames. The attention transfer model is added to generate new convolution parameter weights. The attention transfer model is applied to the subsequent convolution and pooling layer operations of the input speech signal, transferring the previous learning experience to the current speech recognition process to generate a fusion target feature representation. The output result of the last layer is used as the input feature of the SVM classifier to realize the imprinting of the speech signal to be imprinted.
9. The edge computing device for voice imprinting according to claim 1, characterized in that: The memory of the storage unit adopts 2G memory + 16G EMMC.
10. An edge computing method for voice imprinting, characterized in that: An edge computing device for voice imprinting as described in any one of claims 1 to 9, comprising: Collect the voice signal to be transcribed in real time and pre-process it; Extract features from the pre-processed speech signal to be transcribed to obtain key speech feature parameters; A speech synthesis model based on deep learning is used to take the extracted key speech feature parameters as input to generate a speech signal similar to the speech signal to be imitated; The memory is used to store the speech signal to be imitated and the generated similar speech signal.