Voice signal processing method and device, storage medium, electronic device, product
By downsampling the voice signal and performing self-attention calculations on the original and downsampling sequences, the problem of large amount of calculations in voice signal processing is solved, and more efficient computing efficiency is achieved.
Patent Information
- Application Number
- CN202210283378.1
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Filing Date
- 2022-03-22
- Publication Date
- 2025-05-09
- Estimated Expiration
- 2042-03-22
AI Technical Summary
In the prior art, the Transformer model has a large amount of calculation of the self-attention mechanism in speech signal processing, resulting in high requirements for system server performance.
By downsampling the voice signal to be processed, a reduced-length signal sequence is obtained, and self-attention calculation is performed on the original signal sequence and the downsampling sequence to reduce the size of the model processed data.
Reduces the complexity of signal processing, improves computing efficiency, and reduces the requirements for server performance.
Smart Images

Figure CN114550722B_ABST
Abstract
Description
Technical Field
[0001] The present disclosure relates to signal processing technology, and in particular to a speech signal processing method and device, a storage medium, an electronic device, and a product. Background Art
[0002] The implementation of speech recognition, voice wake-up and other applications in the existing technology usually uses the Transformer model to perform keyword recognition and response wake-up. The self-attention mechanism in the conventional Transformer model has a large amount of computation (each frame of speech signal needs to be correlated with all signals, and the computational complexity is O(T2), where T is the number of input speech signal frames), which places high requirements on the system server performance. Summary of the invention
[0003] In order to solve the above technical problems, the present disclosure is proposed. The embodiments of the present disclosure provide a speech signal processing method and device, a storage medium, and an electronic device.
[0004] According to one aspect of an embodiment of the present disclosure, there is provided a method for processing a speech signal, comprising:
[0005] Processing the speech signal to be processed to obtain a first signal sequence of n frames in length; wherein each frame in the first signal sequence is a signal vector of the same length; and n is an integer greater than 1;
[0006] Downsampling the first signal sequence to obtain a second signal sequence with a length of m frames; wherein m is an integer greater than or equal to 1, and n is greater than m;
[0007] Performing self-attention calculation on the first signal sequence and the second signal sequence to obtain a first matrix;
[0008] Based on the first matrix, a feature conversion result of the speech signal is determined.
[0009] Optionally, performing downsampling processing on the first signal sequence to obtain a second signal sequence having a length of m frames includes:
[0010] Divide the first signal sequence into m vector groups; wherein each of the vector groups includes at least one signal vector;
[0011] For each vector group in the m vector groups, down-sampling processing is performed on at least one signal vector included in the vector group to obtain a down-sampled vector;
[0012] Based on the m down-sampling vectors corresponding to the m vector groups, a second signal sequence with a length of m frames is obtained.
[0013] Optionally, dividing the first signal sequence into m vector groups includes:
[0014] The n frames of signal vectors included in the first signal sequence are equally divided into m equal parts to obtain the m vector groups.
[0015] Optionally, dividing the first signal sequence into m vector groups includes:
[0016] Converting the first signal sequence into an n-dimensional segmentation vector; wherein the segmentation vector is composed of 0 and / or 1;
[0017] Determine m groups of target positions based on a position of at least one 1 segmented by 0 in the segmentation vector;
[0018] For each group of target signal positions in the m groups of target positions, a signal vector of a corresponding position in the first signal sequence is determined based on the target position to obtain one of the vector groups.
[0019] Optionally, converting the first signal sequence into an n-dimensional segmentation vector comprises:
[0020] Converting the n signal vectors included in the first signal sequence into numerical representations to obtain an n-dimensional intermediate vector including n numerical values;
[0021] Determine the magnitude relationship between the value of each numerical value in the n-dimensional intermediate vector and a set threshold value;
[0022] The numerical values in the n-dimensional intermediate vector that are greater than or equal to the set threshold are converted to 1, and the numerical values that are less than the set threshold are converted to 0; and the segmentation vector is obtained.
[0023] Optionally, performing self-attention calculation on the first signal sequence and the second signal sequence to obtain a first matrix includes:
[0024] For each of the n signal vectors included in the first signal sequence, perform self-attention calculation on the signal vector and the m down-sampling vectors included in the second signal sequence to obtain a calculation result;
[0025] The first matrix is obtained based on the n calculation results.
[0026] Optionally, before determining the feature conversion result of the speech signal based on the first matrix, the method further includes:
[0027] For each signal vector of the n signal vectors included in the first signal sequence, perform self-attention calculation on the signal vector and a set number of signal vectors adjacent to the signal vector to obtain a second matrix corresponding to the first signal sequence;
[0028] The determining, based on the first matrix, a feature conversion result of the speech signal includes:
[0029] Based on the first matrix and the second matrix, a feature conversion result of the speech signal is determined.
[0030] Optionally, determining a feature conversion result of the speech signal based on the first matrix and the second matrix includes:
[0031] Performing matrix addition on the first matrix and the second matrix to obtain a superposition matrix;
[0032] Perform at least one encoding process and at least one decoding process on the superposition matrix to obtain a feature conversion result of the speech signal.
[0033] Optionally, it also includes:
[0034] Performing feature extraction and feature conversion on the feature conversion result to obtain a wake-up probability value;
[0035] Based on the magnitude relationship between the wake-up probability value and the wake-up threshold, it is determined whether to wake up the preset device.
[0036] Optionally, it also includes:
[0037] Based on the feature conversion result, a first word sequence is obtained by matching with a set word table; wherein the first word sequence includes a plurality of word vectors;
[0038] Based on the first word sequence, a text recognition result corresponding to the speech signal is determined.
[0039] Optionally, determining a text recognition result corresponding to the speech signal based on the first word sequence includes:
[0040] Downsampling the first word sequence to obtain a second word sequence with a reduced length;
[0041] Performing self-attention calculation on the first word sequence and the second word sequence to obtain an intermediate representation result;
[0042] Based on the intermediate representation result, a text recognition result of the speech signal is determined.
[0043] According to another aspect of an embodiment of the present disclosure, there is provided a speech signal processing apparatus, including:
[0044] A signal processing module, used for processing the speech signal to be processed to obtain a first signal sequence of n frames in length; wherein each frame in the first signal sequence is a signal vector of the same length; and n is an integer greater than 1;
[0045] A downsampling module, configured to perform downsampling processing on the first signal sequence to obtain a second signal sequence of m frames in length; wherein m is an integer greater than or equal to 1, and n is greater than m;
[0046] A self-attention module, configured to perform self-attention calculation on the first signal sequence and the second signal sequence to obtain a first matrix;
[0047] A feature conversion module is used to determine a feature conversion result of the speech signal based on the first matrix.
[0048] Optionally, the downsampling module includes:
[0049] A sequence segmentation unit, configured to segment the first signal sequence into m vector groups; wherein each of the vector groups includes at least one signal vector;
[0050] A signal sampling unit, configured to perform downsampling processing on at least one signal vector included in each of the m vector groups to obtain a downsampled vector;
[0051] The second signal sequence unit is used to obtain a second signal sequence with a length of m frames based on the m down-sampling vectors corresponding to the m vector groups.
[0052] Optionally, the sequence segmentation unit is specifically configured to evenly segment n frames of signal vectors included in the first signal sequence into m equal parts to obtain the m vector groups.
[0053] Optionally, the sequence segmentation unit is specifically used to convert the first signal sequence into an n-dimensional segmentation vector; wherein the segmentation vector is composed of 0 and / or 1; based on the position of at least one 1 divided by 0 in the segmentation vector, m groups of target positions are determined; for each group of target signal positions in the m groups of target positions, a signal vector of a corresponding position in the first signal sequence is determined based on the target position, to obtain a vector group.
[0054] Optionally, when converting the first signal sequence into an n-dimensional segmentation vector, the sequence segmentation unit is used to convert the n signal vectors included in the first signal sequence into numerical representations to obtain an n-dimensional intermediate vector including n numerical values; determine the size relationship between the value of each numerical value in the n-dimensional intermediate vector and a set threshold; convert the numerical values in the n-dimensional intermediate vector that are greater than or equal to the set threshold to 1, and convert the numerical values that are less than the set threshold to 0; and obtain the segmentation vector.
[0055] Optionally, the self-attention module is specifically used to perform self-attention calculation on each of the n signal vectors included in the first signal sequence and the m down-sampling vectors included in the second signal sequence to obtain a calculation result; and obtain the first matrix based on the n calculation results.
[0056] Optionally, the device further comprises:
[0057] a local self-attention module, configured to perform self-attention calculation on each of the n signal vectors included in the first signal sequence and a set number of signal vectors adjacent to the signal vector, so as to obtain a second matrix corresponding to the first signal sequence;
[0058] The feature conversion module is specifically used to determine the feature conversion result of the speech signal based on the first matrix and the second matrix.
[0059] Optionally, the feature conversion module is specifically used to perform matrix addition on the first matrix and the second matrix to obtain a superposition matrix; perform at least one encoding process and at least one decoding process on the superposition matrix to obtain a feature conversion result of the speech signal.
[0060] Optionally, the device further comprises:
[0061] The voice wake-up module is used to extract and convert the features of the feature conversion result to obtain a wake-up probability value; based on the size relationship between the wake-up probability value and the wake-up threshold, determine whether to wake up the preset device.
[0062] Optionally, the device further comprises:
[0063] A vocabulary matching module, used for matching the feature conversion result with a set vocabulary to obtain a first word sequence; wherein the first word sequence includes a plurality of word vectors;
[0064] The speech recognition module is used to determine a text recognition result corresponding to the speech signal based on the first word sequence.
[0065] Optionally, the speech recognition module is specifically used to downsample the first word sequence to obtain a second word sequence with a reduced length; perform self-attention calculation on the first word sequence and the second word sequence to obtain an intermediate representation result; and determine the text recognition result of the speech signal based on the intermediate representation result.
[0066] According to another aspect of the embodiments of the present disclosure, a computer-readable storage medium is provided, on which computer program instructions are stored. When the computer program instructions are executed by a processor, the speech signal processing method described in any of the above embodiments is implemented.
[0067] According to another aspect of the embodiments of the present disclosure, an electronic device is provided, the electronic device comprising:
[0068] A memory for storing a computer program product;
[0069] The processor is used to execute the computer program product stored in the memory, and when the computer program product is executed, the speech signal processing method described in any one of the above embodiments is implemented.
[0070] According to another aspect of the embodiments of the present disclosure, a computer program product is provided, including computer program instructions, which, when executed by a processor, implement the speech signal processing method described in any of the above embodiments.
[0071] Based on a speech signal processing method and device, storage medium, electronic device, and product provided by the above embodiments of the present disclosure, the speech signal to be processed is processed to obtain a first signal sequence of n frames in length; wherein each frame in the first signal sequence is a signal vector of the same length; and n is an integer greater than 1; the first signal sequence is downsampled to obtain a second signal sequence of m frames in length; wherein m is an integer greater than or equal to 1, and n is greater than m; self-attention calculation is performed on the first signal sequence and the second signal sequence to obtain a first matrix; based on the first matrix, a feature conversion result of the speech signal is determined; this embodiment reduces the size of model processing data through downsampling processing, greatly reduces the complexity of signal processing, and thus improves the computational efficiency of signal processing.
[0072] The technical solution of the present disclosure is further described in detail below through the accompanying drawings and embodiments. BRIEF DESCRIPTION OF THE DRAWINGS
[0073] The above and other purposes, features and advantages of the present disclosure will become more apparent by describing the embodiments of the present disclosure in more detail in conjunction with the accompanying drawings. The accompanying drawings are used to provide a further understanding of the embodiments of the present disclosure and constitute a part of the specification. Together with the embodiments of the present disclosure, they are used to explain the present disclosure and do not constitute a limitation of the present disclosure. In the drawings, the same reference numerals generally represent the same components or steps.
[0074] Figure 1 It is a flowchart of a speech signal processing method provided by an exemplary embodiment of the present disclosure.
[0075] Figure 2 This disclosure Figure 1 A schematic flow chart of step 104 in the illustrated embodiment.
[0076] Figure 3 This disclosure Figure 2 A flowchart diagram of step 1041 in the embodiment shown.
[0077] Figure 4 It is a structural diagram of a speech signal processing device provided by an exemplary embodiment of the present disclosure.
[0078] Figure 5 is a structural diagram of an electronic device provided by an exemplary embodiment of the present disclosure. DETAILED DESCRIPTION
[0079] Below, the exemplary embodiments according to the present disclosure will be described in detail with reference to the accompanying drawings. Obviously, the described embodiments are only a part of the embodiments of the present disclosure, rather than all the embodiments of the present disclosure, and it should be understood that the present disclosure is not limited to the exemplary embodiments described here.
[0080] It should be noted that the relative arrangement of components and steps, the numerical expressions and numerical values set forth in these embodiments do not limit the scope of the present disclosure unless specifically stated otherwise.
[0081] Those skilled in the art can understand that the terms "first" and "second" in the embodiments of the present disclosure are only used to distinguish different steps, devices or modules, etc., and neither represent any specific technical meaning nor indicate the necessary logical order between them.
[0082] It should also be understood that in the embodiments of the present disclosure, “plurality” may refer to two or more than two, and “at least one” may refer to one, two, or more than two.
[0083] It should also be understood that any component, data or structure mentioned in the embodiments of the present disclosure can generally be understood as one or more, unless explicitly limited or otherwise indicated in the context.
[0084] In addition, the term "and / or" in this disclosure is only a description of the association relationship of associated objects, indicating that there may be three relationships. For example, A and / or B can represent: A exists alone, A and B exist at the same time, and B exists alone. In addition, the character " / " in this disclosure generally indicates that the previous and next associated objects are in an "or" relationship. The data referred to in this disclosure may include unstructured data such as text, images, and videos, and may also be structured data.
[0085] It should also be understood that the description of the various embodiments in the present disclosure focuses on the differences between the various embodiments, and the same or similar aspects thereof can be referenced to each other, and for the sake of brevity, they will not be described one by one.
[0086] At the same time, it should be understood that for the convenience of description, the sizes of the various parts shown in the drawings are not drawn according to the actual proportional relationship.
[0087] The following description of at least one exemplary embodiment is merely illustrative in nature and is in no way intended to limit the present disclosure, its application, or uses.
[0088] Technologies, methods, and equipment known to ordinary technicians in the relevant art may not be discussed in detail, but where appropriate, the technologies, methods, and equipment should be considered part of the specification.
[0089] It should be noted that like reference numerals and letters refer to similar items in the following figures, and therefore, once an item is defined in one figure, it need not be further discussed in subsequent figures.
[0090] The disclosed embodiments can be applied to electronic devices such as terminal devices, computer systems, servers, etc., which can operate with many other general or special computing system environments or configurations. Examples of well-known terminal devices, computing systems, environments and / or configurations suitable for use with electronic devices such as terminal devices, computer systems, servers, etc. include, but are not limited to: personal computer systems, server computer systems, thin clients, thick clients, handheld or laptop devices, microprocessor-based systems, set-top boxes, programmable consumer electronics, network personal computers, small computer systems, large computer systems, and distributed cloud computing technology environments including any of the above systems, etc.
[0091] Electronic devices such as terminal devices, computer systems, servers, etc. can be described in the general context of computer system executable instructions (such as program modules) executed by computer systems. Generally, program modules can include routines, programs, object programs, components, logic, data structures, etc., which perform specific tasks or implement specific abstract data types. Computer systems / servers can be implemented in a distributed cloud computing environment, where tasks are performed by remote processing devices linked through a communication network. In a distributed cloud computing environment, program modules can be located on local or remote computing system storage media including storage devices.
[0092] Exemplary Methods
[0093] Figure 1 is a flow chart of a method for processing a speech signal provided by an exemplary embodiment of the present disclosure. This embodiment can be applied to electronic devices, such as Figure 1 As shown, the following steps are included:
[0094] Step 102: Process the speech signal to be processed to obtain a first signal sequence with a length of n frames.
[0095] Each frame in the first signal sequence is a signal vector with the same length; n is an integer greater than 1.
[0096] In this embodiment, any received voice signal can be processed to obtain a first signal sequence, which includes signal vectors with consistent n frame lengths. Optionally, the signal processing method can be the existing technology, and it is only necessary to process the voice signal into a matrix without limiting the specific processing method.
[0097] Step 104: downsample the first signal sequence to obtain a second signal sequence with a length of m frames.
[0098] Here, m is an integer greater than or equal to 1, and n is greater than m.
[0099] Optionally, downsampling of the first signal sequence can be achieved through a network layer of a deep neural network, or through a function, and each signal vector in the second signal sequence obtained after processing can be one of the signal vectors in the first signal sequence, or obtained by calculation using multiple signal vectors in the first signal sequence.
[0100] Step 106: Perform self-attention calculation on the first signal sequence and the second signal sequence to obtain a first matrix.
[0101] In the prior art, a Transformer structure is usually used to process speech signals, and the association between multiple signal vectors is realized through self-attention calculation in the encoder. In this embodiment, a second signal sequence is obtained by downsampling, and self-attention calculation is performed on the first signal sequence and the second signal sequence. By changing the structure of the initial encoder in the Transformer structure, the amount of information processing is greatly reduced, and the processing speed is improved.
[0102] Step 108: Determine a feature conversion result of the speech signal based on the first matrix.
[0103] Optionally, the first matrix can be processed through subsequent processing in a general Transformer structure to obtain a feature conversion result expressed in matrix form, and further speech tasks such as speech wake-up and speech recognition can be performed based on the feature conversion result.
[0104] The above-mentioned embodiment of the present disclosure provides a speech signal processing method, which processes the speech signal to be processed to obtain a first signal sequence with a length of n frames; wherein each frame in the first signal sequence is a signal vector with the same length; the n is an integer greater than 1; the first signal sequence is downsampled to obtain a second signal sequence with a length of m frames; wherein m is an integer greater than or equal to 1, and the n is greater than the m; self-attention calculation is performed on the first signal sequence and the second signal sequence to obtain a first matrix; based on the first matrix, the feature conversion result of the speech signal is determined; this embodiment reduces the size of the model processing data through downsampling processing, greatly reduces the complexity of signal processing, and thus improves the computational efficiency of signal processing.
[0105] like Figure 2 As shown in the above Figure 1 Based on the illustrated embodiment, step 104 may include the following steps:
[0106] Step 1041: divide the first signal sequence into m vector groups.
[0107] Each vector group includes at least one signal vector.
[0108] In this embodiment, the first signal sequence may be divided into m vector groups by equal division, unequal division, or dynamic division. This embodiment does not limit the division method, and it is only necessary to divide n frames of signal vectors into m groups.
[0109] Step 1042: for each vector group in the m vector groups, down-sample at least one signal vector included in the vector group to obtain a down-sampled vector.
[0110] Step 1043: Based on the m down-sampling vectors corresponding to the m vector groups, a second signal sequence with a length of m frames is obtained.
[0111] This embodiment reduces the length of the first signal sequence by processing at least one signal vector included in each vector group into a down-sampling vector. The optional down-sampling methods may include but are not limited to: a. mean method, that is, averaging at least one signal vector to obtain a down-sampling vector; b. maximum method, that is, taking the maximum value in at least one signal vector as the down-sampling vector; c. function conversion method, that is, inputting at least one signal vector into a function and outputting 1 frame. The function may include but is not limited to: learnable parameters, selectable linear functions, neural network structures and other forms; the above are only three examples of down-sampling in this embodiment, and are not used to limit the down-sampling method in this embodiment. The existing technology and subsequent methods that can achieve down-sampling can be applied to the present embodiment to achieve down-sampling processing.
[0112] Optionally, based on the above embodiment, step 1041 may include:
[0113] The n frames of signal vectors included in the first signal sequence are evenly divided into m equal parts to obtain m vector groups.
[0114] The present embodiment is a feasible way to obtain m vector groups, by evenly dividing n frames of signal vectors to obtain m vector groups. Of course, there is a situation where n cannot be divided by m. In this case, the divisibility can be achieved by padding the n frames of signal vectors with zeros, or at least one signal vector that cannot be divided by the last one is used as a vector group, with the final m vector groups being the standard; and, since m is not a fixed value, m can be adjusted according to the value of n so that n can be divided by m. The present embodiment implements downsampling with a fixed window length. This sampling method is simple and fast, but the window length of each processing unit is fixed, while the duration of each phoneme in the speech pronunciation is not fixed. If a fixed window is used to evenly divide the speech signal, it is easy to cause the information in a single window to contain information of more than one phoneme, which is not suitable for downsampling. Based on this, the present application proposes the following Figure 3 The illustrated embodiment solves the above-mentioned problem by using a dynamic window.
[0115] like Figure 3 As shown in the above Figure 2 Based on the illustrated embodiment, step 1041 may include the following steps:
[0116] Step 301: convert a first signal sequence into an n-dimensional segmentation vector.
[0117] The segmentation vector is composed of 0 and / or 1.
[0118] In this embodiment, by converting each signal vector in the first signal sequence into a 0 or 1 representation, an n-dimensional segmentation vector is obtained, and each dimension in the segmentation vector corresponds to a signal vector in the first signal sequence according to its position.
[0119] Step 302: determine m groups of target positions based on the position of at least one 1 segmented by 0 in the segmentation vector.
[0120] Optionally, m groups of target positions are determined based on positions corresponding to at least one continuous 1 in the segmentation vector, that is, positions with continuous values of 1 in the segmentation vector are divided into a window, and positions with values of 0 can be directly ignored, thereby realizing dynamic window segmentation.
[0121] Step 303: for each group of target signal positions in the m groups of target positions, determine the signal vector of the corresponding position in the first signal sequence based on the target position to obtain a vector group.
[0122] In this embodiment, since the position corresponding to each value (0 or 1) in the segmentation vector corresponds one-to-one to the position of the signal vector in the first signal sequence, when at least one dynamic window that is continuously 1 is determined, at least one vector group can be obtained, and the signal vector in the first signal sequence corresponding to 0 in the segmentation vector can be deleted when determining the vector group, that is, the position corresponding to 0 in the segmentation vector in the first signal sequence does not participate in the composition of the vector group. Of course, there are special cases. When the segmentation vector is all 0 or all 1, the first signal sequence can be downsampled by means of a fixed window.
[0123] Optionally, based on the above embodiment, step 301 may include:
[0124] a. Convert the n signal vectors included in the first signal sequence into numerical representations to obtain an n-dimensional intermediate vector including n numerical values.
[0125] Optionally, the numerical conversion can be achieved by the following formula (1):
[0126] G = softmax(f(H)) Formula (1)
[0127] Wherein, H represents the first signal sequence with a length of n, G represents the intermediate vector with the same length as H and in the form of [g1, g2, …, gn], each element has a value between [0, 1], and f is the parameter function to be learned (for example: conventional neural network, linear transformation, etc.).
[0128] b. Determine the relationship between the value of each numerical value in the n-dimensional intermediate vector and the set threshold.
[0129] c. Convert the values in the n-dimensional intermediate vector that are greater than or equal to the set threshold to 1, and convert the values that are less than the set threshold to 0; and obtain the segmentation vector.
[0130] In this embodiment, the segmentation vector can be represented as T, which is in the form of [t1, t2, …, tn], and each element takes a value of 0 or 1. The process of obtaining the segmentation vector is determined based on the set threshold and the intermediate vector. When an element in the intermediate vector is greater than or equal to the set threshold, the value of the position in the segmentation vector is set to 1. When an element in the intermediate vector is less than the set threshold, the value of the position in the segmentation vector is set to 0; it can be expressed as: ti=1 if gi>threshold, else (other) 0, where threshold represents the set threshold, which can be set according to the actual application scenario, and the size of the set threshold can be adjusted to avoid the situation where all values in the segmentation vector are 0 or all values are 1. The positions in T where the continuous value is 1 are divided into a window, and the data at the corresponding positions in H are downsampled (mean method, maximum method, function conversion method, etc.). For example, in an optional example: if T = [0000111100111000111100000], 3 windows can be obtained, then the data at the corresponding positions of H are downsampled to obtain 3 frames, and the data corresponding to the value 0 in T is directly discarded.
[0131] Optionally, based on the above embodiment, step 106 may include:
[0132] For each of the n signal vectors included in the first signal sequence, a self-attention calculation is performed on the signal vector and the m down-sampled vectors included in the second signal sequence to obtain a calculation result.
[0133] A first matrix is obtained based on the n calculation results.
[0134] The self-attention mechanism in the prior art is to perform self-attention calculation on each signal vector among n signal vectors and other n-1 signal vectors, and the calculation complexity is: O(n 2 ); In this embodiment, by adding downsampling processing, self-attention calculation is performed on the first signal sequence H and the downsampled second signal sequence H′, and its calculation complexity is O(mn), which is m / n of the original self-attention calculation complexity. The self-attention calculation process is as follows:
[0135] Q=HW Q
[0136] K=H′W K
[0137] V=H′W V
[0138]
[0139] G=AV Formula (2)
[0140] Where H represents the first signal sequence, H′ represents the second signal sequence, and W Q , W K , W V represents the parameters to be learned (weight matrix), in matrix form, Q, K and V represent the query matrix, key matrix and value matrix respectively; A represents the intermediate calculation result, d K Represents the length of K. The output A of the Softmax function is the learned self-attention value, and G is the self-attention result output.
[0141] Due to the above downsampling process, a lot of information is inevitably reduced in the second signal sequence relative to the first signal sequence. This embodiment compensates for the information loss that may be caused by downsampling through local self-attention. Optionally, before step 108, the following may also be included:
[0142] For each signal vector of the n signal vectors included in the first signal sequence, perform self-attention calculation on the signal vector and a set number of signal vectors adjacent to the signal vector to obtain a second matrix corresponding to the first signal sequence;
[0143] At this time, step 108 may include: determining a feature conversion result of the speech signal based on the first matrix and the second matrix.
[0144] This embodiment implements a local self-attention mechanism by performing self-attention calculation on each signal vector and a set number of adjacent signal vectors. That is, for each signal vector in the first signal sequence, self-attention calculation is performed on it and its k-frame context, where k is much smaller than n. The calculation method can refer to the self-attention calculation process for determining the first matrix, and the specific formula includes:
[0145] Q′=HW Q
[0146] K′=H″W K
[0147] V′=H″W V
[0148]
[0149] G′=A′V Formula (3)
[0150] Where H represents the first signal sequence, H″ represents the signal sequence consisting of k frames of context, and W Q , W K , W Vrepresents the parameters to be learned (weight matrix), in matrix form, Q′, K′ and V′ represent the query matrix, key matrix and value matrix respectively; A′ represents the intermediate calculation result, d K Represents the length of K′. The output A′ of the Softmax function is the learned self-attention value, and G′ is the local self-attention result output.
[0151] Optionally, the obtained second sentence is accumulated with the first matrix to form the final output of the downsampled self-attention. The computational complexity added by this embodiment is only O(nk). In this embodiment, the total computational complexity obtained by combining the above computational complexity is O(n(k+m)), which is much smaller than the original computational complexity O(n 2 ); This embodiment enhances the self-attention relationship in the local range of the input sequence through the local self-attention mechanism, while slightly increasing the amount of calculation and compensating for the information loss that may be caused by downsampling.
[0152] Optionally, determining a feature conversion result of the speech signal based on the first matrix and the second matrix includes:
[0153] Performing matrix addition on the first matrix and the second matrix to obtain a superposition matrix;
[0154] Perform at least one encoding process and at least one decoding process on the superposition matrix to obtain a feature conversion result of the speech signal.
[0155] In this embodiment, the first matrix and the second matrix are of the same size, and the superposition of the two is achieved by adding each element in the matrix to obtain a superposition matrix; subsequent processing based on a general Transformer structure can obtain the feature conversion result of the speech signal. This embodiment combines downsampling processing and local self-attention mechanism to improve the general Transformer structure. The improved Transformer structure significantly reduces the amount of calculation without reducing the calculation accuracy, thereby improving the processing efficiency of the speech signal.
[0156] In some optional embodiments, the method provided in this embodiment further includes:
[0157] Perform feature extraction and feature conversion on the feature conversion result to obtain the wake-up probability value;
[0158] Based on the magnitude relationship between the wake-up probability value and the wake-up threshold, it is determined whether to wake up the preset device.
[0159] This embodiment is an application scenario of the feature conversion result obtained by the above-mentioned speech signal processing method, speech wake-up, and the obtained feature conversion result is processed using a fully connected neural network as a feature extractor, and the processing result of the fully connected neural network (expressed as a numerical value) is normalized (for example, a numerical value between 0 and 1 is obtained by Sigmoid function processing), and the numerical value of the normalized result is used as the wake-up probability value. When the wake-up probability value is greater than or equal to the wake-up threshold, it is determined to wake up the preset device; otherwise, the preset device is not woken up. Different wake-up thresholds can be set for different scenarios and different preset devices. Therefore, this embodiment can be applicable to a variety of speech wake-up scenarios.
[0160] In some optional embodiments, the method provided in this embodiment further includes:
[0161] Based on the feature conversion result, the feature conversion result is matched with the set word list to obtain a first word sequence, wherein the first word sequence includes multiple word vectors.
[0162] Optionally, the vocabulary is set to include a large number of words (or characters) that may be used, and each saved word includes a corresponding word vector. The matching in this embodiment can be achieved by decomposing the matrix of the feature conversion result into multiple vector expressions, and matching the vector with the word vector in the set vocabulary to obtain multiple word vectors to form a first word sequence.
[0163] Based on the first word sequence, a text recognition result corresponding to the speech signal is determined.
[0164] The input speech signal is acoustically processed to obtain acoustic features; acoustic encoding is performed using the acoustic features and acoustic models to obtain a syllable sequence; vocabulary matching is performed based on the syllable sequence and the vocabulary to obtain a word sequence; the language decoding result is output based on the word sequence and the language model, thus completing the speech recognition process.
[0165] The computational complexity of the improved Transformer structure is significantly reduced, and it can be used for signal processing in acoustic models and language models to improve the response speed of application devices to voice signals, thereby achieving faster voice recognition tasks.
[0166] Optionally, determining a text recognition result corresponding to the speech signal based on the first word sequence includes:
[0167] Downsampling the first word sequence to obtain a second word sequence with reduced length;
[0168] Perform self-attention calculation on the first word sequence and the second word sequence to obtain an intermediate representation result;
[0169] Based on the intermediate representation result, the text recognition result of the speech signal is determined.
[0170] In this embodiment, the first word sequence is based on Figure 1 The speech signal processing method provided is processed by a method similar to that provided, firstly, the first word sequence is down-sampled by any achievable method in the above embodiments to obtain a second word sequence with reduced length, and the second sequence is processed by using other parts in the Transformer structure to obtain an intermediate representation result. Optionally, the text recognition result of the speech signal is determined by the intermediate representation result; or, the first word sequence is processed based on the local self-attention mechanism disclosed in the above embodiments to obtain a second representation result, and the text recognition result is determined based on the result of adding the intermediate representation result and the second representation result matrix; this embodiment processes the first signal sequence and the first word sequence respectively based on the latter Transformer structure, which greatly reduces the computational complexity, reduces the consumption of computing resources, and improves the efficiency of speech recognition, so that the method provided in this embodiment can be applied to more hardware devices with smaller computing space.
[0171] The speech signal processing method provided in the present application can also be applied to other speech processing applications. The above embodiments are only for ease of understanding, and provide two speech processing applications, speech wake-up and speech recognition, and are not intended to limit the scope of application of the method provided in the present application.
[0172] Any speech signal processing method provided in the embodiments of the present disclosure may be executed by any appropriate device with data processing capabilities, including but not limited to: a terminal device and a server, etc. Alternatively, any speech signal processing method provided in the embodiments of the present disclosure may be executed by a processor, such as the processor executing any speech signal processing method mentioned in the embodiments of the present disclosure by calling corresponding instructions stored in a memory. This will not be described in detail below.
[0173] Exemplary Devices
[0174] Figure 4 FIG. 1 is a schematic diagram of the structure of a speech signal processing device provided by an exemplary embodiment of the present disclosure. Figure 4 As shown, the device provided in this embodiment includes:
[0175] The signal processing module 41 is used to process the speech signal to be processed to obtain a first signal sequence with a length of n frames.
[0176] Each frame in the first signal sequence is a signal vector with the same length; n is an integer greater than 1.
[0177] The down-sampling module 42 is used to perform down-sampling processing on the first signal sequence to obtain a second signal sequence with a length of m frames.
[0178] Here, m is an integer greater than or equal to 1, and n is greater than m.
[0179] The self-attention module 43 is used to perform self-attention calculation on the first signal sequence and the second signal sequence to obtain a first matrix.
[0180] The feature conversion module 44 is used to determine a feature conversion result of the speech signal based on the first matrix.
[0181] The above-mentioned embodiment of the present disclosure provides a speech signal processing device, which processes the speech signal to be processed to obtain a first signal sequence of n frames in length; wherein each frame in the first signal sequence is a signal vector of the same length; and n is an integer greater than 1; down-sampling is performed on the first signal sequence to obtain a second signal sequence of m frames in length; wherein m is an integer greater than or equal to 1, and n is greater than m; self-attention calculation is performed on the first signal sequence and the second signal sequence to obtain a first matrix; based on the first matrix, a feature conversion result of the speech signal is determined; this embodiment reduces the size of the model processing data through down-sampling, greatly reduces the complexity of signal processing, and thus improves the computational efficiency of signal processing.
[0182] In some optional embodiments, the downsampling module 42 includes:
[0183] A sequence segmentation unit, used to segment the first signal sequence into m vector groups; wherein each vector group includes at least one signal vector;
[0184] A signal sampling unit, configured to perform down-sampling processing on at least one signal vector included in each of the m vector groups to obtain a down-sampled vector;
[0185] The second signal sequence unit is used to obtain a second signal sequence with a length of m frames based on the m down-sampling vectors corresponding to the m vector groups.
[0186] In some optional embodiments, the sequence segmentation unit is specifically configured to segment n frames of signal vectors included in the first signal sequence into m equal parts to obtain m vector groups.
[0187] In some other optional embodiments, the sequence segmentation unit is specifically used to convert the first signal sequence into an n-dimensional segmentation vector; wherein the segmentation vector is composed of 0 and / or 1; based on the position of at least one 1 divided by 0 in the segmentation vector, m groups of target positions are determined; for each group of target signal positions in the m groups of target positions, a signal vector of the corresponding position in the first signal sequence is determined based on the target position to obtain a vector group.
[0188] Optionally, when converting the first signal sequence into an n-dimensional segmentation vector, the sequence segmentation unit is used to convert the n signal vectors included in the first signal sequence into numerical representations to obtain an n-dimensional intermediate vector including n numerical values; determine the size relationship between the value of each numerical value in the n-dimensional intermediate vector and a set threshold; convert the numerical values in the n-dimensional intermediate vector that are greater than or equal to the set threshold to 1, and convert the numerical values that are less than the set threshold to 0; and obtain the segmentation vector.
[0189] In some optional embodiments, the self-attention module 43 is specifically used to perform self-attention calculation on each signal vector of the n signal vectors included in the first signal sequence and the m down-sampling vectors included in the second signal sequence to obtain a calculation result; and obtain a first matrix based on the n calculation results.
[0190] In some optional embodiments, the device provided by this embodiment further includes:
[0191] A local self-attention module, configured to perform self-attention calculation on each of the n signal vectors included in the first signal sequence and a set number of signal vectors adjacent to the signal vector to obtain a second matrix corresponding to the first signal sequence;
[0192] The feature conversion module is specifically used to determine the feature conversion result of the speech signal based on the first matrix and the second matrix.
[0193] Optionally, the feature conversion module is specifically used to perform matrix addition on the first matrix and the second matrix to obtain a superposition matrix; perform at least one encoding process and at least one decoding process on the superposition matrix to obtain a feature conversion result of the speech signal.
[0194] In some optional embodiments, the device provided by this embodiment further includes:
[0195] The voice wake-up module is used to extract and transform the feature conversion results to obtain a wake-up probability value; based on the relationship between the wake-up probability value and the wake-up threshold, determine whether to wake up the preset device.
[0196] In some other optional embodiments, the device provided in this embodiment further includes:
[0197] A vocabulary matching module, used for matching the feature conversion result with a set vocabulary to obtain a first word sequence; wherein the first word sequence includes a plurality of word vectors;
[0198] The speech recognition module is used to determine a text recognition result corresponding to the speech signal based on the first word sequence.
[0199] Optionally, the speech recognition module is specifically used to downsample the first word sequence to obtain a second word sequence with a reduced length; perform self-attention calculation on the first word sequence and the second word sequence to obtain an intermediate representation result; and determine the text recognition result of the speech signal based on the intermediate representation result.
[0200] Exemplary Electronic Devices
[0201] Below, reference Figure 5 The electronic device according to the embodiment of the present disclosure is described. The electronic device may be any one or both of the first device 100 and the second device 200, or a stand-alone device independent of them, and the stand-alone device may communicate with the first device and the second device to receive the collected input signals from them.
[0202] Figure 5 A block diagram of an electronic device according to an embodiment of the present disclosure is illustrated.
[0203] like Figure 5 As shown, the electronic device 50 includes one or more processors 51 and a memory 52 .
[0204] The processor 51 may be a central processing unit (CPU) or other forms of processing units having data processing capabilities and / or instruction execution capabilities, and may control other components in the electronic device 50 to perform desired functions.
[0205] The memory 52 may store one or more computer program products, and the memory 52 may include various forms of computer-readable storage media, such as volatile memory and / or non-volatile memory. The volatile memory may include, for example, a random access memory (RAM) and / or a cache memory (cache), etc. The non-volatile memory may include, for example, a read-only memory (ROM), a hard disk, a flash memory, etc. One or more computer program products may be stored on the computer-readable storage medium, and the processor 51 may run the computer program products to implement the speech signal processing methods of the various embodiments of the present disclosure described above and / or other desired functions. Various contents such as input signals, signal components, noise components, etc. may also be stored in the computer-readable storage medium.
[0206] In one example, the electronic device 50 may further include: an input device 53 and an output device 54, and these components are interconnected via a bus system and / or other forms of connection mechanisms (not shown).
[0207] For example, when the electronic device is the first device 100 or the second device 200, the input device 53 may be the microphone or microphone array described above, for capturing input signals from a sound source. When the electronic device is a stand-alone device, the input device 53 may be a communication network connector, for receiving collected input signals from the first device 100 and the second device 200.
[0208] In addition, the input device 53 may also include, for example, a keyboard, a mouse, etc.
[0209] The output device 54 can output various information to the outside, including the determined distance information, direction information, etc. The output device 54 can include, for example, a display, a speaker, a printer, a communication network and a remote output device connected thereto, and the like.
[0210] Of course, to simplify, Figure 5 Only some of the components related to the present disclosure in the electronic device 50 are shown, and components such as a bus, an input / output interface, etc. are omitted. In addition, according to specific application situations, the electronic device 50 may also include any other appropriate components.
[0211] In addition to the above-mentioned methods and devices, an embodiment of the present disclosure may also be a computer program product, which includes computer program instructions, which, when executed by a processor, enable the processor to execute the steps of the speech signal processing method according to various embodiments of the present disclosure described in the above-mentioned "Exemplary Method" section of this specification.
[0212] The computer program product may be written in any combination of one or more programming languages to write program code for performing the operations of the disclosed embodiments, including object-oriented programming languages such as Java, C++, etc., and conventional procedural programming languages such as "C" or similar programming languages. The program code may be executed entirely on the user computing device, partially on the user device, as a separate software package, partially on the user computing device and partially on a remote computing device, or entirely on a remote computing device or server.
[0213] In addition, an embodiment of the present disclosure may also be a computer-readable storage medium on which computer program instructions are stored. When the computer program instructions are executed by a processor, the processor executes the steps of the speech signal processing method according to various embodiments of the present disclosure described in the above "Exemplary Method" section of this specification.
[0214] The computer readable storage medium can adopt any combination of one or more readable media. The readable medium can be a readable signal medium or a readable storage medium. The readable storage medium can include, for example, but is not limited to, a system, device or device of electricity, magnetism, light, electromagnetic, infrared, or semiconductor, or any combination of the above. More specific examples (non-exhaustive list) of readable storage media include: an electrical connection with one or more wires, a portable disk, a hard disk, a random access memory (RAM), a read-only memory (ROM), an erasable programmable read-only memory (EPROM or flash memory), an optical fiber, a portable compact disk read-only memory (CD-ROM), an optical storage device, a magnetic storage device, or any suitable combination of the above.
[0215] The basic principles of the present disclosure are described above in conjunction with specific embodiments. However, it should be noted that the advantages, strengths, effects, etc. mentioned in the present disclosure are only examples and not limitations, and it cannot be considered that these advantages, strengths, effects, etc. are required by each embodiment of the present disclosure. In addition, the specific details disclosed above are only for the purpose of illustration and ease of understanding, and are not limiting. The above details do not limit the present disclosure to the necessity of adopting the above specific details to be implemented.
[0216] Each embodiment in this specification is described in a progressive manner, and each embodiment focuses on the differences from other embodiments. The same or similar parts between the embodiments can be referred to each other. For the system embodiment, since it basically corresponds to the method embodiment, the description is relatively simple, and the relevant parts can be referred to the partial description of the method embodiment.
[0217] The block diagrams of the devices, apparatuses, equipment, and systems involved in this disclosure are only illustrative examples and are not intended to require or imply that they must be connected, arranged, and configured in the manner shown in the block diagrams. As will be appreciated by those skilled in the art, these devices, apparatuses, equipment, and systems can be connected, arranged, and configured in any manner. Words such as "including," "comprising," "having," and the like are open words, referring to "including but not limited to," and can be used interchangeably therewith. The words "or" and "and" used herein refer to the words "and / or," and can be used interchangeably therewith, unless the context clearly indicates otherwise. The word "such as" used herein refers to the phrase "such as but not limited to," and can be used interchangeably therewith.
[0218] The method and apparatus of the present disclosure may be implemented in many ways. For example, the method and apparatus of the present disclosure may be implemented by software, hardware, firmware, or any combination of software, hardware, and firmware. The above order of steps for the method is for illustration only, and the steps of the method of the present disclosure are not limited to the order specifically described above, unless otherwise specifically stated. In addition, in some embodiments, the present disclosure may also be implemented as a program recorded in a recording medium, which includes machine-readable instructions for implementing the method according to the present disclosure. Therefore, the present disclosure also covers a recording medium storing a program for executing the method according to the present disclosure.
[0219] It should also be noted that in the apparatus, device and method of the present disclosure, each component or each step can be decomposed and / or recombined. Such decomposition and / or recombination should be regarded as equivalent solutions of the present disclosure.
[0220] The above description of the disclosed aspects is provided to enable any person skilled in the art to make or use the present disclosure. Various modifications to these aspects will be readily apparent to those skilled in the art, and the general principles defined herein may be applied to other aspects without departing from the scope of the present disclosure. Therefore, the present disclosure is not intended to be limited to the aspects shown herein, but rather to the widest scope consistent with the principles and novel features disclosed herein.
[0221] The above description has been given for the purpose of illustration and description. In addition, this description is not intended to limit the embodiments of the present disclosure to the forms disclosed herein. Although multiple example aspects and embodiments have been discussed above, those skilled in the art will recognize certain variations, modifications, changes, additions and sub-combinations thereof.
Claims
1. A speech signal processing method, characterized in that: include: Processing the speech signal to be processed to obtain a first signal sequence of n frames in length; wherein each frame in the first signal sequence is a signal vector of the same length; and n is an integer greater than 1; Downsampling the first signal sequence to obtain a second signal sequence with a length of m frames; wherein m is an integer greater than or equal to 1, and n is greater than m; each signal vector in the second signal sequence is one of the signal vectors in the first signal sequence, or is calculated from multiple signal vectors in the first signal sequence; Performing self-attention calculation on the first signal sequence and the second signal sequence to obtain a first matrix; Based on the first matrix, a feature conversion result of the speech signal is determined.
2. The method according to claim 1, characterized in that The downsampling the first signal sequence to obtain a second signal sequence with a length of m frames includes: Divide the first signal sequence into m vector groups; wherein each of the vector groups includes at least one signal vector; For each vector group in the m vector groups, down-sampling processing is performed on at least one signal vector included in the vector group to obtain a down-sampled vector; Based on the m down-sampling vectors corresponding to the m vector groups, a second signal sequence with a length of m frames is obtained.
3. The method according to claim 2, characterized in that The dividing the first signal sequence into m vector groups comprises: The n frames of signal vectors included in the first signal sequence are equally divided into m equal parts to obtain the m vector groups.
4. The method according to claim 2, characterized in that: The dividing the first signal sequence into m vector groups comprises: Converting the first signal sequence into an n-dimensional segmentation vector; wherein the segmentation vector is composed of 0 and / or 1; Determine m groups of target positions based on a position of at least one 1 segmented by 0 in the segmentation vector; For each group of target signal positions in the m groups of target positions, a signal vector of a corresponding position in the first signal sequence is determined based on the target position to obtain one of the vector groups.
5. The method according to claim 4, characterized in that The converting the first signal sequence into an n-dimensional segmentation vector comprises: Converting the n signal vectors included in the first signal sequence into numerical representations to obtain an n-dimensional intermediate vector including n numerical values; Determine the magnitude relationship between the value of each numerical value in the n-dimensional intermediate vector and a set threshold value; The numerical values in the n-dimensional intermediate vector that are greater than or equal to the set threshold are converted to 1, and the numerical values that are less than the set threshold are converted to 0; and the segmentation vector is obtained.
6. The method according to any one of claims 1 to 5, characterized in that: The performing self-attention calculation on the first signal sequence and the second signal sequence to obtain a first matrix includes: For each of the n signal vectors included in the first signal sequence, perform self-attention calculation on the signal vector and the m down-sampling vectors included in the second signal sequence to obtain a calculation result; The first matrix is obtained based on the n calculation results.
7. The method according to any one of claims 1 to 5, characterized in that: Before determining the feature conversion result of the speech signal based on the first matrix, the method further includes: For each signal vector of the n signal vectors included in the first signal sequence, perform self-attention calculation on the signal vector and a set number of signal vectors adjacent to the signal vector to obtain a second matrix corresponding to the first signal sequence; The determining, based on the first matrix, a feature conversion result of the speech signal includes: Based on the first matrix and the second matrix, a feature conversion result of the speech signal is determined.
8. The method according to claim 7, characterized in that The determining, based on the first matrix and the second matrix, a feature conversion result of the speech signal includes: Performing matrix addition on the first matrix and the second matrix to obtain a superposition matrix; Perform at least one encoding process and at least one decoding process on the superposition matrix to obtain a feature conversion result of the speech signal.
9. The method according to any one of claims 1 to 5, characterized in that: Also includes: Performing feature extraction and feature conversion on the feature conversion result to obtain a wake-up probability value; Based on the magnitude relationship between the wake-up probability value and the wake-up threshold, it is determined whether to wake up the preset device.
10. The method according to any one of claims 1 to 5, characterized in that: Also includes: Based on the feature conversion result, a first word sequence is obtained by matching with a set word table; wherein the first word sequence includes a plurality of word vectors; Based on the first word sequence, a text recognition result corresponding to the speech signal is determined.
11. The method according to claim 10, characterized in that The determining, based on the first word sequence, a text recognition result corresponding to the speech signal includes: Downsampling the first word sequence to obtain a second word sequence with a reduced length; Performing self-attention calculation on the first word sequence and the second word sequence to obtain an intermediate representation result; Based on the intermediate representation result, a text recognition result of the speech signal is determined.
12. A speech signal processing device, characterized in that: include: A signal processing module, used for processing the speech signal to be processed to obtain a first signal sequence of n frames in length; wherein each frame in the first signal sequence is a signal vector of the same length; and n is an integer greater than 1; a downsampling module, configured to perform downsampling processing on the first signal sequence to obtain a second signal sequence of m frames in length; wherein m is an integer greater than or equal to 1, and n is greater than m; and each signal vector in the second signal sequence is one of the signal vectors in the first signal sequence, or is calculated from multiple signal vectors in the first signal sequence; A self-attention module, configured to perform self-attention calculation on the first signal sequence and the second signal sequence to obtain a first matrix; A feature conversion module is used to determine a feature conversion result of the speech signal based on the first matrix.
13. A computer-readable storage medium, characterized in that: Computer program instructions are stored thereon, and it is characterized in that when the computer program instructions are executed by the processor, the speech signal processing method described in any one of claims 1-11 above is implemented.
14. An electronic device, characterized in that: The electronic device comprises: a memory for storing a computer program product; A processor is used to execute the computer program product stored in the memory, and when the computer program product is executed, the speech signal processing method described in any one of claims 1 to 11 is implemented.
15. A computer program product comprising computer program instructions, characterized in that When the computer program instructions are executed by a processor, the speech signal processing method described in any one of claims 1 to 11 is implemented.
Citation Information
Patent Citations
Wake-up voice determination method and device, equipment and medium
CN111933112A
Image format conversion method and device, equipment, storage medium and program product
CN113487524A
Image processing method and device, equipment, storage medium and computer program product
CN113642585A