Speech recognition method, device, equipment and storage medium

By introducing the directional characteristics of the voice signal into the far-field speech recognition system and the output results of multiple hidden layers of the neural network for computing, the problem of inconsistent optimization goals of the front-end enhancement module and the back-end recognition system is solved, the recognition rate and accuracy are improved, and the end-to-end speech recognition processing is realized.

CN114267359BActive Publication Date: 2025-08-08BEIJING BAIDU NETCOM SCI & TECH CO LTD
View PDF 2 Cites 0 Cited by

Patent Information

Application Number
CN202111570340.4
Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Filing Date
2021-12-21
Publication Date
2025-08-08
Estimated Expiration
2041-12-21

AI Technical Summary

Technical Problem

In the existing far-field voice recognition system, the optimization goals of the front-end enhancement module and the back-end recognition system are not unified, resulting in the recognition rate being undesirable.

Method used

By introducing the directional characteristics of the speech signal and the output results of multiple hidden layers of the neural network, the recognition rate and recognition accuracy of speech recognition are improved.

Benefits of technology

It greatly improves the recognition rate and recognition accuracy of speech recognition, suppresses the interference of environmental noise, and realizes end-to-end speech recognition processing.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN114267359B_ABST
    Figure CN114267359B_ABST
Patent Text Reader

Abstract

The present disclosure provides a speech recognition method, apparatus, device, and storage medium, relating to the fields of artificial intelligence technology, specifically deep learning and speech recognition technology. A specific implementation scheme comprises: determining a directional feature of a target speech based on azimuth information; the azimuth information is used to indicate the direction of the sound source of the target speech; determining an intermediate feature of the target speech; the intermediate feature comprising at least one of N intermediate results before a final result is obtained by recognizing the target speech; N being an integer not less than 1; and determining a recognition result of the target speech based on the directional feature and the intermediate feature.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present disclosure relates to the field of artificial intelligence technology, specifically deep learning and speech recognition technology. Background Art

[0002] In the field of far-field speech recognition, non-end-to-end speech recognition models are typically used to achieve speech enhancement for specific sound source directions. The front-end enhancement module in the recognition model utilizes speech signal processing technology to enhance the target signal and improve signal quality. The back-end recognition module uses data acquired from the front-end to train and optimize the speech recognition model and output recognition results. However, the misalignment of optimization objectives between the front-end speech enhancement module and the back-end recognition system results in suboptimal recognition rates for the entire recognition system.

[0003] To this end, how to develop an end-to-end speech recognition model and improve the recognition rate of the model based on directional features have become technical problems that need to be solved. Summary of the Invention

[0004] The present disclosure provides a speech recognition method, apparatus, device, and storage medium.

[0005] According to one aspect of the present disclosure, a speech recognition method is provided, which may include the following steps:

[0006] Determine the directional characteristics of the target speech based on the directional information; the directional information is used to indicate the direction of the sound source of the target speech;

[0007] Determining an intermediate feature of the target speech; the intermediate feature includes at least one of N intermediate results before obtaining a final result by recognizing the target speech; N is an integer not less than 1;

[0008] The recognition result of the target speech is determined based on the directional features and the intermediate features.

[0009] According to another aspect of the present disclosure, a speech recognition device is provided, which may include:

[0010] A directional feature determination module is used to determine the directional features of the target speech based on the directional information; the directional information is used to indicate the direction of the sound source of the target speech;

[0011] An intermediate feature determination module, configured to determine an intermediate feature of the target speech; the intermediate feature comprises at least one of N intermediate results obtained before the final result is obtained by recognizing the target speech; N is an integer not less than 1;

[0012] The recognition module is used to determine the recognition result of the target speech based on the directional features and the intermediate features.

[0013] According to another aspect of the present disclosure, there is provided an electronic device, comprising:

[0014] at least one processor; and

[0015] a memory communicatively connected to the at least one processor; wherein,

[0016] The memory stores instructions that can be executed by the at least one processor. The instructions are executed by the at least one processor to enable the at least one processor to perform the method in any embodiment of the present disclosure.

[0017] According to another aspect of the present disclosure, a non-transitory computer-readable storage medium storing computer instructions is provided. The computer instructions are used to cause a computer to execute the method in any embodiment of the present disclosure.

[0018] According to another aspect of the present disclosure, a computer program product is provided, including a computer program, which implements the method in any embodiment of the present disclosure when executed by a processor.

[0019] The disclosed technical solution incorporates the directional characteristics of the sound source corresponding to the speech signal and performs corresponding calculations on the output results of multiple hidden layers in the neural network, ultimately obtaining output features that include directional characteristics. This significantly improves the recognition rate and accuracy of speech recognition.

[0020] It should be understood that the contents described in this section are not intended to identify the key or important features of the embodiments of the present disclosure, nor are they intended to limit the scope of the present disclosure. Other features of the present disclosure will become readily understood through the following description. BRIEF DESCRIPTION OF THE DRAWINGS

[0021] The accompanying drawings are provided to facilitate a better understanding of the present invention and do not constitute a limitation of the present disclosure.

[0022] Figure 1 is a flow chart of a speech recognition method according to the present disclosure;

[0023] Figure 2 is a flow chart of a method for determining position information according to the present disclosure;

[0024] Figure 3 is a flow chart of a method for determining a theoretical phase difference according to the present disclosure;

[0025] Figure 4 is a schematic diagram of determining the angle of arrival according to the present disclosure;

[0026] Figure 5 is a flow chart of a method for determining an actual phase difference according to the present disclosure;

[0027] Figure 6 is a flow chart of a method for determining intermediate features according to the present disclosure;

[0028] Figure 7 is a flow chart of a method for determining recognition results according to the present disclosure;

[0029] Figure 8 is a schematic diagram of performing a vector dot product operation according to the present disclosure;

[0030] Figure 9 This is a schematic diagram of performing matrix merging operations according to the present disclosure. Figure 1 ;

[0031] Figure 10 This is a schematic diagram of performing matrix merging operations according to the present disclosure. Figure 2 ;

[0032] Figure 11 is a structural diagram of a speech recognition device according to the present disclosure;

[0033] Figure 12 A block diagram of an electronic device implementing feature image processing according to an embodiment of the present disclosure. DETAILED DESCRIPTION

[0034] The following description of exemplary embodiments of the present disclosure is made in conjunction with the accompanying drawings, including various details of the embodiments of the present disclosure to facilitate understanding. These details should be considered as merely exemplary. Therefore, those skilled in the art will recognize that various changes and modifications may be made to the embodiments described herein without departing from the scope and spirit of the present disclosure. Similarly, for the sake of clarity and conciseness, descriptions of well-known functions and structures are omitted in the following description.

[0035] like Figure 1 As shown, the present disclosure relates to a speech recognition method, which may include the following steps:

[0036] S101: Determine the directional characteristics of the target speech according to the directional information; the directional information is used to indicate the direction of the sound source of the target speech;

[0037] S102: Determine an intermediate feature of the target speech; the intermediate feature includes at least one of N intermediate results before obtaining a final result after recognizing the target speech; N is an integer not less than 1;

[0038] S103: Determine the recognition result of the target speech according to the directional features and the intermediate features.

[0039] This embodiment can be applied to computer devices, which may specifically include but are not limited to servers, desktop computers, laptop computers, cloud computers, or a server set consisting of multiple servers. This application does not limit the product type of the computer device.

[0040] In one implementation, the embodiment may be executed by a speech recognition neural network model in a computer device, comprising a stacked convolutional neural network, a recurrent neural network, and a sequence-to-sequence network. The model may output a corresponding speech recognition result based on an input target speech signal, wherein the speech recognition result may be text information corresponding to the target speech.

[0041] The target voice may be any form of voice signal from a fixed direction, such as a human voice or voice played by an electronic device, which is not limited here.

[0042] Direction information indicates the direction of the target speech's source. Specifically, it can be the location of the target speech's source relative to the location of the audio receiver. For example, this can be determined by determining the distance or angle of the sound source relative to the audio receiver. Based on this direction information input into the speech recognition neural network model, the corresponding directional features can be encoded.

[0043] After determining the directional features of the target speech, intermediate features can be further determined based on the speech recognition neural network model. Specifically, the N intermediate results output by the N intermediate layers within the speech recognition neural network model can be used as the intermediate features of the target speech, where N can be set as needed, for example, N = 1, 2, 3, etc., and is not limited here.

[0044] The recognition result of the target speech is determined according to the directional features and the intermediate features. The directional features and the intermediate features can be calculated based on a preset fusion strategy, and then the final speech recognition result is determined based on the calculation results.

[0045] Through the above process, the speech recognition neural network model integrates the orientation information into the intermediate features based on the directional features, thereby improving the accuracy of speech recognition.

[0046] like Figure 2 In one embodiment, the method for determining the position information includes:

[0047] S201: Determine a theoretical phase difference and an actual phase difference of the target speech relative to the receiving device;

[0048] S202: Determine the orientation information using the theoretical phase difference and the actual phase difference.

[0049] The sound receiving device may be an electronic device including a plurality of microphone arrays. The number of microphones may be set to 2, 3, 4, etc. as needed, and is not limited here.

[0050] The microphone array of a sound receiving device can be arranged in different ways. Based on the different arrangements, sound receiving devices can be divided into linear arrays, circular arrays, and stereo arrays. Among them, linear array sound receiving devices can be arranged in a straight line, circular array microphones can be arranged in a circle, and stereo arrays are arranged in three dimensions, such as tetrahedral arrays, cube arrays, spherical arrays, etc., which are not exhaustive here.

[0051] Based on the positional relationship between the microphone array in the sound receiving device and the sound source of the target speech, a theoretical phase difference of the target speech relative to the multiple microphones in the sound receiving device can be determined. The number of theoretical phase differences obtained is K, where K is a positive integer not less than 1.

[0052] When the sound receiving device detects the target speech signal, the actual phase difference can be determined based on the actual phase values of the target speech relative to the multiple microphones in the sound receiving device. The actual phase value is the phase value detected by the microphone after the target speech is affected by the ambient noise.

[0053] At the same time, the actual phase difference also corresponds to a specific microphone array in the sound receiving device, and the number of actual phase differences is also K, where K is a positive integer not less than 1.

[0054] The direction information is determined by using the theoretical phase difference and the actual phase difference. The theoretical phase difference and the corresponding actual phase difference can be used to perform cosine similarity calculation, and the calculated direction-steering vector is used as the direction information of the target speech.

[0055]

[0056] in, Used to indicate the location information of the target speech. and They represent the theoretical phase difference and actual phase difference of the target speech relative to the receiving device respectively.

[0057] Through the above process, the azimuth information of the target speech relative to the sound receiving device can be determined, and then the target speech can be enhanced based on the azimuth information, thereby reducing the impact of environmental noise on the target recognition model and improving the recognition accuracy of the speech recognition model.

[0058] like Figure 3 As shown, in one embodiment, step S201 includes the following sub-steps:

[0059] S301: Determine the angle of arrival of the target speech;

[0060] S302: Determine the theoretical phase difference of the target speech relative to the sound receiving device using the arrival angle of the target speech.

[0061] The angle of arrival is used to indicate the angle of the target speech relative to the sound receiving device. Figure 4 As shown, taking mic1 (the first microphone) and mic2 (the second microphone) as an example, first determine the center point O of mic1 and mic2, the angle between the line connecting point A where the sound source of the target speech is located and point O and the line connecting mic1 and mic2, that is, the arrival angle of the target speech relative to mic1 and mic2.

[0062] Then, according to traditional signal theory, the determined arrival angle of the target speech is substituted into the following formula (2) to calculate the theoretical phase difference.

[0063]

[0064] in, represents the theoretical phase difference, f represents the frequency of the target speech, d represents the distance between the two microphones, θ represents the angle of arrival, and c represents the speed of light. The frequency f of the target speech can be an M-dimensional vector, where each dimension represents the frequency value of the target speech. The value of M can be 68, 128, 256, and is not limited here. The calculated theoretical phase difference is a vector with the same dimensions as the frequency of the target speech.

[0065] When the number of microphones in the sound receiving device is greater than 2, such as Figure 4 For example, the sound receiving device further includes mic3 (a third microphone). Each pair of microphones from mic1, mic2, and mic3 forms a group, resulting in a total of K microphone arrays. The target speech corresponds to an angle of arrival relative to each microphone array, and K theoretical phase differences are then determined, where K is a positive integer not less than 1.

[0066] like Figure 5 As shown, in one embodiment, step S202 includes the following sub-steps:

[0067] S501: Converting the time domain signal of the target speech into a frequency domain signal;

[0068] S502: Determine at least two phase values of the target speech relative to the sound receiving device based on the frequency domain signal;

[0069] S503: Determine an actual phase difference of the target speech relative to the sound receiving device using at least two phase values.

[0070] The target speech received by the sound receiving device includes the pure sound emitted from the sound source position and the ambient noise. After calculating the theoretical phase difference of the target speech, it is necessary to further determine the actual phase difference of the target speech.

[0071] First, a Fourier transform is performed on the target speech received by the microphone to convert the time domain signal of the target speech into a frequency domain signal. Specifically, a short-time Fourier transform (STFT) or a fast Fourier transform (FFT) can be used, but the two are not limited here. Specifically, the Fourier transform can be performed according to the following formula (3).

[0072]

[0073] Among them, y(n) represents the discrete time domain signal obtained by sampling, Y t,f represents the frequency domain signal obtained by Fourier transform of y(n), N represents the signal sampling point, n represents the number of sampling points, w represents the window function, and f represents the frequency of the target speech.

[0074] Secondly, at least two phase values of the target speech relative to the sound receiving device are determined based on the frequency domain signal. Specifically, the number of phase values can be set as needed, for example, 2, 3, 4, etc., which is not limited here.

[0075] When the number of phase values obtained by the radio equipment is equal to 2, such as Figure 4 As shown, the target speech has a phase value relative to each microphone in the sound receiving device. Target speech A corresponds to a phase value relative to mic1 and mic2 respectively, and the phase values and the frequency vector of the target speech are vectors with the same dimension. According to the following formula (4), the difference between the two phase values is taken as the actual phase difference of the target speech relative to the sound receiving device.

[0076]

[0077] in, Indicates the actual phase difference, Indicates the actual phase value of the target voice relative to mic1. Indicates the actual phase value of the target speech relative to mic2.

[0078] If the number of phase values obtained by the sound receiving device is greater than 2, the differences between each of the multiple phase values can be used to determine K initial phase differences. The number of initial phase differences (the value of K) is the same as the number of theoretical phase differences. The K initial phase differences are then used to determine the actual phase difference of the target speech relative to the sound receiving device, where K is a positive integer not less than 3. The K phase differences and the actual phase difference are both vectors with the same dimensions as the frequency vector of the target speech.

[0079] For example, if Figure 4As shown, when the sound receiving device includes three microphones, mic1, mic2, and mic3, the phase values of the target speech A relative to mic1, mic2, and mic3 can be expressed as φ1, φ2, and φ3, respectively. The three initial phase differences are obtained by taking the difference between any two of the three phase values. Among them, the phase difference of the target speech A relative to mic1 and mic2 is φ 12 , the phase difference of the target speech A relative to mic2 and mic3 is φ 23 , the phase difference of the target speech A relative to mic1 and mic3 is φ 13 Among them, φ 12 、φ 23 and φ 13 are both vectors with the same dimension as the frequency vector of the target speech.

[0080] The actual phase difference of the target speech relative to the sound receiving device is determined by using the K initial phase differences. The K initial phase differences can be directly used as the actual phase difference, or the actual phase difference can be obtained by calculating multiple initial phase differences.

[0081] In one embodiment, the actual phase difference is calculated from multiple initial phase differences. This may be achieved by averaging the vectors using the multiple initial phase differences and using the average as the actual phase difference. Alternatively, this may be achieved by weighted summing the vectors corresponding to the multiple initial phase differences and using the sum as the actual phase difference. Other calculation methods are also possible and are not limited herein.

[0082] The number of phase values obtained based on the radio equipment can also be 4, 5, 6, etc., which are not exhaustive here.

[0083] Through the above process, based on determining the theoretical phase difference and actual phase difference of the target speech relative to the receiving device, the azimuth information of the target speech relative to the receiving device can be determined, and then the target speech can be enhanced based on the azimuth information to improve the recognition accuracy of the speech recognition model.

[0084] In one embodiment, step S101 may include the following sub-steps:

[0085] The orientation information is input into a pre-trained first branch network model, which then outputs the orientation features. The first branch network is a branch network in the speech recognition neural network model, which may include multiple convolutional layers, recursive layers, and fully connected layers.

[0086] Based on multiple convolutional and recursive layers, the directional information of the target speech can be extracted to obtain K directional features. The directional features are also M-dimensional vectors.

[0087] Through the above process, the corresponding direction vector can be determined based on the orientation information. In this way, the direction vector can be fused with the intermediate features of the target speech to achieve end-to-end speech recognition processing and improve the efficiency and accuracy of speech recognition.

[0088] like Figure 6 As shown, in one embodiment, step S102 includes the following sub-steps:

[0089] S601: Extracting spectral features of the target speech;

[0090] S602: Input the spectral features into a pre-trained second branch network model; the second branch network includes N hidden layers;

[0091] S603: Taking the intermediate result output by the i-th hidden layer as the i-th intermediate feature; i is an integer not less than 1 and not greater than N;

[0092] The spectral features may be frequency features distributed along the time axis, for example, mel spectrum, log mel spectrum energy, mel spectrum cepstral coefficients, etc., which are not limited here.

[0093] Extracting the spectral features of the target speech can be implemented using a feature extraction module. The feature extraction module includes a Fourier transform layer and a beamforming submodule. The beamforming submodule can include a complex convolution layer, a complex bias layer, and a complex linear transform layer. The number of intermediate layers can be set as needed and is not limited here.

[0094] The obtained spectral features are input into a pre-trained second-branch network model, and the multiple hidden layers in the pre-trained second-branch network model output intermediate features. The second-branch network is another branch network in the speech recognition neural network model and may include multiple hidden layers. For example, convolutional layers, low-frame-rate feature extraction layers, long short-term memory networks (LSTMs), and streaming multi-level truncated attention (SMLTA) layers, etc., are not limited here.

[0095] The number of hidden layers in the second branch network is N, and the intermediate result output by the i-th hidden layer is used as the i-th intermediate feature. The number of each hidden layer can be set as needed and is not limited here.

[0096] For the extraction of the spectral feature, the intermediate feature obtained can correspond to a matrix of size M×H, where its dimension in the frequency direction is the same as the frequency dimension of the target speech.

[0097] Through the above process, the intermediate features of the target speech can be determined based on the second branch network model. This allows the directional features to be integrated with the intermediate features of the target speech, achieving end-to-end speech recognition processing and improving the efficiency and accuracy of speech recognition.

[0098] like Figure 7 As shown, in one embodiment, step S103 includes the following sub-steps:

[0099] S701: Perform a preset operation using the directional feature and the i-th intermediate feature to obtain a preset operation result;

[0100] S702: When the preset operation result meets the predetermined conditions, the preset operation result is input to the output layer, and the output layer determines the recognition result of the target speech.

[0101] The directional features may be K M-dimensional vectors, the i-th intermediate feature may be a matrix of size M×H, and a preset operation is performed using the directional features and the i-th intermediate feature to obtain a preset operation result.

[0102] In one embodiment, the preset operation includes at least one of a vector dot product operation and a matrix merge operation.

[0103] Specifically, when K=1, that is, when the number of directional features is 1, only the vector dot product operation or the matrix merging operation may be used, or both operations may be used simultaneously.

[0104] To simplify the description, for example, the directional feature is a 4-dimensional vector (M=4), and the i-th intermediate feature is a matrix of size 4×3. Figure 8 As shown, the vector dot product operation is performed using the directional feature and the i-th intermediate feature, that is, each element in the directional feature is multiplied in sequence with a column of elements of the i-th intermediate feature, and the result of the multiplication operation is used to overwrite the original column, and the same operation is performed on each column in sequence, so that the new matrix obtained is used as the preset operation result.

[0105] like Figure 9 As shown, a matrix merge operation is performed using the directional features and the i-th intermediate features. That is, the directional features are combined as a new column with the i-th intermediate features. Specifically, the merge can be performed before the first column or after the last column, which is not limited here. The merged matrix is used as the preset operation result.

[0106] In one embodiment, the directional feature and the i-th intermediate feature may be matrix-merged and then further subjected to a vector dot product operation. The specific process is the same as above and will not be described in detail here.

[0107] When K is greater than 1, that is, when the number of directional features is greater than 1, K M-dimensional directional features can be merged to obtain an M×K matrix. At this time, the preset operation can be a merge operation between two matrices, and the M×(K+H) matrix obtained by merging the matrices is used as the preset operation result.

[0108] like Figure 10 As shown, the first and second directional features are both 4-dimensional directional features. A matrix merge operation is performed with the i-th intermediate feature matrix of size 4×3, i.e., M = 4, K = 2, H = 3, resulting in a result matrix of size 4×5. Similarly, the merge position can be before the first column or after the last column, which is not limited here.

[0109] If the preset operation result meets the predetermined condition, the preset operation result is input to the output layer, and the output layer obtains the recognition result of the target speech. The predetermined condition can be that the value of each element in the operation result is less than a preset threshold. For example, the preset threshold can be 5, 10, 20, etc., and the list is not exhaustive here. The predetermined condition can also be that the value of each element is within a preset range, such as (10, 20), (20, 30), etc., and the list is not exhaustive here.

[0110] If the preset operation result does not meet the predetermined conditions, the preset operation result is input to the i+1th hidden layer. At this time, the above operation is repeated using the intermediate features and directional features output by the i+1th hidden layer until the operation result meets the predetermined conditions. This is not detailed here.

[0111] In one embodiment, the operation method of the directional feature and the (i+1)th intermediate feature can be the same as that of the previous layer, or it can be different. For example, if the directional feature and the intermediate feature of the i-th layer are operated using vector dot products, then the directional feature and the intermediate feature of the (i+1)th layer can be operated using vector dot products, or it can be changed to matrix merge operations, which is not limited here.

[0112] By introducing the directional features of the target speech and performing flexible fusion operations with the intermediate features of the target speech, the interference of non-target direction speech signals is suppressed, further improving the accuracy of speech recognition.

[0113] like Figure 11 As shown, the present disclosure relates to a speech recognition device, which may include:

[0114] Directional feature determination module 1101, for determining the directional feature of the target speech according to the directional information; the directional information is used to indicate the direction of the sound source of the target speech;

[0115] An intermediate feature determination module 1102 is configured to determine an intermediate feature of the target speech; the intermediate feature includes at least one of N intermediate results obtained before a final result is obtained by recognizing the target speech; N is an integer not less than 1;

[0116] The recognition module 1103 is configured to determine a recognition result of the target speech according to the directional feature and the intermediate feature.

[0117] In one embodiment, the direction feature determination module 1101 may further include:

[0118] A phase difference determination submodule, configured to determine a theoretical phase difference and an actual phase difference of the target speech relative to the sound receiving device;

[0119] The azimuth information determination submodule is configured to determine the azimuth information using the theoretical phase difference and the actual phase difference.

[0120] In one embodiment, the phase difference determination submodule may further include:

[0121] An angle of arrival determination submodule, configured to determine the angle of arrival of the target speech;

[0122] The theoretical phase difference determination submodule is used to determine the theoretical phase difference of the target speech relative to the sound receiving device using the angle of arrival of the target speech.

[0123] In one embodiment, the phase difference determination submodule may further include:

[0124] A signal conversion submodule, configured to convert the time domain signal of the target speech into a frequency domain signal;

[0125] a phase value determination submodule, configured to determine at least two phase values of the target speech relative to the sound receiving device based on the frequency domain signal;

[0126] The actual phase difference determination submodule is configured to determine an actual phase difference of the target speech relative to the sound receiving device using the at least two phase values.

[0127] In one embodiment, the directional feature determination module is configured to input the orientation information into a pre-trained first branch network model, and the pre-trained first branch network model outputs the directional feature.

[0128] In one embodiment, the intermediate feature determination module may further include:

[0129] A spectrum feature extraction module, used to extract the spectrum features of the target speech;

[0130] An input module, configured to input the spectral features into a pre-trained second branch network model; the second branch network includes N hidden layers;

[0131] An intermediate feature determination submodule, configured to use the intermediate result output by the i-th hidden layer as the i-th intermediate feature; i is an integer not less than 1 and not greater than N;

[0132] In one embodiment, the identification module may further include:

[0133] an operation module, configured to perform a preset operation using the directional feature and the i-th intermediate feature to obtain a preset operation result;

[0134] The recognition result execution submodule is used to input the preset operation result to the output layer when the preset operation result meets the predetermined condition, and the output layer determines the recognition result of the target speech.

[0135] In one embodiment, the identification module further includes:

[0136] The secondary input word module is used to input the preset operation result into the (i+1)th hidden layer when the preset operation result does not meet the predetermined condition.

[0137] In one embodiment, the preset operation includes at least one of a vector dot product operation and a matrix merge operation.

[0138] In the technical solutions disclosed herein, the acquisition, storage, and application of user personal information involved comply with the provisions of relevant laws and regulations and do not violate public order and good morals.

[0139] According to an embodiment of the present disclosure, the present disclosure also provides an electronic device, a readable storage medium, and a computer program product.

[0140] Figure 12 A schematic block diagram of an example electronic device 1200 that can be used to implement embodiments of the present disclosure is shown. The electronic device is intended to represent various forms of digital computers, such as laptop computers, desktop computers, workstations, personal digital assistants, servers, blade servers, mainframe computers, and other suitable computers. The electronic device can also represent various forms of mobile devices, such as personal digital assistants, cellular phones, smartphones, wearable devices, and other similar computing devices. The components shown herein, their connections and relationships, and their functions are provided as examples only and are not intended to limit the implementation of the present disclosure described and / or claimed herein.

[0141] like Figure 12As shown, the device 1200 includes a computing unit 1201, which can perform various appropriate actions and processes according to a computer program stored in a read-only memory (ROM) 1202 or a computer program loaded from a storage unit 1208 into a random access memory (RAM) 1203. Various programs and data required for the operation of the device 1200 can also be stored in the RAM 1203. The computing unit 1201, the ROM 1202, and the RAM 1203 are connected to each other via a bus 1204. An input / output (I / O) interface 1205 is also connected to the bus 1204.

[0142] Various components in device 1200 are connected to I / O interface 1205, including an input unit 1206, such as a keyboard and mouse; an output unit 1207, such as various types of displays and speakers; a storage unit 1208, such as a magnetic disk and optical disk; and a communication unit 1209, such as a network card, a modem, a wireless communication transceiver, etc. Communication unit 1209 allows device 1200 to exchange information / data with other devices via a computer network such as the Internet and / or various telecommunication networks.

[0143] The computing unit 1201 can be a variety of general-purpose and / or specialized processing components with processing and computing capabilities. Some examples of the computing unit 1201 include, but are not limited to, a central processing unit (CPU), a graphics processing unit (GPU), various dedicated artificial intelligence (AI) computing chips, various computing units that run machine learning model algorithms, digital signal processors (DSPs), and any appropriate processors, controllers, microcontrollers, etc. The computing unit 1201 performs the various methods and processes described above, such as the method of speech recognition. For example, in some embodiments, the method of speech recognition can be implemented as a computer software program that is tangibly contained in a machine-readable medium, such as the storage unit 1208. In some embodiments, part or all of the computer program can be loaded and / or installed on the device 1200 via the ROM 1202 and / or the communication unit 1209. When the computer program is loaded into the RAM 1203 and executed by the computing unit 1201, one or more steps of the method of speech recognition described above can be performed. Alternatively, in other embodiments, the computing unit 1201 can be configured to perform the method of speech recognition by any other appropriate means (e.g., by means of firmware).

[0144] Various embodiments of the systems and techniques described herein can be implemented in digital electronic circuit systems, integrated circuit systems, field programmable gate arrays (FPGAs), application specific integrated circuits (ASICs), application specific standard products (ASSPs), system-on-chip systems (SOCs), programmable logic devices (CPLDs), computer hardware, firmware, software, and / or combinations thereof. These various embodiments can include being implemented in one or more computer programs that are executable and / or interpreted on a programmable system that includes at least one programmable processor, which can be a special purpose or general purpose programmable processor that can receive data and instructions from a storage system, at least one input device, and at least one output device, and transmit data and instructions to the storage system, the at least one input device, and the at least one output device.

[0145] The program code for implementing the method of the present disclosure can be written in any combination of one or more programming languages. These program codes can be provided to a processor or controller of a general-purpose computer, a special-purpose computer, or other programmable data processing device so that when the program code is executed by the processor or controller, the functions / operations specified in the flow chart and / or block diagram are implemented. The program code can be executed entirely on the machine, partially on the machine, as a stand-alone software package, partially on the machine and partially on a remote machine, or entirely on a remote machine or server.

[0146] In the context of the present disclosure, a machine-readable medium can be a tangible medium that can contain or store a program for use by or in conjunction with an instruction execution system, device or equipment. A machine-readable medium can be a machine-readable signal medium or a machine-readable storage medium. A machine-readable medium can include, but is not limited to, an electronic, magnetic, optical, electromagnetic, infrared, or semiconductor system, device or equipment, or any suitable combination of the foregoing. A more specific example of a machine-readable storage medium can include an electrical connection based on one or more lines, a portable computer disk, a hard disk, a random access memory (RAM), a read-only memory (ROM), an erasable programmable read-only memory (EPROM or flash memory), an optical fiber, a portable compact disk read-only memory (CD-ROM), an optical storage device, a magnetic storage device, or any suitable combination of the foregoing.

[0147] To provide interaction with a user, the systems and techniques described herein can be implemented on a computer having: a display device (e.g., a CRT (cathode ray tube) or LCD (liquid crystal display) monitor) for displaying information to the user; and a keyboard and pointing device (e.g., a mouse or trackball) through which the user can provide input to the computer. Other types of devices can also be used to provide interaction with the user; for example, the feedback provided to the user can be any form of sensory feedback (e.g., visual feedback, auditory feedback, or tactile feedback); and input from the user can be received in any form (including acoustic input, voice input, or tactile input).

[0148] The systems and techniques described herein can be implemented in a computing system that includes back-end components (e.g., as a data server), or a computing system that includes middleware components (e.g., an application server), or a computing system that includes front-end components (e.g., a user computer having a graphical user interface or a web browser through which a user can interact with implementations of the systems and techniques described herein), or a computing system that includes any combination of such back-end components, middleware components, or front-end components. The components of the system can be interconnected by any form or medium of digital data communication (e.g., a communication network). Examples of communication networks include a local area network (LAN), a wide area network (WAN), and the Internet.

[0149] A computer system may include a client and a server. The client and server are generally remote from each other and typically interact through a communication network. The client-server relationship arises through computer programs running on the respective computers and having a client-server relationship with each other. The server may be a cloud server, a server in a distributed system, or a server integrated with a blockchain.

[0150] It should be understood that the various forms of the processes shown above can be used to reorder, add, or delete steps. For example, the steps described in this disclosure can be performed in parallel, sequentially, or in a different order, as long as the desired results of the technical solutions disclosed in this disclosure can be achieved. This is not a limitation herein.

[0151] The above specific embodiments do not constitute a limitation on the scope of protection of this disclosure. Those skilled in the art will appreciate that various modifications, combinations, sub-combinations, and substitutions may be made based on design requirements and other factors. Any modifications, equivalent substitutions, and improvements made within the spirit and principles of this disclosure shall be included within the scope of protection of this disclosure.

Claims

1. A speech recognition method, comprising: Determine the directional characteristics of the target speech based on the direction information; The direction information is used to indicate the sound source direction of the target speech; Determining an intermediate feature of the target speech; the intermediate feature includes at least one of N intermediate results before obtaining a final result by recognizing the target speech; N is an integer not less than 1; Performing a preset operation on the directional feature and the intermediate feature to obtain a preset operation result; If the preset operation result meets a predetermined condition, the preset operation result is input to the output layer, and the output layer determines the recognition result of the target speech; If the preset operation result does not meet the predetermined condition, input the preset operation result into the next hidden layer, update the intermediate feature based on the output result of the hidden layer, and return to the step of performing the preset operation on the directional feature and the intermediate feature to obtain the preset operation result, until the preset operation result meets the predetermined condition; The determining of the intermediate features of the target speech includes: Extract the spectral features of the target speech; input the spectral features into a pre-trained second branch network model; the second branch network includes N hidden layers; the intermediate result output by the i-th hidden layer is used as the i-th intermediate feature; i is an integer not less than 1 and not greater than N.

2. The method according to claim 1, wherein the method for determining the position information comprises: Determining a theoretical phase difference and an actual phase difference of the target speech relative to a sound receiving device; The orientation information is determined using the theoretical phase difference and the actual phase difference.

3. The method according to claim 2, wherein the method for determining the theoretical phase difference comprises: determining the angle of arrival of the target speech; The theoretical phase difference of the target speech relative to the sound receiving device is determined using the angle of arrival of the target speech.

4. The method according to claim 2, wherein the actual phase difference is determined by: Converting the target speech's time domain signal into a frequency domain signal; Determine at least two phase values of the target speech relative to the sound receiving device based on the frequency domain signal; The actual phase difference of the target speech relative to the sound receiving device is determined using the at least two phase values.

5. The method according to claim 1, wherein Determining the directional characteristics of the target speech according to the direction information includes: The orientation information is input into a pre-trained first branch network model, and the pre-trained first branch network model outputs the direction feature.

6. The method according to claim 1, wherein determining the recognition result of the target speech according to the directional feature and the intermediate feature comprises: A preset operation is performed using the directional feature and the i-th intermediate feature to obtain a preset operation result. 7 . The method according to claim 6 , wherein the preset operation comprises at least one of a vector dot product operation and a matrix merge operation.

8. A speech recognition device comprising: A directional feature determination module is used to determine the directional features of the target speech according to the direction information; The direction information is used to indicate the sound source direction of the target speech; an intermediate feature determination module, configured to determine an intermediate feature of the target speech; the intermediate feature comprising at least one of N intermediate results obtained before a final result is obtained by recognizing the target speech; N being an integer not less than 1; an identification module, configured to perform a preset operation on the directional feature and the intermediate feature to obtain a preset operation result; If the preset operation result meets a predetermined condition, the preset operation result is input to the output layer, and the output layer determines the recognition result of the target speech; If the preset operation result does not meet the predetermined condition, input the preset operation result into the next hidden layer, update the intermediate feature based on the output result of the hidden layer, and return to the step of performing the preset operation on the directional feature and the intermediate feature to obtain the preset operation result, until the preset operation result meets the predetermined condition; The intermediate feature determination module includes: A spectrum feature extraction module, used to extract the spectrum features of the target speech; An input module, configured to input the spectral features into a pre-trained second branch network model; the second branch network includes N hidden layers; The intermediate feature determination submodule is used to use the intermediate result output by the i-th hidden layer as the i-th intermediate feature; i is an integer not less than 1 and not greater than N.

9. The apparatus according to claim 8, wherein the directional feature determination module comprises: A phase difference determination submodule, configured to determine a theoretical phase difference and an actual phase difference of the target speech relative to the sound receiving device; The azimuth information determination submodule is configured to determine the azimuth information using the theoretical phase difference and the actual phase difference.

10. The apparatus according to claim 9, wherein the phase difference determination submodule comprises: An angle of arrival determination submodule, configured to determine the angle of arrival of the target speech; The theoretical phase difference determination submodule is used to determine the theoretical phase difference of the target speech relative to the sound receiving device using the angle of arrival of the target speech.

11. The device according to claim 9, wherein The phase difference determination submodule comprises: A signal conversion submodule, configured to convert the time domain signal of the target speech into a frequency domain signal; a phase value determination submodule, configured to determine at least two phase values of the target speech relative to the sound receiving device based on the frequency domain signal; The actual phase difference determination submodule is configured to determine an actual phase difference of the target speech relative to the sound receiving device using the at least two phase values.

12. The device according to claim 8, wherein The directional feature determination module is used to input the orientation information into a pre-trained first branch network model, and the pre-trained first branch network model outputs the directional feature.

13. The device according to claim 8, wherein The identification module includes: The operation module is used to perform a preset operation using the directional feature and the i-th intermediate feature to obtain a preset operation result. The apparatus according to claim 13 , wherein the preset operation comprises at least one of a vector dot product operation and a matrix merge operation.

15. An electronic device comprising: at least one processor; as well as a memory communicatively connected to the at least one processor; wherein, The memory stores instructions that can be executed by the at least one processor, and the instructions are executed by the at least one processor to enable the at least one processor to perform the method according to any one of claims 1 to 7.

16. A non-transitory computer-readable storage medium storing computer instructions, wherein: The computer instructions are used to cause the computer to execute the method according to any one of claims 1 to 7.

17. A computer program product comprising a computer program, which, when executed by a processor, implements the method according to any one of claims 1 to 7.

Citation Information

Patent Citations

  • Voice recognition method, device and equipment and computer readable storage medium

    CN110992974A

  • Voice processing method and device and electronic equipment

    CN112466327A