Speech recognition method and device, electronic equipment and storage medium
By replacing the address book decoding diagram in the voice recognition system with the reserved location of the general decoding diagram, the target decoding diagram is formed, and the problem of low recognition accuracy of general links and equipment links is solved, and high-precision recognition of multi-task voice commands is achieved.
Patent Information
- Application Number
- CN202311755394.7
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2023-12-19
- Publication Date
- 2025-06-20
AI Technical Summary
In the vehicle voice recognition application scenario, the general voice recognition link has low recognition accuracy for device-related instructions, while the device voice recognition link has low recognition accuracy for general instructions, resulting in low accuracy for final recognition results.
By replacing the reserved position of the second data decoding map with the first data decoding map, the target decoding map is obtained, and the task information of at least one task type is identified from the input voice data using the decoder and the target decoding map.
The fusion of data of different task types is realized in a decoded graph, and the task information of different task types is supported at the same time, which improves the recognition accuracy of audio containing multiple voice commands.
Smart Images

Figure CN120183384A_ABST
Abstract
Description
Technical Field
[0001] The present disclosure relates to the field of speech recognition technology, and in particular, to a speech recognition method, apparatus, electronic device, and storage medium. Background Art
[0002] Speech recognition technology is a technology that collects speech signals and decodes the speech using an acoustic model and a language model to identify the text and language structure therein. During the speech recognition process, a decoding graph in the language model can be used to perform a matching search on the acoustic scores output by the acoustic model to select the optimal encoding path and obtain the final recognition result; among them, in the in-vehicle speech recognition application scenario, for general instructions applicable to different vehicles, such as vehicle control instructions like opening the window and turning on the navigation, a general decoding graph can be pre-configured in the in-vehicle terminal, and a general speech recognition link established based on the general decoding graph is used to recognize such instructions; while for device-related instructions related to the user device, such as making a call and playing audio in the mobile phone, etc., after the user device is connected to the in-vehicle terminal, relevant data in the user device needs to be obtained, such as the user address book, audio files stored in the user device, etc., and a device decoding graph is established based on the obtained data, and a device speech recognition link established based on the device decoding graph is used to recognize such instructions. Since the audio to be recognized may include task instructions that match different speech recognition links, for example, an audio may simultaneously include a call instruction for making a call and a control instruction for opening the window, therefore, as Figure 1 shown, after the in-vehicle terminal receives the input audio, it needs to perform recognition in the general speech recognition link and the device speech recognition link respectively, and arbitrate the recognition results of the two links to select the optimal recognition result as the final recognition result.
[0003] However, since the general speech recognition link performs recognition based on the general decoding graph, the general speech recognition link has low recognition accuracy for device-related instructions. And since the device speech recognition link performs recognition based on the device decoding graph, the device speech recognition link has low recognition accuracy for general instructions, resulting in a low accuracy of the final recognition result. Summary of the Invention
[0004] The present disclosure provides a speech recognition method, apparatus, electronic device, and storage medium.
[0005] According to a first aspect of the present disclosure, an information method is provided, the method comprising:
[0006] Determine a first decoding graph according to first data, and determine a second decoding graph according to second data; wherein, the second decoding graph includes a reserved first string, and the task types of the first data and the second data do not match;
[0007] Replace the first string position of the second decoded graph with the first decoded graph to obtain a target decoded graph;
[0008] Use a decoder and the target decoded graph to identify task information of at least one task type from the input voice data to obtain an identification result.
[0009] In the above solution, the determining the second decoded graph according to the second data includes:
[0010] Determine an initial decoded graph according to the second data;
[0011] Reserve a first position in the initial decoded graph with a first string to obtain a second decoded graph; the first string includes an identifier for identifying that the first position stores the first decoded graph.
[0012] In the above solution, the determining the first decoded graph according to the first data includes:
[0013] Identify text units in the first data according to a preset word list;
[0014] According to a preset dictionary, obtain the serial number and language score of the pronunciation unit corresponding to the text unit, where the language score is used to represent the probability that the pronunciation unit maps to the text unit;
[0015] Construct a first decoded graph according to the text unit, the serial number of the pronunciation unit, and the language score.
[0016] In the above solution, the using a decoder and the target decoded graph to identify task information of at least one task type from the input voice data to obtain an identification result includes:
[0017] Extract acoustic features from the input voice data to determine the acoustic features of the voice data;
[0018] Use an acoustic model to process the acoustic features of the voice data to obtain the serial number and acoustic score of the pronunciation unit corresponding to the voice data, where the acoustic score is used to represent the probability that the voice data maps to the pronunciation unit;
[0019] Use the decoder to search for an identification path of the voice data in the target decoded graph based on the acoustic score and a dynamically extended threshold to obtain a target identification path;
[0020] Based on the target identification path, determine task information of at least one task type from the text units corresponding to the target path to obtain an identification result.
[0021] In the above solution, using the decoder to search for the recognition path of the speech data in the target decoding graph based on the acoustic score and the dynamic expansion threshold to obtain the target recognition path includes:
[0022] Traverse the target decoding graph according to the serial number and acoustic score of the pronunciation unit corresponding to the speech data in a preset time series. The target decoding graph includes nodes and paths between nodes, and the paths of the target decoding graph include the serial number of the pronunciation unit, the text unit, and the language score;
[0023] Based on the dynamic expansion threshold, generate a node expansion list corresponding to each time point in the time series. The node expansion list includes the nodes selected on the target decoding graph at each time point;
[0024] According to the node expansion list corresponding to each time point in the time series, determine the optional recognition paths of the speech data;
[0025] Select the path with the highest total score from the optional recognition paths as the target recognition path, where the total score is the sum of the scores of all paths between nodes in the optional recognition path, and the score of the path between nodes includes the acoustic score and the language score.
[0026] In the above solution, generating a node expansion list corresponding to each time point in the time series based on the dynamic expansion threshold includes:
[0027] Judge whether the number of nodes in the node expansion list corresponding to the first time point is less than the preset number of nodes;
[0028] When the number of nodes is less than the preset number of nodes, set an upper limit for the dynamic expansion threshold, and perform pruning processing on the node expansion list based on the restricted dynamic expansion threshold to obtain the nodes in the node expansion list corresponding to the second time point, where the second time point is after the first time point.
[0029] In the above solution, the first decoding graph is an address book decoding graph, and the second decoding graph is a general decoding graph. Determining the first decoding graph according to the first data and determining the second decoding graph according to the second data includes:
[0030] Determine the address book decoding graph according to the address book data, and determine the general decoding graph according to the general data, where the general decoding graph includes a reserved first string;
[0031] Replacing the position of the first string in the second decoding graph with the first decoding graph to obtain the target decoding graph includes:
[0032] Replace the first string position of the general decoding graph with the address book decoding graph to obtain a target decoding graph.
[0033] According to a second aspect of the present disclosure, there is provided a speech recognition device, the device comprising:
[0034] A first processing unit, configured to determine a first decoding graph according to first data and determine a second decoding graph according to second data; wherein, the second decoding graph includes a reserved first string, and the task types corresponding to the first data and the second data are different;
[0035] A second processing unit, configured to add the first decoding graph to the first string position of the second decoding graph to obtain a target decoding graph;
[0036] A third processing unit, configured to use a decoder and the target decoding graph to identify task information of at least one task type from input speech data to obtain an identification result.
[0037] According to a third aspect of the present disclosure, there is provided an electronic device, comprising:
[0038] At least one processor; and
[0039] A memory communicatively connected to the at least one processor; wherein,
[0040] The memory stores instructions executable by the at least one processor, and the instructions are executed by the at least one processor so that the at least one processor can execute the method described in the foregoing first aspect.
[0041] According to a fourth aspect of the present disclosure, there is provided a non-transitory computer-readable storage medium storing computer instructions, wherein the computer instructions are used to cause the computer to execute the method described in the foregoing first aspect.
[0042] According to a fifth aspect of the present disclosure, there is provided a computer program product, comprising a computer program, where the computer program implements the method described in the foregoing first aspect when executed by a processor.
[0043] The speech recognition method, device, electronic device and storage medium provided by the present disclosure determine a first decoding graph according to first data and determine a second decoding graph according to second data; wherein, the second decoding graph includes a reserved first string, and the task types of the first data and the second data do not match; replace the position of the first string in the second decoding graph with the first decoding graph to obtain a target decoding graph; use a decoder and the target decoding graph to recognize task information of at least one task type from the input speech data to obtain a recognition result. The technical solution provided by the embodiments of the present disclosure can obtain a multi-task decoding graph by adding the decoding graph of the first data to the reserved position of the second data decoding graph, and can realize data fusion of different task types in one decoding graph, so that during the speech recognition process, it is possible to support simultaneous recognition of task information of different task types and improve the recognition accuracy of audio containing multiple speech commands.
[0044] It should be understood that the content described in this part is not intended to identify the key or important features of the embodiments of the present application, nor is it used to limit the scope of the present application. Other features of the present application will become easily understood through the following description. BRIEF DESCRIPTION OF THE DRAWINGS
[0045] The drawings are used to better understand the solution and do not constitute a limitation to the present disclosure. Among them:
[0046] Figure 1 is a flowchart of a speech recognition method in the related art;
[0047] Figure 2 is a flowchart of a speech recognition method provided by an embodiment of the present disclosure;
[0048] Figure 3 is a flowchart of another speech recognition provided by an embodiment of the present disclosure;
[0049] Figure 4 is a flowchart of a cross-domain multi-task recognition method provided by an application embodiment of the present disclosure;
[0050] Figure 5 is a schematic structural diagram of a speech recognition device provided by an embodiment of the present disclosure;
[0051] Figure 6 is a schematic block diagram of an exemplary electronic device 600 provided by an embodiment of the present disclosure. DETAILED DESCRIPTION OF THE EMBODIMENTS
[0052] The exemplary embodiments of the present disclosure will be described below in conjunction with the accompanying drawings. Various details of the embodiments of the present disclosure are included to facilitate understanding, and they should be considered merely exemplary. Therefore, those of ordinary skill in the art should recognize that various changes and modifications can be made to the embodiments described herein without departing from the scope and spirit of the present disclosure. Similarly, descriptions of well-known functions and structures are omitted in the following description for clarity and conciseness.
[0053] The speech recognition processing method, device, electronic device, and storage medium according to the embodiments of the present disclosure will be described below with reference to the accompanying drawings.
[0054] Figure 2 A speech recognition method is provided, which is applied to an in-vehicle terminal. Specifically, it can be applied to a car machine. In actual application, the car machine can also be referred to as HU. The embodiments of the present disclosure do not limit this, as long as its functions can be realized.
[0055] As Figure 2 shown, the method may include:
[0056] Step 201: Determine a first decoding graph according to first data, and determine a second decoding graph according to second data; wherein, the second decoding graph includes a reserved first string, and the task types of the first data and the second data do not match.
[0057] In actual application, the first data may be data of a user device. Specifically, it may be data obtained by the in-vehicle terminal from the user device after the user device is connected to the in-vehicle terminal. Exemplarily, the first data may be address book data.
[0058] In an embodiment, the method may further include: obtaining the first data from the user device.
[0059] In actual application, the user device may be a mobile phone, a tablet computer, a smart wearable device, etc.
[0060] In actual application, the second data may be general data related to vehicle control. For example, control data related to in-vehicle devices such as windows and seats, and control data related to in-vehicle terminal devices such as navigation and audio players. Exemplarily, the first data may include vehicle control data and POI-related data; the second data may also be referred to as general data. The embodiments of the present disclosure do not limit this, as long as its functions can be realized.
[0061] In an embodiment, the determining a first decoding graph according to first data and determining a second decoding graph according to second data includes: determining an initial decoding graph according to the second data; reserving a first position in the initial decoding graph with the first string to obtain a second decoding graph; the first string includes an identifier, and the identifier is used to identify that the first decoding graph is stored in the first position
[0062] In practical applications, the second data is general data applicable to different vehicles, and the first data is address book data. An address book decoding graph is determined based on the address book data, and a general decoding graph is determined based on the general data. In the general decoding graph, a preset string is used for placeholder filling, and the position for placeholder filling is random and can be any empty position in the general decoding graph. An identifier is used to identify this position, and this position is used to store the address book decoding graph. For example, in the general decoding graph, the special string "@CONTACT" is used for placeholder filling at any empty position, and the identifier is 1101, indicating that this position is used to store the address book decoding graph. The constructed general decoding graph is pre-configured in the vehicle-mounted terminal.
[0063] In one embodiment, a decoding graph is constructed using knowledge sources such as a pronunciation dictionary, an acoustic model, and a language model. The decoding graph includes nodes and paths, and the paths include pronunciation unit serial numbers, text units, and language scores. Among them, the pronunciation dictionary is a bridge between the acoustic model and the language model. The trained pronunciation dictionary includes pronunciation unit serial numbers, pronunciation units, corresponding text units, and language scores.
[0064] In one embodiment, determining a first decoding graph based on the first data includes: identifying text units in the first data according to a preset word list; obtaining the serial number and language score of the pronunciation unit corresponding to the text unit according to a preset dictionary, where the language score is used to represent the probability that the pronunciation unit is mapped to the text unit; and constructing the first decoding graph based on the text unit, the serial number of the pronunciation unit, and the language score.
[0065] In one embodiment, determining an initial decoding graph based on the second data includes: identifying text units in the second data according to a preset word list; obtaining the serial number and language score of the pronunciation unit corresponding to the text unit according to a preset dictionary, where the language score is used to represent the probability that the pronunciation unit is mapped to the text unit; and constructing the initial decoding graph based on the text unit, the serial number of the pronunciation unit, and the language score.
[0066] In practical applications, the first decoding graph can also be referred to as the device decoding graph. Further, the first decoding graph can be an address book decoding graph. The embodiments of the present disclosure do not limit this, as long as its function can be realized.
[0067] In practical applications, the second decoding graph can also be referred to as the general decoding graph. The embodiments of the present disclosure do not limit this, as long as its function can be realized.
[0068] In actual application, the task type may include any one of vehicle control tasks, making a call tasks, and POI-related tasks. Among them, the vehicle control task may include control tasks for vehicle internal devices. For example, instructions such as opening the window, adjusting the seat, and turning on the windshield wipers. Making a call may be a control instruction for a user device connected to the vehicle terminal. Specifically, it is a control instruction to query a phone number from the user's address book after reading the user's address book and control the user device to make a call. The POI-related task may include a task for identifying an interesting location or place mentioned in the voice. Specifically, it can identify the location mentioned in the input audio, and by comparing the extracted location information with the data in the navigation system, determine the specific longitude and latitude coordinates, so as to determine the location of the POI.
[0069] In actual application, the task type may also include other tasks related to the user device or tasks related to vehicle device control, which are not limited in this disclosure.
[0070] Step 202: Replace the first string position of the second decoded graph with the first decoded graph to obtain a target decoded graph.
[0071] In actual application, when the first data is obtained from the user device, the first data can be converted into a first decoded graph and added to the reserved position of the second decoded graph, that is, replace the preset string in the second decoded graph with the first decoded graph to form a complete cross-domain multi-task decoded graph.
[0072] In actual application, the target decoded graph can also be called a multi-task decoded graph, which includes data related to multiple task types.
[0073] Step 203: Use the decoder and the target decoded graph to identify task information of at least one task type from the input voice data to obtain an identification result.
[0074] In actual application, an acoustic model can be used to extract features from the input audio to obtain the sequence number and acoustic score of the pronunciation unit. Then, use the sequence number and acoustic score of the pronunciation unit to perform a matching search in the target decoded graph. According to the sequence number of the pronunciation unit, the corresponding text unit, and the language score on the path in the decoded graph, select the path with the largest total score. The text on the path with the largest score, that is, the task information of each task, so as to obtain the final identification result.
[0075] In some embodiments, preferably, the first decoded graph is an address book decoded graph, and the second decoded graph is a general decoded graph. Through the above method, the problem that the recognition accuracy of names in the general link recognition link is not high and the recognition effect of vehicle control instructions in the address book recognition link is very poor can be solved.
[0076] In summary, the speech recognition method provided by the embodiments of the present disclosure obtains a cross-domain multi-task decoding graph by replacing the reserved position of the decoding graph of the second data with the decoding graph of the first data, and can realize data fusion of different task types in one decoding graph, so that during the speech recognition process, it can support the simultaneous recognition of task information of different task types, and improve the recognition accuracy of audio containing multiple speech commands.
[0077] Based on Figure 2 the embodiments shown, Figure 3 FIG. further shows a schematic flowchart of a speech recognition method provided by the embodiments of the present disclosure. As Figure 3 shown, the method may include:
[0078] Step 301: Determine a first decoding graph according to the first data, and determine a second decoding graph according to the second data; wherein, the second decoding graph includes a reserved first string, and the task types of the first data and the second data do not match.
[0079] In practical applications, the first data may be data of a user device. Specifically, it may be data obtained by a vehicle-mounted terminal from the user device after the user device establishes a connection with the vehicle-mounted terminal. Exemplarily, the first data may be address book data.
[0080] In one embodiment, the method may further include: obtaining the first data from the user device.
[0081] In practical applications, the user device may be a mobile phone, a tablet computer, a smart wearable device, etc.
[0082] In one embodiment, a decoding graph is constructed using knowledge sources such as a pronunciation dictionary, an acoustic model, and a language model. The decoding graph includes nodes and paths. The paths include pronunciation unit serial numbers, text units, and language scores. Among them, the pronunciation dictionary is a bridge between the acoustic model and the language model. The trained pronunciation dictionary includes pronunciation unit serial numbers, pronunciation units, corresponding text units, and language scores.
[0083] In one embodiment, determining the first decoding graph according to the first data includes: identifying the text unit in the first data according to a preset word list; obtaining the serial number and language score of the pronunciation unit corresponding to the text unit according to a preset dictionary, where the language score is used to represent the probability that the pronunciation unit is mapped to the text unit; constructing the first decoding graph according to the text unit, the serial number of the pronunciation unit, and the language score.
[0084] In practical applications, if the first data is address book data and the first data includes "Liu Daqiang", according to the preset word list, two words "Liu" and "Daqiang" can be recognized. According to the preset dictionary, the corresponding pronunciation and corresponding language scores can be searched from the dictionary. For example, the probability of the pronunciation "liu" to the word "Liu" is 0.69315, and the serial number of the pronunciation "liu" is 15883. The pronunciation units corresponding to the obtained text units can be multiple.
[0085] In practical applications, the second data can be general data related to vehicle control. For example, control data related to in-vehicle devices such as windows and seats, and control data related to in-vehicle terminal devices such as navigation and audio players; Exemplarily, the first data may include vehicle control data and POI-related data; The second data can also be referred to as general data, and the embodiments of the present disclosure do not limit this, as long as its function can be realized.
[0086] In one embodiment, determining the first decoding graph according to the first data and determining the second decoding graph according to the second data includes: determining an initial decoding graph according to the second data; reserving a first position in the initial decoding graph with a first string to obtain a second decoding graph; The first string includes an identifier, and the identifier is used to identify that the first decoding graph is stored in the first position.
[0087] In one embodiment, determining the initial decoding graph according to the second data includes: identifying text units in the second data according to the preset word list; obtaining the serial number and language score of the pronunciation unit corresponding to the text unit according to the preset dictionary, and the language score is used to represent the probability that the pronunciation unit is mapped to the text unit; constructing an initial decoding graph according to the text unit, the serial number of the pronunciation unit, and the language score.
[0088] In practical applications, since the second data is general data and applicable to different vehicles, therefore, the second data can be pre-converted into a second decoding graph, and during the process of constructing the second decoding graph, a preset string is used for placeholder. The first string includes an identifier, and the identifier is used to identify that the first decoding graph is stored in the first position. For example, in the second decoding graph, a special string "@CONTACT" is used to replace the decoding graph of the address book data for placeholder, and an identifier is used to mark the placeholder position of the special string "@CONTACT". The identifier of the special string "@CONTACT" is 1101, and 1101 indicates that the first decoding graph is stored in this position. Then, the constructed second decoding graph is pre-configured in the in-vehicle terminal.
[0089] In practical applications, the first decoding graph can also be referred to as a device decoding graph. Further, the first decoding graph can be an address book decoding graph. The embodiments of the present disclosure do not limit this, as long as its function can be realized.
[0090] In practical applications, the second decoded graph can also be referred to as a general decoded graph. The embodiments of the present disclosure do not limit this, as long as its functions can be achieved.
[0091] In practical applications, the task type can include any one of vehicle control tasks, making a call tasks, and POI-related tasks. Among them, vehicle control tasks can include control tasks for vehicle internal devices. For example, instructions such as opening the window, adjusting the seat, and turning on the windshield wipers. Making a call can be a control instruction for a user device connected to the vehicle terminal. Specifically, it is a control instruction to query a phone number from the user's address book after reading the user's address book and control the user device to make a call. POI-related tasks can include tasks for identifying interesting locations or places mentioned in the voice. Specifically, it can identify the locations mentioned in the input audio, and by comparing the extracted location information with the data in the navigation system, determine the specific longitude and latitude coordinates, so as to determine the location of the POI.
[0092] In practical applications, the task type can also include other tasks related to user devices or tasks related to vehicle device control. The present disclosure does not limit this.
[0093] Step 302: Replace the first string position of the second decoded graph with the first decoded graph to obtain a target decoded graph.
[0094] In one embodiment, step 302 includes: in response to obtaining the second decoded graph, obtaining the first position according to the identifier; replacing the first string in the first position with the first decoded graph to obtain a target decoded graph.
[0095] In one embodiment, the target decoded graph simultaneously includes text units of first data and second data, serial numbers of pronunciation units, and language scores.
[0096] In practical applications, after obtaining the first data from the user device, the first data can be converted into a first decoded graph, and the first decoded graph can be added to the reserved position of the second decoded graph. That is, the placeholder position is obtained through the identifier, and the preset string at the placeholder position in the second decoded graph is replaced with the second decoded graph to form a complete cross-domain multi-task decoded graph.
[0097] In practical applications, the target decoded graph can also be referred to as a multi-task decoded graph and includes data related to multiple task types.
[0098] Step 303: Extract acoustic features from the input voice data to determine the acoustic features of the voice data.
[0099] In practical applications, the input audio signal is first divided into a series of frames, each frame can be from 10 ms to 20 ms, and then, each frame of the signal is processed by a window function to extract the effective features of the signal, and information reflecting the speech features is extracted from each frame of the signal.
[0100] Step 304: Process the acoustic features of the speech data by using an acoustic model to obtain the sequence number and acoustic score of the pronunciation unit corresponding to the speech data, where the acoustic score is used to represent the probability that the speech data is mapped to the pronunciation unit.
[0101] In practical applications, an acoustic model is trained by using acoustic features and corresponding labels. During the processing of the acoustic model, it is used to establish the corresponding probability distribution between the acoustic signal and the pronunciation unit. The pronunciation unit or the modeling unit includes HMM (Hidden Markov model) states, phonemes, syllables, characters, etc. The acoustic model can adopt model structures such as GMM-HMM, DNN-HMM, etc., which are not limited in the present disclosure.
[0102] Step 304: Use the decoder to search for the recognition path of the speech data in the target decoding graph based on the acoustic score and the dynamic expansion threshold to obtain the target recognition path.
[0103] In one embodiment, the step of using the decoder to search for the recognition path of the speech data in the target decoding graph based on the acoustic score and the dynamic expansion threshold to obtain the target recognition path may include:
[0104] Traverse the target decoding graph according to the sequence number and acoustic score of the pronunciation unit corresponding to the speech data in a preset time series, where the target decoding graph includes nodes and paths between the nodes, and the paths of the target decoding graph include the sequence number of the pronunciation unit, the text unit, and the language score;
[0105] Generate a node expansion list corresponding to each time point in the time series based on the dynamic expansion threshold, where the node expansion list includes the nodes selected on the target decoding graph at each time point;
[0106] Determine the optional recognition paths of the speech data according to the node expansion list corresponding to each time point in the time series;
[0107] Select the path with the highest total score as the target recognition path from the optional recognition paths, where the total score is the sum of the scores of all the paths between the nodes of the optional recognition path, and the score of the path between the nodes includes the acoustic score and the language score.
[0108] In practical applications, the Viterbi algorithm can be used to search for the globally optimal path in the target decoding graph. The path search algorithm is a dynamic programming pruning algorithm. The adaptive beam is used to dynamically adjust the number of nodes at the next moment during the decoding process to achieve path pruning. When dynamically adjusting the number of nodes during the decoding process, all the nodes at the current moment form the node expansion list at the current moment, and the node expansion list at the next moment is generated based on the node expansion list at the current moment and the value of the parameter adaptive beam.
[0109] In practical applications, when the number of nodes in the current coding list is less than a certain number, for example, when the number of nodes in the coding list corresponding to the current moment is less than 3, the adaptive beam at the next moment is infinite, resulting in a relatively large number of nodes in the coding list corresponding to the next moment. A large number of nodes need to be traversed when traversing the coding list corresponding to the next moment, which greatly affects the decoding speed.
[0110] Based on this, in one embodiment, generating the node expansion list at the next moment according to the node expansion list at the current moment and the value of the parameter adaptive beam includes:
[0111] Determine whether the number of nodes in the node expansion list corresponding to the first time point is less than the preset number of nodes;
[0112] In the case where the number of nodes is less than the preset number of nodes, set an upper limit threshold for the adaptive beam, and perform pruning processing on the node expansion list based on the restricted adaptive beam to obtain the nodes in the node expansion list corresponding to the second time point, where the first time point is the current moment and the second time point is the next moment of the current moment.
[0113] In the embodiments of the present disclosure, when the number of coding nodes corresponding to the current moment is less than the preset number, an upper limit threshold is set for the value of the parameter adaptive beam to prevent the value of the adaptive beam from being infinite. Based on the restricted value of the adaptive beam, the second node expansion list corresponding to the next moment is generated, avoiding the problem of a relatively large number of nodes in the second node expansion list, thereby reducing the number of coding nodes traversed in the target decoding graph during the decoding process, accelerating the decoding speed, and improving the real-time rate of decoding.
[0114] In practical applications, for each node expansion list, by traversing the nodes in the node expansion list, the optimal node in the node expansion list can be determined, and the target decoding path can be obtained according to the optimal nodes of each node expansion list.
[0115] Step 305: Based on the target recognition path, determine task information of at least one task type from the text units corresponding to the target path to obtain a recognition result.
[0116] In practical applications, the task information can also be referred to as task instruction information. The embodiments of the present disclosure do not limit this, as long as its function can be achieved.
[0117] In summary, for the speech recognition method provided by the embodiments of the present disclosure, by replacing the reserved position of the decoding graph of the second data with the decoding graph of the first data, a cross-domain multi-task decoding graph is obtained, which can realize data fusion of different task types in one decoding graph. Therefore, during the speech recognition process, it can support the simultaneous recognition of task information of different task types, improve the recognition accuracy of audio containing multiple speech commands, and by setting an upper threshold for the parameter adaptive beam during the decoding process, the number of node traversals in the decoding graph during the decoding process is reduced, thereby accelerating the decoding speed.
[0118] The technical solution of the present disclosure will be further described in detail below with reference to application embodiments.
[0119] Figure 4 FIG. is a schematic flowchart of a cross-domain multi-task recognition method provided by an application embodiment of the present disclosure; as Figure 4 shown, the method includes:
[0120] Step 1: Generate a general decoding graph and a contact book decoding graph.
[0121] Among them, during the construction process of the general decoding graph, a special string such as "@CONTACT" is used to replace the contact sub-graph for placeholder; when receiving a user speech recognition request and obtaining contact book data, the contact book data will be converted through a conversion process into a contact book decoding graph.
[0122] Step 2: Replace the "@CONTACT" string in the general decoding graph with the contact book decoding graph to form a complete cross-domain multi-task decoding graph.
[0123] Step 3: Use an acoustic model to calculate the acoustic probability of the audio (i.e., speech data).
[0124] Among them, Step 3 can be executed before Step 2, after Step 2, or simultaneously with Step 2.
[0125] Step 4: Combine the acoustic probability output by the acoustic model and decode the cross-domain multi-task decoding graph in a decoder to obtain the text and related information of the speech recognition.
[0126] Among them, the related information may include the position of the text or vocabulary in the text in the audio, specifically, the time information in the audio.
[0127] Among them, during the execution of step 4, when the number of nodes in the node expansion list corresponding to the current moment is less than 3, the upper limit of the parameter adaptive beam is set to generate the state list corresponding to the next moment, so as to reduce the number of node traversals in the decoding graph during the decoding process, thereby accelerating the decoding speed.
[0128] In summary, the application embodiment of the present disclosure has the following advantages:
[0129] (1) By constructing the address book information into an address book decoding graph and then merging the address book decoding graph into the general decoding graph, cross-domain multi-task recognition including the names in the address book can be realized in one audio, which can solve the problems in the related technology that the recognition accuracy of names in the general link recognition link is not high and the recognition effect of vehicle control instructions in the address book recognition link is very poor;
[0130] (2) Limiting the upper limit of adaptive beam in the decoder reduces the number of node traversals in the decoding graph during the decoding process, thereby accelerating the decoding speed.
[0131] Corresponding to the above speech recognition method, the present invention also proposes a speech recognition device. Since the device embodiment of the present invention corresponds to the above method embodiment, for the details not disclosed in the device embodiment, reference may be made to the above method embodiment, and the present invention will not be elaborated herein.
[0132] Figure 5 It is a schematic structural diagram of a speech recognition device provided by an embodiment of the present disclosure, as Figure 5 shown. The device may include:
[0133] A first processing unit 501, configured to determine a first decoding graph according to first data and determine a second decoding graph according to second data; wherein, the second decoding graph includes a reserved first string, and the task types corresponding to the first data and the second data are different;
[0134] A second processing unit 502, configured to replace the position of the first string in the second decoding graph with the first decoding graph to obtain a target decoding graph;
[0135] A third processing unit 503, configured to use a decoder and the target decoding graph to identify task information of at least one task type from the input speech data to obtain a recognition result.
[0136] In an embodiment, the first processing unit 501 may specifically be configured to:
[0137] Determine an initial decoding graph according to the second data;
[0138] Reserve a first position in the decoding graph with a first string to obtain a second decoding graph; the first string includes an identifier for identifying that the first position stores a first decoding graph.
[0139] In one embodiment, the first processing unit 501 may specifically be configured to:
[0140] Identify text units in the first data according to a preset word list;
[0141] Obtain the serial number and language score of the pronunciation unit corresponding to the text unit according to a preset dictionary, where the language score is used to represent the probability that the pronunciation unit is mapped to the text unit;
[0142] Construct a first decoding graph according to the text unit, the serial number of the pronunciation unit, and the language score.
[0143] In one embodiment, the third processing unit 503 may specifically be configured to:
[0144] Extract acoustic features from the input speech data to determine the acoustic features of the speech data;
[0145] Process the acoustic features of the speech data using an acoustic model to obtain the serial number and acoustic score of the pronunciation unit corresponding to the speech data, where the acoustic score is used to represent the probability that the speech data is mapped to the pronunciation unit;
[0146] Use the decoder to search for the recognition path of the speech data in the target decoding graph based on the acoustic score and the dynamic expansion threshold to obtain a target recognition path;
[0147] Based on the target recognition path, determine task information of at least one task type from the text units corresponding to the target path to obtain a recognition result.
[0148] In one embodiment, the third processing unit 503 may specifically be configured to:
[0149] Traverse the target decoding graph according to the serial number and acoustic score of the pronunciation unit corresponding to the speech data in a preset time series, where the target decoding graph includes nodes and paths between nodes, and the serial number of the pronunciation unit, the text unit, and the language score are included on the path of the target decoding graph;
[0150] Generate a node expansion list corresponding to each time point in the time series based on the dynamic expansion threshold, where the node expansion list includes the nodes selected on the target decoding graph at each time point;
[0151] Determine an optional recognition path of the speech data according to the node expansion list corresponding to each time point in the time series;
[0152] Select the path with the highest total score from the optional recognition paths as the target recognition path, where the total score is the sum of the scores of all the paths between the nodes of the optional recognition path, and the score of the path between the nodes includes the acoustic score and the language score.
[0153] In one embodiment, the third processing unit 503 can also be used to:
[0154] Determine whether the number of nodes in the node expansion list corresponding to the first time point is less than a preset number of nodes;
[0155] In the case where the number of the nodes is less than the preset number of nodes, set an upper limit for the dynamic expansion threshold, and perform pruning processing on the node expansion list based on the restricted dynamic expansion threshold to obtain the nodes in the node expansion list corresponding to the second time point, where the second time point is after the first time point.
[0156] In one embodiment, the first decoding graph is an address book decoding graph, and the second decoding graph is a general decoding graph. The first processing unit 501 can specifically be used to:
[0157] Determine an address book decoding graph according to address book data, and determine a general decoding graph through general data training, where the general decoding graph includes a reserved first string;
[0158] The obtaining the target decoding graph by replacing the first string position of the second decoding graph with the first decoding graph includes:
[0159] Replace the first string position of the general decoding graph with the address book decoding graph to obtain the target decoding graph.
[0160] It should be noted that the foregoing explanations of the method embodiments also apply to the device in this embodiment, with the same principle, and will not be limited in this embodiment.
[0161] According to an embodiment of the present disclosure, the present disclosure also provides an electronic device, a readable storage medium, and a computer program product.
[0162] Figure 6FIG. 0 is a schematic block diagram of an exemplary electronic device 600 that may be used to implement embodiments of the present disclosure. The electronic device is intended to represent various forms of digital computers, such as, for example, laptop computers, desktop computers, workstations, personal digital assistants, servers, blade servers, mainframe computers, and other suitable computers. The electronic device may also represent various forms of mobile devices, such as, for example, personal digital processors, cellular telephones, smart phones, wearable devices, and other similar computing devices. The components shown herein, their connections and relationships, and their functions are merely exemplary and are not intended to limit the implementations of the present disclosure described and / or claimed herein.
[0163] As Figure 6 shown, the device 600 includes a computing unit 601 that may perform various appropriate actions and processes in accordance with a computer program stored in a ROM (Read-Only Memory) 602 or a computer program loaded from a storage unit 608 into a RAM (Random Access Memory) 603. In the RAM 603, various programs and data required for the operation of the device 600 may also be stored. The computing unit 601, the ROM 602, and the RAM 603 are connected to each other via a bus 604. An I / O (Input / Output) interface 605 is also connected to the bus 604.
[0164] A plurality of components in the device 600 are connected to the I / O interface 605, including: an input unit 606, such as a keyboard, a mouse, etc.; an output unit 607, such as various types of displays, speakers, etc.; a storage unit 608, such as a magnetic disk, an optical disk, etc.; and a communication unit 609, such as a network card, a modem, a wireless communication transceiver, etc. The communication unit 609 allows the device 600 to exchange information / data with other devices via a computer network such as the Internet and / or various telecommunication networks.
[0165] The computing unit 601 can be various general and / or special processing components with processing and computing capabilities. Some examples of the computing unit 601 include, but are not limited to, CPU (Central Processing Unit), GPU (Graphic Processing Units), various dedicated AI (Artificial Intelligence) computing chips, various computing units running machine learning model algorithms, DSP (Digital Signal Processor), and any suitable processor, controller, microcontroller, etc. The computing unit 601 executes the various methods and processes described above, such as the speech recognition method. For example, in some embodiments, the speech recognition method can be implemented as a computer software program tangibly embodied in a machine-readable medium, such as the storage unit 608. In some embodiments, part or all of the computer program can be loaded and / or installed onto the device 600 via the ROM 602 and / or the communication unit 609. When the computer program is loaded into the RAM 603 and executed by the computing unit 601, one or more steps of the methods described above can be performed. Alternatively, in other embodiments, the computing unit 601 can be configured to execute the foregoing speech recognition method in any other suitable manner (e.g., by means of firmware).
[0166] The various embodiments of the systems and techniques described above in this document can be implemented in digital electronic circuit systems, integrated circuit systems, FPGA (Field Programmable Gate Array), ASIC (Application-Specific Integrated Circuit), ASSP (Application Specific Standard Product), SOC (System On Chip), CPLD (Complex Programmable Logic Device), computer hardware, firmware, software, and / or combinations thereof. These various embodiments can include: being implemented in one or more computer programs that can be executed and / or interpreted on a programmable system including at least one programmable processor, which can be a dedicated or general programmable processor, and can receive data and instructions from a storage system, at least one input device, and at least one output device, and transmit the data and instructions to the storage system, the at least one input device, and the at least one output device.
[0167] The program code for implementing the methods of the present disclosure can be written in any combination of one or more programming languages. These program codes can be provided to a processor or controller of a general purpose computer, a special purpose computer, or other programmable data processing device, such that the program codes, when executed by the processor or controller, cause the functions / operations specified in the flowchart and / or block diagram to be implemented. The program codes can be executed entirely on the machine, partially on the machine, as a stand-alone software package partially on the machine and partially on a remote machine, or entirely on a remote machine or server.
[0168] In the context of the present disclosure, a machine-readable medium can be a tangible medium that can contain or store a program for use by or in connection with an instruction execution system, apparatus, or device. A machine-readable medium can be a machine-readable signal medium or a machine-readable storage medium. A machine-readable medium can include, but is not limited to, electronic, magnetic, optical, electromagnetic, infrared, or semiconductor systems, apparatus, or devices, or any suitable combination of the foregoing. More specific examples of a machine-readable storage medium would include an electrical connection based on one or more wires, a portable computer diskette, a hard disk, a RAM, a ROM, an EPROM (Electrically Programmable Read-Only-Memory), or flash memory, an optical fiber, a CD-ROM (Compact Disc Read-Only Memory), an optical storage device, a magnetic storage device, or any suitable combination of the foregoing.
[0169] In order to provide interaction with a user, the systems and techniques described herein can be implemented on a computer having: a display device (e.g., a CRT (Cathode-Ray Tube) or LCD (Liquid Crystal Display) monitor) for displaying information to the user; and a keyboard and a pointing device (e.g., a mouse or a trackball) through which the user can provide input to the computer. Other kinds of devices can also be used to provide interaction with the user; for example, the feedback provided to the user can be any form of sensory feedback (e.g., visual feedback, auditory feedback, or tactile feedback); and input from the user can be received in any form (including acoustic input, speech input, or tactile input).
[0170] The systems and techniques described herein can be implemented in a computing system that includes backend components (e.g., as a data server), or a computing system that includes middleware components (e.g., an application server), or a computing system that includes frontend components (e.g., a user computer having a graphical user interface or a web browser through which a user can interact with an implementation of the systems and techniques described herein), or a computing system that includes any combination of such backend components, middleware components, or frontend components. The components of the system can be interconnected to each other by digital data communication in any form or medium (e.g., a communication network). Examples of communication networks include: LAN (Local Area Network), WAN (Wide Area Network), the Internet, and blockchain networks.
[0171] A computer system can include a client and a server. The client and the server are generally far from each other and usually interact through a communication network. The client-server relationship is created by computer programs that run on the respective computers and have a client-server relationship with each other. The server can be a cloud server, also known as a cloud computing server or a cloud host, which is a host product in the cloud computing service system, and solves the defects of difficult management and weak business scalability existing in traditional physical hosts and VPS services ("Virtual Private Server", or simply "VPS"). The server can also be a server of a distributed system, or a server combined with blockchain.
[0172] It should be noted that artificial intelligence is a discipline that studies to make a computer simulate certain thinking processes and intelligent behaviors of humans (such as learning, reasoning, thinking, planning, etc.), and there are both hardware-level technologies and software-level technologies. Artificial intelligence hardware technologies generally include technologies such as sensors, dedicated artificial intelligence chips, cloud computing, distributed storage, and big data processing; artificial intelligence software technologies mainly include several major directions such as computer vision technology, speech recognition technology, natural language processing technology, and machine learning / deep learning, big data processing technology, and knowledge graph technology.
[0173] It should be understood that various forms of processes shown above can be used, with steps reordered, added, or deleted. For example, the steps recited in this disclosure can be executed in parallel, sequentially, or in a different order, as long as the desired results of the technical solutions disclosed in this disclosure can be achieved, and no limitations are imposed herein.
[0174] The above specific embodiments do not constitute a limitation to the protection scope of the present disclosure. Those skilled in the art should understand that various modifications, combinations, sub-combinations and substitutions can be made according to design requirements and other factors. Any modifications, equivalent substitutions and improvements made within the spirit and principle of the present disclosure shall be included within the protection scope of the present disclosure.
Claims
1. A speech recognition method, characterized in that, including: determining a first decoded graph according to first data, and determining a second decoded graph according to second data; wherein, the second decoded graph includes a reserved first string, and the task types corresponding to the first data and the second data do not match; replacing the position of the first string in the second decoded graph with the first decoded graph to obtain a target decoded graph; using a decoder and the target decoded graph to identify task information of at least one task type from the input voice data to obtain an identification result.
2. The method according to claim 1, characterized in that, The determining the second decoded graph according to the second data includes: determining an initial decoded graph according to the second data; reserving a first position with a first string in the initial decoded graph to obtain a second decoded graph; the first string includes an identifier for identifying that the first position stores the first decoded graph.
3. The method according to claim 1, characterized in that, The determining the first decoded graph according to the first data includes: identifying text units in the first data according to a preset word list; acquiring the serial number and language score of the pronunciation unit corresponding to the text unit according to a preset dictionary, where the language score is used to represent the probability that the pronunciation unit is mapped to the text unit; constructing a first decoded graph according to the text unit, the serial number of the pronunciation unit, and the language score.
4. The method according to claim 3, characterized in that, The using the decoder and the target decoded graph to identify task information of at least one task type from the input voice data to obtain an identification result includes: extracting acoustic features from the input voice data to determine the acoustic features of the voice data; processing the acoustic features of the voice data by using an acoustic model to obtain the serial number and acoustic score of the pronunciation unit corresponding to the voice data, where the acoustic score is used to represent the probability that the voice data is mapped to the pronunciation unit; using the decoder, based on the acoustic score and a dynamic expansion threshold, to search for an identification path of the voice data in the target decoded graph to obtain a target identification path; based on the target identification path, determining task information of at least one task type from the text units corresponding to the target path to obtain an identification result.
5. The method according to claim 4, characterized in that, The using the decoder, based on the acoustic score and a dynamic expansion threshold, to search for an identification path of the voice data in the target decoded graph to obtain a target identification path includes: traversing the target decoded graph according to the serial number and acoustic score of the pronunciation unit corresponding to the voice data in a preset time series, where the target decoded graph includes nodes and paths between the nodes, and the path of the target decoded graph includes the serial number of the pronunciation unit, the text unit, and the language score; generating a node expansion list corresponding to each time point in the time series based on the dynamic expansion threshold, where the node expansion list includes the nodes selected on the target decoded graph at each time point; determining an optional identification path of the voice data according to the node expansion list corresponding to each time point in the time series; Select the path with the highest total score from the optional recognition paths, where the total score is the sum of the scores of all the paths between the nodes of the optional recognition path, and the score of the path between the nodes includes the acoustic score and the language score.
6. The method according to claim 5, characterized in that, Generating the node expansion list corresponding to each time point in the time series based on the dynamic expansion threshold includes: Determine whether the number of nodes in the node expansion list corresponding to the first time point is less than a preset number of nodes; In the case where the number of the nodes is less than the preset number of nodes, set an upper limit for the dynamic expansion threshold, and perform pruning processing on the node expansion list based on the restricted dynamic expansion threshold to obtain the nodes in the node expansion list corresponding to the second time point, where the second time point is after the first time point.
7. The method according to any one of claims 1-6, characterized in that, The first decoding graph is an address book decoding graph, and the second decoding graph is a general decoding graph. Determining the first decoding graph according to the first data and the second decoding graph according to the second data includes: Determine the address book decoding graph according to the address book data, and determine the general decoding graph according to the general data, where the general decoding graph includes a reserved first string; Replacing the first string position of the second decoding graph with the first decoding graph to obtain the target decoding graph includes: Replace the first string position of the general decoding graph with the address book decoding graph to obtain the target decoding graph.
8. A speech recognition device, characterized in that, Includes: A first processing unit for determining a first decoding graph according to the first data and a second decoding graph according to the second data; where the second decoding graph includes a reserved first string, and the task types corresponding to the first data and the second data are different; A second processing unit for replacing the first string position of the second decoding graph with the first decoding graph to obtain the target decoding graph; A third processing unit for using a decoder and the target decoding graph to identify task information of at least one task type from the input voice data to obtain a recognition result.
9. An electronic device, characterized in that, Includes: At least one processor; And A memory communicatively connected to the at least one processor; where The memory stores instructions executable by the at least one processor, and the instructions are executed by the at least one processor so that the at least one processor can execute the method according to any one of claims 1 to 7.
10. A non-transitory computer-readable storage medium storing computer instructions, characterized in that, The computer instructions are used to cause the computer to execute the method according to any one of claims 1 to 7.
11. A vehicle, characterized in that, Includes the voice recognition device according to claim 8.