Data processing method, speech recognition method and device
By segmenting and accumulating feature recognition of real-time speech data, the problem of balancing latency and accuracy in streaming speech recognition is solved, achieving low latency and high accuracy speech recognition results.
Patent Information
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Filing Date
- 2021-04-21
- Publication Date
- 2026-03-27
AI Technical Summary
Existing speech recognition methods cannot simultaneously balance recognition latency and recognition accuracy. In particular, in streaming speech recognition, incomplete speech fragment information leads to low recognition accuracy, while waiting for complete data to be recognized increases latency.
The real-time input speech data is segmented to obtain a first speech data with a smaller number of frames for preliminary recognition. Multi-frame speech features are accumulated to update the recognition results. The same model is used for two recognitions to reduce latency and improve accuracy.
By segmenting and processing real-time voice data and accumulating feature recognition, a balance between low latency and high recognition accuracy is achieved, ensuring the real-time nature and accuracy of the recognition results.
Smart Images

Figure CN115223546B_ABST
Abstract
Description
TECHNICAL FIELD
[0001] Embodiments of the present application relate to the technical field of data processing, and in particular, to a data processing method, a speech recognition method and a device. BACKGROUND
[0002] Currently, there is a widespread demand for converting data of a first category into data of a second category, such as in machine translation, speech recognition and other application scenarios.
[0003] Taking the speech recognition scenario as an example, it is required to convert speech data into text, and streaming speech recognition, i.e., speech recognition on real-time input speech data, has certain requirements on recognition delay rate and recognition accuracy. However, the existing recognition manner cannot simultaneously satisfy the delay rate and the recognition accuracy. SUMMARY
[0004] Embodiments of the present application provide a data processing method, a speech recognition method and a device to solve the technical problem that the prior art cannot simultaneously consider recognition delay rate and recognition accuracy.
[0005] In a first aspect, a speech recognition method is provided in embodiments of the present application, comprising:
[0006] processing real-time input speech data to obtain first speech data of a first frame number;
[0007] extracting first speech features of the first speech data by using a speech recognition model and recognizing first text corresponding to the first speech features;
[0008] outputting the first text;
[0009] accumulating a plurality of first speech features to obtain second speech data of a second frame number;
[0010] recognizing second text corresponding to the second speech data by using the speech recognition model based on the second speech data;
[0011] updating the output text by using the second text.
[0012] Optionally, the method further comprises:
[0013] accumulating a plurality of first speech data to obtain third speech data of a second frame number;
[0014] The recognizing second text corresponding to the second speech data by using the speech recognition model based on the second speech data comprises:
[0015] recognizing second text corresponding to the second speech data and the third speech data by using the speech recognition model based on the second speech data and the third speech data.
[0016] Optionally, the step of recognizing the corresponding second text using the speech recognition model based on the second speech data and the third speech data includes:
[0017] The second and third speech data are subjected to feature splicing and capacity reduction processing to obtain fourth speech data of a third capacity.
[0018] Based on the fourth speech data, the corresponding second text is identified using the speech recognition model.
[0019] Optionally, the method further includes:
[0020] By accumulating the first text corresponding to the multiple first speech features, a third text is obtained;
[0021] The step of recognizing the corresponding second text based on the second speech data using the speech recognition model includes:
[0022] Based on the second speech data and the third text, the corresponding second text is identified using the speech recognition model.
[0023] Optionally, the step of segmenting the real-time target data to obtain the first audio data of the first frame number includes:
[0024] The real-time input voice data is segmented according to the chronological order to obtain the fifth voice data of the fourth frame.
[0025] Acquire the sixth voice data after the fifth voice data for a first predetermined number of frames, and the seventh voice data that was historically generated before the fifth voice data for a second predetermined number of frames;
[0026] The first voice data is composed of the fifth voice data, the sixth voice data, and the seventh voice data. Secondly, this application provides a data processing method, including:
[0027] The real-time input target data is segmented to obtain the first data of the first capacity;
[0028] The recognition model is used to extract the first feature corresponding to the first data and to identify the first text corresponding to the first feature.
[0029] Output the first text;
[0030] Accumulate multiple first features to obtain second data of second capacity;
[0031] Based on the second data, the corresponding second text is identified using the recognition model;
[0032] Update the output text using the second text.
[0033] Optionally, the method further includes:
[0034] Accumulate multiple first data points to obtain third data points of a second capacity;
[0035] The step of identifying the corresponding second text using the recognition model based on the second data includes:
[0036] Based on the second data and the third data, the corresponding second text is identified using the recognition model.
[0037] Optionally, the step of recognizing the corresponding second text using the recognition model based on the second data and the third data includes:
[0038] The second data and the third data are subjected to feature splicing and capacity reduction processing to obtain the fourth data with the third capacity;
[0039] Based on the fourth data, the corresponding second text is identified using the recognition model.
[0040] Optionally, it also includes:
[0041] By accumulating the first texts corresponding to the multiple first features, a third text is obtained;
[0042] The step of identifying the corresponding second text using the recognition model based on the second data includes:
[0043] Based on the second data and the third text, the corresponding second text is identified using the recognition model.
[0044] Optionally, the step of segmenting the real-time target data to obtain first data of a first capacity includes:
[0045] The real-time input target data is segmented according to the chronological order to obtain the fifth data of the fourth capacity;
[0046] The sixth data of the first predetermined capacity is obtained after the fifth data, and the seventh data of the second predetermined capacity is generated historically before the fifth data;
[0047] The first data is composed of the fifth data, the sixth data, and the seventh data.
[0048] Thirdly, this application provides a data processing method, including:
[0049] Determine the target sample data and corresponding target text for the second capacity;
[0050] The target sample data is segmented to obtain multiple sub-sample data of a first capacity;
[0051] The recognition model is trained by using multiple subsample data as model inputs and the target text as model labels.
[0052] Optionally, the recognition model includes a first encoder and a first decoder, as well as a second encoder and a second decoder;
[0053] The step of training the recognition model by using the multiple sub-sample data as model inputs and the target text as model labels includes:
[0054] Based on the multiple subsample data, the first encoder is used to obtain the first encoded feature and the first decoder is used to obtain the corresponding first predicted text.
[0055] Determine the second sample data generated by accumulating the first encoded features corresponding to the multiple subsample data;
[0056] Based on the second sample data, the second encoder is used to obtain the second encoded feature and the second decoder is used to obtain the corresponding second predicted text.
[0057] The recognition model is trained based on the first predicted text, the second predicted text, and the target text.
[0058] Optionally, the recognition model further includes a first prediction module and a first attention mechanism module;
[0059] The step of obtaining a first encoded feature using the first encoder and obtaining the corresponding first predicted text using the first decoder based on the plurality of subsample data includes:
[0060] Based on the multiple sub-sample data, the first encoder is used to extract the first encoded features corresponding to the multiple sub-sample data respectively;
[0061] The first prediction module is used to predict the number of text elements corresponding to multiple first coding features in sequence.
[0062] Based on the prediction results of the first prediction module, the first attention mechanism module is used to focus on the first encoded feature currently participating in training.
[0063] Based on the output of the first attention mechanism module, the corresponding first predicted text is obtained using the first decoder.
[0064] Optionally, the recognition model further includes a third encoder; the method further includes:
[0065] Based on the first predicted text corresponding to multiple subsample data, the third encoder is used to encode and obtain the third encoded feature.
[0066] The step of obtaining the corresponding second encoded feature using the second encoder and obtaining the corresponding second predicted text using the second decoder based on the second sample data includes:
[0067] Based on the second sample data, the second encoder is used to encode and obtain the second encoded feature;
[0068] Based on the second encoding feature and the third encoding feature, the corresponding second predicted text is obtained using the second decoder.
[0069] Optionally, the recognition model further includes a capacity reduction module; the step of obtaining the second encoded features using the second encoder and obtaining the corresponding second predicted text using the second decoder based on the second sample data includes:
[0070] Using the aforementioned reduction module, feature concatenation and reduction processing are performed on the second sample data and the target sample data to obtain the second sample features;
[0071] The second sample features are used as input data for the second encoder, and the second encoder is used to obtain the second encoded features.
[0072] Based on the second encoded feature, the corresponding second predicted text is identified using the second decoder.
[0073] Fourthly, embodiments of this application provide a speech recognition method, including:
[0074] Real-time collection of user voice data;
[0075] The voice data is processed to obtain the first voice data of the first frame number;
[0076] The speech recognition model is used to extract the first speech feature of the first speech data and to identify the first subtitle corresponding to the first speech feature.
[0077] Display the first subtitle;
[0078] Accumulate multiple first speech features to obtain second speech data for the second frame number;
[0079] Based on the second speech data, the corresponding second text is identified using the speech recognition model.
[0080] Update the displayed subtitles using the second subtitle.
[0081] Fifthly, embodiments of this application provide a computing device, including a processing component and a storage component;
[0082] The storage component stores one or more computer instructions; the one or more computer instructions are invoked and executed by the processing component, the processing component being used to:
[0083] The real-time input target data is segmented to obtain the first data of the first capacity;
[0084] The recognition model is used to extract the first feature corresponding to the first data and to identify the first text corresponding to the first feature.
[0085] Output the first text;
[0086] Accumulate multiple first features to obtain second data of second capacity;
[0087] Based on the second data, the corresponding second text is identified using the recognition model;
[0088] Update the output text using the second text.
[0089] Sixthly, embodiments of this application provide an electronic device, including a processing component, a storage component, and a display component;
[0090] The storage component stores one or more computer instructions; the one or more computer instructions are invoked and executed by the processing component, the processing component being used to:
[0091] The real-time input voice data is processed to obtain the first voice data of the first frame.
[0092] The speech recognition model is used to extract the first speech feature of the first speech data and to identify the first text corresponding to the first speech feature.
[0093] Output the first text;
[0094] Accumulate multiple first speech features to obtain second speech data for the second frame number;
[0095] Based on the second speech data, the corresponding second text is identified using the speech recognition model.
[0096] Update the output text using the second text.
[0097] Seventhly, embodiments of this application provide a computing device, including a processing component and a storage component;
[0098] The storage component stores one or more computer instructions; the one or more computer instructions are invoked and executed by the processing component, the processing component being used to:
[0099] Determine the target sample data and corresponding target text for the second capacity;
[0100] The target sample data is segmented to obtain multiple sub-sample data of a first capacity;
[0101] The recognition model is trained by using multiple subsample data as model inputs and the target text as model labels.
[0102] In this embodiment, the real-time input speech data is segmented to obtain a first frame of speech data. A speech recognition model is then used to extract the first speech features corresponding to the first speech data and recognize the corresponding first text, which is then output. Simultaneously, multiple first speech features are accumulated to obtain a second frame of speech data. Based on the second speech data, the speech recognition model recognizes the corresponding second text, which is then used to update the output text. In this embodiment, by segmenting the real-time input speech data to obtain a smaller number of first speech data frames, speech recognition can be performed, and the corresponding first text can be output, reducing the recognition latency. Simultaneously, the second frame of speech data is accumulated, and the corresponding second text is recognized. The output result is then refreshed using the second text. Since the second speech data is obtained by accumulating multiple first speech features, its information completeness is greater than that of the first speech data, thus ensuring the accuracy of the final output result and guaranteeing recognition precision. The technical solution of this embodiment simultaneously considers latency and recognition accuracy.
[0103] These or other aspects of this application will become more apparent in the following description of the embodiments. Attached Figure Description
[0104] To more clearly illustrate the technical solutions in the embodiments of this application or the prior art, the drawings used in the description of the embodiments or the prior art will be briefly introduced below. Obviously, the drawings described below are some embodiments of this application. For those skilled in the art, other drawings can be obtained based on these drawings without creative effort.
[0105] Figure 1 A flowchart of one embodiment of a data processing method provided in this application is shown;
[0106] Figure 2 A flowchart of yet another embodiment of a data processing method provided in this application is shown;
[0107] Figure 3 A flowchart of yet another embodiment of a data processing method provided in this application is shown;
[0108] Figure 4A flowchart of yet another embodiment of a data processing method provided in this application is shown;
[0109] Figure 5a This paper illustrates a schematic diagram of the structure of the recognition model provided in an embodiment of this application in a practical application.
[0110] Figure 5b This illustration shows a model training diagram in a practical application based on an embodiment of this application;
[0111] Figure 6 A flowchart of yet another embodiment of a data processing method provided in this application is shown;
[0112] Figure 7 A flowchart of one embodiment of the speech recognition method provided in this application is shown;
[0113] Figure 8 This illustration shows a scenario interaction diagram of an embodiment of this application in a practical application;
[0114] Figure 9 This invention provides a schematic diagram of the structure of one embodiment of a data processing apparatus.
[0115] Figure 10 This application provides a schematic diagram illustrating the structure of one embodiment of a computing device.
[0116] Figure 11 This application provides a schematic diagram illustrating the structure of an embodiment of an electronic device.
[0117] Figure 12 This illustration shows a structural schematic diagram of yet another embodiment of a data processing apparatus provided in this application;
[0118] Figure 13 A schematic diagram of another embodiment of a computing device provided in this application is shown. Detailed Implementation
[0119] To enable those skilled in the art to better understand the present application, the technical solutions in the embodiments of the present application will be clearly and completely described below with reference to the accompanying drawings.
[0120] In some of the processes described in the specification, claims, and accompanying drawings of this application, multiple operations appearing in a specific order are included. However, it should be clearly understood that these operations may not be executed in the order they appear herein, or may be executed in parallel. The operation numbers, such as 101, 102, etc., are merely used to distinguish different operations and do not themselves represent any execution order. Furthermore, these processes may include more or fewer operations, and these operations may be executed sequentially or in parallel. It should be noted that the descriptions such as "first," "second," etc., in this document are used to distinguish different messages, devices, modules, etc., and do not represent a chronological order, nor do they limit "first" and "second" to different types.
[0121] The technical solutions of this application can be applied to application scenarios that convert data of the first category into data of the second category, such as machine translation scenarios, translating voice or text input in the source language into text of the target language, or speech recognition scenarios, recognizing voice data as text, etc., and are particularly suitable for real-time recognition scenarios.
[0122] Taking speech recognition as an example, real-time speech recognition, also known as streaming speech recognition, aims to return recognition results while the speaker is speaking. Current speech recognition methods mostly employ end-to-end approaches, pre-training end-to-end speech recognition models. Based on these models, input speech data can be directly output as corresponding text. In streaming recognition scenarios, due to requirements for latency and recognition accuracy, existing streaming speech recognition methods typically begin speech recognition immediately after acquiring a speech segment to ensure timely output of results. However, since speech segments contain incomplete information, this inevitably affects recognition accuracy. Waiting for complete speech data acquisition before recognition would inevitably increase latency.
[0123] To balance latency and recognition accuracy, the inventors, through a series of studies, proposed the technical solution of this application. In the embodiments of this application, the real-time input target data is segmented to obtain first data of a first capacity. Using a recognition model, first features of the first data are extracted, and the corresponding first recognition result is identified, which can then be output. Simultaneously, multiple first features are accumulated to obtain second data of a second capacity. Based on the second data, the recognition model identifies the corresponding second recognition result. The output result is then updated using the second recognition result. By segmenting the real-time input target data to obtain a smaller capacity of first data, recognition can be performed, and the corresponding first recognition result can be output, thereby reducing the recognition latency. Simultaneously, the second data of a second capacity is accumulated, and the corresponding second recognition result is identified. The output result is refreshed using the second recognition result. The second data contains more complete information than the first data, thus ensuring the accuracy of the final output result and guaranteeing recognition precision. The technical solution of this application uses the same model to simultaneously balance latency and recognition accuracy, ensuring both low latency and high recognition precision.
[0124] The technical solutions of the embodiments of this application will be clearly and completely described below with reference to the accompanying drawings. Obviously, the described embodiments are only some embodiments of this application, and not all embodiments. Based on the embodiments of this application, all other embodiments obtained by those skilled in the art without creative effort are within the scope of protection of this application.
[0125] Figure 1 A flowchart of one embodiment of a data processing method provided in this application is shown. The method may include the following steps:
[0126] 101: The target data input in real time is segmented to obtain the first data of the first capacity.
[0127] Real-time input target data can be understood as streaming input data, acquired during the streaming input process. For example, real-time input speech data can refer to speech segments acquired during speech acquisition.
[0128] Specifically, the real-time input target data can be segmented based on a first capacity and processed in chronological order. This can be achieved by using each accumulated data of the first capacity as the first data for subsequent operations, or by using a corresponding duration value for each accumulated data duration as the first data for subsequent operations.
[0129] Here, "first capacity" can refer to the number of unit data contained in the first data. For example, the first data can be a sequence of data composed of element feature data, and the first capacity can refer to the number of elements contained in the first data. For instance, if the target data is speech data and the first data is an audio frame sequence, the capacity can refer to the number of frames; or if the target data is text and the first data is assumed to be a text sequence composed of single characters, the capacity can refer to the number of characters.
[0130] To further improve recognition accuracy, the first data may include historical data and future data. Therefore, optionally, step 101 may specifically be:
[0131] The real-time input target data is segmented according to the chronological order to obtain the fifth data of the fourth capacity;
[0132] The sixth data of the first predetermined capacity after obtaining the fifth data, and the seventh data of the second predetermined capacity produced historically before the fifth data;
[0133] The first data consists of the fifth, sixth, and seventh data.
[0134] For example, when the target data is voice data, the capacity can specifically refer to the number of frames. The first data is assumed to be 20 frames, where the fourth capacity can be 10 frames. After obtaining 10 frames of voice data in real time, the most recent 5 frames of voice data can be obtained from the most recently obtained voice data, and then 5 more frames of voice data can be accumulated to form 20 frames of voice data before processing.
[0135] 102: Use the recognition model to extract the first feature corresponding to the first data and recognize the first recognition result corresponding to the first feature.
[0136] The target data can refer to data of a first category, and the first recognition result can be data of a second category, thus converting data of the first category into data of the second category. For example, in a speech recognition scenario, the target data can refer to speech data, and the first recognition result is the corresponding speech recognition result, i.e., text. In a machine recognition scenario, the target data can refer to the source language text, and the first recognition result is the target language text.
[0137] The first data can be sequence data of a first category, consisting of feature data of at least one element of the first category; the first recognition result is sequence data of a second category, consisting of feature data of at least one element of the second category. In a speech recognition scenario, the element of the first category can refer to an audio frame, and the element of the second category can refer to a text element, such as a word composed of a single character or multiple characters in the case of Chinese text; in a machine translation scenario, the element of the first category can refer to text elements of the source language, and the element of the second category can refer to text elements of the target language.
[0138] The first feature can have the same capacity as the first data, and the first feature can be obtained by transforming the feature data of each element in the first data.
[0139] 103: Output the first recognition result.
[0140] In this embodiment, the first data can be identified using the recognition model to obtain the corresponding first recognition result, which can then be output in a timely manner.
[0141] Since the first data will be continuously acquired and the first recognition results will be continuously output, multiple first recognition results will be arranged in chronological order.
[0142] 104: Accumulate multiple first features to obtain second data of second capacity.
[0143] Since the first data is continuously acquired, steps 102 and 103 are performed for each piece of first data. Therefore, first features are continuously generated. In this embodiment, these first features can be accumulated, and multiple first features form second data with a second capacity. These multiple first features can be arranged in chronological order to obtain the second data. The second capacity can be greater than the first capacity.
[0144] The first and second capacities can be set according to the actual application.
[0145] By accumulating multiple first features to obtain second data, the computational load of the recognition model can be reduced, thereby improving recognition efficiency.
[0146] 105: Based on the second data, use the recognition model to identify the corresponding second recognition result.
[0147] 106: Update the output results using the second recognition result.
[0148] Specifically, the second recognition result can be used to update the corresponding first recognition results in the already output results.
[0149] The multiple first identification results corresponding to the second identification result can refer to the first identification results corresponding to the multiple first features of the second data.
[0150] In this embodiment, the recognition model first identifies the first data and outputs the corresponding first recognition result. Then, the second data is identified to obtain the second recognition result. The output result is updated using the second recognition result. Since the first data has a smaller capacity, the recognition result can be output in real time. The second data contains more complete information than the first data, thus ensuring the accuracy of the final output result and guaranteeing recognition precision. This embodiment of the application simultaneously considers latency and recognition accuracy.
[0151] The recognition model can be pre-trained using a second volume of target sample data and the corresponding target recognition results. For example... Figure 2 The data processing method shown, from the perspective of model training, can be described as follows:
[0152] 201: Determine the target sample data of the second capacity and the corresponding target recognition results.
[0153] 202: The target sample data is segmented to obtain multiple sub-sample data of the first size.
[0154] Optionally, each subsample data of the second capacity may consist of a first subdata of the fourth capacity, a second subdata of the first predetermined capacity arranged before the first subdata, and a third subdata of the second predetermined capacity arranged after the first subdata.
[0155] 203: Train the recognition model by using multiple subsample data as model inputs and the target recognition results as model labels.
[0156] This recognition model can be used to identify target data input in real time, balancing latency and recognition accuracy. For details on the recognition process, please refer to [link / reference]. Figure 1 The embodiments shown will not be described in detail here.
[0157] The recognition model trained through the embodiments of this application can perform two recognitions, while ensuring both latency and recognition accuracy. Compared with the method of using two models for two recognitions, it can reduce the computational load and training workload of the model.
[0158] The target sample data and target recognition results can be sequence data. The recognition model is used to realize the sequence-to-sequence conversion. Optionally, the recognition model can be implemented using an Encoder-Decoder model framework to solve the sequence-to-sequence conversion problem.
[0159] The recognition model may include a first recognition network and a second recognition network, both of which can be Encoder-Decoder structures. The first recognition network in the trained recognition model is used to recognize the first recognition result corresponding to the first data, and the second recognition network is used to recognize the second recognition result corresponding to the second data.
[0160] The first recognition network may include a first encoder and a first decoder, and the second recognition network may include a second encoder and a second decoder. The first encoder and the second encoder may employ neural network models such as DFSMN (Deep Feed-Forward Sequential Memory Network), CNN (Convolutional Neural Networks), LSTM (Long Short-Term Memory), LSTM (bidirectional LSTM), and Transformer, etc., and this application does not impose specific limitations on them.
[0161] The first and second decoders can be implemented using, for example, DFSMN, CNN, LSTM, BLSTM, or Transformer.
[0162] During model training, the first encoder can encode multiple subsample data to obtain first encoded features, and the first decoder can identify the corresponding first prediction result based on the first encoded features. The second encoder can encode second sample data generated by accumulating the first encoded features of multiple subsample data to obtain second encoded features; the second decoder can identify the corresponding second prediction result based on the second encoded features. Then, based on the first prediction result, the second prediction result, and the target recognition result, the model parameters are adjusted using a loss function, thus achieving the training of the recognition model.
[0163] To further improve model accuracy, the first recognition network may also include a first prediction module and a first attention mechanism module. The first prediction module can be used to identify the number of elements in the second category corresponding to the data of the first category, and it can be obtained through pre-training. When training the recognition model, the number of elements in the second category corresponding to each subsample data can be pre-labeled. The first recognition network is trained sequentially using multiple subsample data, and the first prediction module can sequentially predict the number of elements in the second category corresponding to each of the multiple subsample data. The first attention mechanism module can determine the subsample data currently participating in training based on the prediction results of the first prediction module, and thus pay attention to this subsample data. Specifically, it can mask the first encoded features of other subsample data except the subsample data currently participating in training, and output only the first encoded features corresponding to the subsample data trained with the current parameters, so that the first decoder can identify the corresponding first prediction result, etc. The first recognition network introduces an attention mechanism during decoding, so that the trained first recognition network can achieve accurate recognition without complete information. In addition, the first attention mechanism module is also used to train and obtain the elements of interest in its input data, so as to assign different weight coefficients to different elements. During the model training process, the weight coefficients can be adjusted to further ensure recognition accuracy.
[0164] Optionally, the second recognition network may include a second prediction module and a second attention mechanism. The second prediction module is pre-trained and can predict the number of elements in the corresponding second category based on the second encoded features. The second attention mechanism can determine the elements of the first category that need attention in the second encoded features based on the prediction results of the second prediction module. It can then assign different weight coefficients to different elements in the second encoded features and adjust these weight coefficients during model training to achieve self-training. This enables the second recognition network to achieve accurate recognition.
[0165] Furthermore, in order to further improve the model recognition accuracy, the target sample data and the second sample data can be used together as the input data of the second recognition network to train the recognition model. Therefore, multiple sub-sample data and target sample data can be used as the model input of the recognition model and the target text can be used as the model label to train the recognition model.
[0166] Furthermore, the recognition model may also include a scaling-down module for reducing the size and dimensionality of the second and target sample data to decrease the computational load. The scaling-down module can first perform feature concatenation on the second and target sample data before scaling down. For example, the second sample data might contain 100 frames of 50-dimensional feature vectors, and the target sample data might also contain 100 frames of 50-dimensional feature vectors. After feature concatenation, 100 frames of 100-dimensional feature vectors can be obtained, and then scaling down can yield 50 frames of 50-dimensional feature vectors. This scaling-down module can consist of one or more neural network layers, such as one or more convolutional neural network layers.
[0167] Furthermore, the recognition model may also include a third encoder, which can encode third coded features based on the first prediction results corresponding to multiple subsample data; the second decoder can predict the second predicted text based on the third coded features and the second decoded features. This third encoder can be implemented using DFSMN, CNN, BLSTM, Transformer, etc.
[0168] In practical applications, the technical solutions of this application can be applied to recognition scenarios such as machine translation or speech recognition. The target data input in real time can refer to speech data or text, which can be converted and processed using a recognition model to obtain the corresponding text. In one or more embodiments below, the technical solutions of this application are mainly described using a recognition model to recognize text as an example.
[0169] Figure 3 A flowchart illustrating another embodiment of a data processing method provided in this application is provided. The method may include the following steps:
[0170] 301: The target data input in real time is segmented to obtain the first data of the first capacity.
[0171] Alternatively, the first data can be obtained in the following specific manner:
[0172] The real-time input target data is segmented according to the chronological order to obtain the fifth data of the fourth capacity;
[0173] The sixth data of the first predetermined capacity after obtaining the fifth data, and the seventh data of the second predetermined capacity produced historically before the fifth data;
[0174] The first data consists of the fifth, sixth, and seventh data.
[0175] 302: Use the recognition model to extract the first feature of the first data and recognize the first text corresponding to the first feature.
[0176] In speech recognition scenarios, the target data can be speech data, while in machine translation scenarios, the target data can be speech data or text, etc.
[0177] 303: Output the first text.
[0178] The first text may include at least one text element.
[0179] 304: Accumulate multiple first features to obtain second data of a second capacity;
[0180] 305: Based on the second data, use the recognition model to identify the corresponding second text.
[0181] 306: Update the output text using the second text.
[0182] Specifically, this can be achieved by updating multiple corresponding first texts using the second text. The multiple first texts corresponding to the second text are also the first texts corresponding to the multiple first features of the second data.
[0183] In this embodiment, by segmenting the real-time input target data to obtain a smaller volume of first data, recognition can be performed, and the corresponding first text can be output to reduce the recognition latency. Simultaneously, a second volume of second data is accumulated, and the corresponding second text is recognized. The output result is refreshed using this second text. Since the second data contains more complete information than the first data, the accuracy of the final output result can be guaranteed, ensuring recognition precision. The technical solution of this application embodiment simultaneously considers latency and recognition accuracy, ensuring both low latency and high recognition precision.
[0184] The recognition model can be trained based on a second volume of target sample data and the corresponding target text. For example... Figure 4 The diagram shown is a flowchart of a data processing method provided in another embodiment of this application. The technical solution of this application is described from the perspective of model training. The method may include the following steps:
[0185] 401: Determine the target sample data and corresponding target text for the second capacity.
[0186] 402: The target sample data is segmented to obtain multiple sub-sample data of the first size.
[0187] 403: Train the recognition model by using multiple subsample data as model inputs and the target text as model labels.
[0188] Specifically, it can be based on multiple subsample data to obtain corresponding first coding features and corresponding first predicted text based on multiple first coding features; and it can accumulate second sample data generated by the first coding features corresponding to multiple subsample data; based on the second sample data, obtain second coding features and corresponding second predicted text based on the second coding features; and then train the recognition model based on the first predicted text, the second predicted text, and the target text.
[0189] This recognition model can be used to identify target data input in real time, balancing latency and recognition accuracy. For details of the recognition process, please refer to [link / reference]. Figure 3 The embodiments shown will not be described in detail here.
[0190] The recognition model trained through the embodiments of this application can perform two recognitions, while ensuring both latency and recognition accuracy. Compared with the method of using two models for two recognitions, it can reduce the computational load and training workload of the model.
[0191] Both the target sample data and the target text can be sequence data. The recognition model realizes the sequence-to-sequence conversion. This recognition model can be implemented using an Encoder-Decoder structure to solve the sequence-to-sequence conversion problem.
[0192] In some embodiments, the recognition model may include a first recognition network and a second recognition network. During model training, multiple sub-sample data are used as input to the recognition model, and the target text is used as the model label. Training the recognition model may include:
[0193] Multiple subsample data are input into the first recognition network, the first recognition network is used to obtain the first coding features corresponding to the multiple subsample data respectively, and the corresponding first predicted text is predicted based on the first coding features.
[0194] The first encoded features corresponding to multiple subsample data are accumulated to generate the second sample data;
[0195] The second sample data is input into the second recognition network, the second recognition network is used to obtain the second coding features, and the corresponding second predicted text is predicted based on the second coding features;
[0196] The recognition network is trained based on the first predicted text, the second predicted text, and the target text. Specifically, a loss function can be used to calculate the loss information between the first predicted text, the second predicted text, and the target text, respectively. The model parameters of the recognition model are then adjusted based on the loss information to train the recognition model.
[0197] When identifying data, the process of extracting the first feature of the first data and identifying the first text corresponding to the first feature using the identification model can be as follows: Based on the first data, the first feature is obtained and the first text corresponding to the first feature is identified using the first identification network in the identification model.
[0198] Based on the second data, the recognition model can be used to identify the corresponding second text, which can be done by: using the second recognition network to obtain the second feature and identify the second text corresponding to the second feature.
[0199] The first recognition network may include a first encoder and a first decoder; the second recognition network may include a second encoder and a second decoder; for ease of understanding, Figure 5a The diagram illustrates a model structure for a practical application. The recognition model may include a first encoder 501, a first decoder 502, a second encoder 503, and a second decoder 504. The first encoder, second encoder, first decoder, and second decoder may be implemented using technologies such as DFSMN, CNN, BLSTM, or Transformer.
[0200] During model training, multiple subsample data are used as the model input of the recognition model and the target text is used as the model label. The training of the recognition model can be as follows: based on multiple subsample data, a first encoder is used to obtain a first encoded feature and a first decoder is used to obtain the corresponding first predicted text; second sample data is generated by accumulating the first encoded features corresponding to the multiple subsample data; based on the second sample data, a second encoder is used to obtain a second encoded feature and a second decoder is used to obtain the corresponding second predicted text; and the recognition model is trained based on the first predicted text, the second predicted text, and the target text.
[0201] When recognizing data, extracting the first feature of the first data and recognizing the first text corresponding to the first feature using the recognition model can be: based on the first data, encoding the first feature using the first encoder, and based on the first feature, recognizing the corresponding first text using the first decoder;
[0202] The recognition of the corresponding second text based on the second data can be achieved by: using the second encoder to encode the second feature based on the second data, and using the second decoder to recognize the corresponding second text based on the second feature.
[0203] like Figure 5a As shown, the first decoder 502 can output the first text in real time, and the second text output by the second decoder 504 can update the text sequence consisting of multiple first texts in the already output text, which can replace the multiple first texts corresponding to it with the second text. The input of the first encoder 501 can be the first data obtained by segmentation.
[0204] To further improve the model's recognition accuracy, in some embodiments, during model training, the first predicted text obtained by the second decoder can be accumulated, and the corresponding second predicted text can be obtained based on the first predicted text corresponding to the second sample data and multiple subsample data respectively.
[0205] Optionally, the recognition model may also include a third encoder, such as Figure 5a The third encoder 505 in a; the third encoder can be implemented using DFSMN, CNN, BLSTM, or Transformer, etc.
[0206] During model training, the method may also include: obtaining third encoded features by encoding the first predicted text corresponding to multiple subsample data respectively using a third encoder;
[0207] The process of obtaining the corresponding second encoded features using the second encoder and the corresponding second predicted text using the second decoder based on the second sample data can include: obtaining the second encoded features using the second encoder based on the second sample data.
[0208] Based on the second and third coding features, the corresponding second predicted text is obtained using the second decoder.
[0209] In this context, the first predicted text corresponding to each of the multiple subsample data is also the first predicted text corresponding to each of the multiple first encoded features.
[0210] During data recognition, multiple first texts corresponding to the first features can be accumulated to obtain a third text; this third text can be obtained by arranging the first texts corresponding to the multiple first features in chronological order.
[0211] Based on the second data, the corresponding second text can be identified using the recognition model.
[0212] Based on the second data and the third text, the corresponding second text is identified using a recognition model. Specifically, this can be done as follows: based on the second data, a second encoder is used to obtain a second feature; based on the third text, a third encoder is used to obtain a third feature; based on the second feature and the third feature, a second decoder is used to identify the corresponding second text.
[0213] Furthermore, in some embodiments, target sample data can be used as model input for model training. During model training, multiple sub-sample data and the target sample data can be used as model inputs to the recognition model, and the target text can be used as the model label to train the recognition model. Specifically, based on multiple sub-sample data, a first encoder can be used to obtain the respective first encoded features of each sub-sample data; the first encoded features corresponding to the multiple sub-sample data can be accumulated to generate second sample data.
[0214] The second encoder can obtain the second feature based specifically on the target sample data and the second sample data.
[0215] During data recognition, multiple sets of first data can be accumulated to obtain a third set of data with a second capacity. Then, based on the second data, the recognition model can be used to identify the corresponding second text. Specifically, this can be achieved by using a second encoder to encode and obtain the second feature based on the second and third data.
[0216] To reduce the computational load of the model and improve recognition efficiency, in some embodiments, the recognition model may also include a scaling-down module, such as... Figure 5a The scaling-down module 506 shown can perform feature concatenation and scaling-down processing on the second sample data and the target sample data to obtain second sample features; and obtain second encoded features based on the second sample features. Alternatively, the second sample features can be used as input data for the second encoder to obtain the second encoded features.
[0217] During data recognition, the recognition of the corresponding second text based on the second and third data may include: performing feature concatenation and capacity reduction on the second and third data to obtain a fourth data of a third capacity; and using the recognition model to recognize the corresponding second text based on the fourth data.
[0218] Specifically, this could involve using a second encoder to extract the second feature corresponding to the fourth data; and then using a second decoder to identify the corresponding second text based on the second feature. For example... Figure 5a As shown, the input of the second encoder 503 is specifically the fourth data. The second feature obtained by encoding is used to input the second decoder 504. At the same time, the third feature obtained by encoding the accumulated third text through the third encoder 505 is also input to the second decoder 504. The second encoder 503 converts the second feature and the third feature to obtain the corresponding second text.
[0219] To further improve the model's recognition accuracy, in some embodiments, the first recognition network in the recognition model may further include a first prediction module and a first attention mechanism module; during model training, based on multiple sub-sample data, obtaining first encoded features using a first encoder and obtaining the corresponding first predicted text using a first decoder includes:
[0220] Based on multiple sub-sample data, the first encoder is used to extract the first encoded features corresponding to the multiple sub-sample data respectively;
[0221] The first prediction module is used to predict the number of text elements corresponding to multiple first coding features in sequence.
[0222] Based on the prediction results of the first prediction module, the first attention mechanism module is used to focus on the first encoded feature currently participating in training;
[0223] Based on the output of the first attention mechanism module, the corresponding first predicted text is obtained using the first decoder.
[0224] Using multiple subsample data as input data for the first recognition network to obtain the first encoded features based on the first encoder and the corresponding first predicted text based on the first decoder may include:
[0225] An attention mechanism is used to train the recognition model, enabling the first decoder to accurately recognize text without requiring complete information.
[0226] For ease of understanding, Figure 5b The diagram shows a schematic of the model training for the first recognition network, as shown below. Figure 5b As shown, assuming there are M subsample data x1, x2...xM, where M is greater than 1, the M subsample data are encoded by the first encoder 501 to obtain the corresponding M first features c1, c2...cM. The M first features are used by the first prediction module 507 and the first attention mechanism module 508 to determine the first feature that needs to be focused on. The first feature that needs to be focused on is then transformed by the first decoder 502. The target text is used as the model label of the first decoder 502 to train the first recognition network.
[0227] Furthermore, the second recognition network in the recognition model can also include a second prediction module and a second attention mechanism module. The second prediction module is pre-trained and can predict the number of corresponding text elements based on the second encoded features. The second attention mechanism can determine the elements that need attention in the second encoded features based on the prediction results of the second prediction module, and can assign different weight coefficients to different elements in the second encoded features accordingly. The weight coefficients can be adjusted during model training to achieve self-training. Thus, accurate recognition by the second recognition network can be achieved.
[0228] During data recognition, the fourth encoding feature can be obtained by recalculating the second encoding feature based on the weight coefficient of the second attention mechanism module. Specifically, the fourth encoding feature and the third encoding feature are input into the second decoder to obtain the corresponding second text.
[0229] Furthermore, the first decoder and the second decoder can adopt an autoregressive approach. The input data of the first decoder and the second decoder can also include their respective historical output data. For example, the input data of the first decoder can also include the historical output first text, and the input data of the second decoder can also include the historical output second text, etc. This application will not elaborate further on the autoregressive approach.
[0230] As yet another embodiment, this application also provides a data processing method, such as... Figure 6 As shown, the method may include the following steps:
[0231] 601: The target data input in real time is segmented to obtain the first data of the first capacity.
[0232] 602: Use the recognition model to identify the first text corresponding to the first data.
[0233] In speech recognition scenarios, the target data can be speech data, while in machine translation scenarios, the target data can be speech data or text, etc.
[0234] 603: Output the first text.
[0235] The first text may include at least one text element.
[0236] 604: Accumulate multiple first data points to obtain the third data point of the second capacity.
[0237] 605: Based on the third data, the corresponding second text is identified using a recognition model.
[0238] 606: Update the output text using the second text.
[0239] That is, the embodiments of this application may also identify the corresponding second text based solely on the third data obtained by accumulating multiple second data.
[0240] The corresponding recognition model can be trained as follows:
[0241] Determine the target sample data and corresponding target text for the second capacity;
[0242] The target sample data is segmented to obtain multiple sub-sample data of the first size;
[0243] The recognition model is trained by using multiple subsample data and target sample data as inputs and the target text as the model label.
[0244] When the recognition model consists of a first recognition network and a second recognition network, multiple sub-sample data serve as input data for the first recognition network, while the target sample data serves as input data for the second recognition network.
[0245] Other identical or corresponding operations in this embodiment can be found in detail in [link to documentation]. Figure 3 As described in the embodiments, it will not be repeated here.
[0246] In a practical application, the technical solution of this application can be applied to streaming speech recognition scenarios. The following section uses a streaming speech recognition scenario as an example to introduce the technical solution of this application. Figure 7 As shown in the illustration, this application embodiment also provides a speech recognition method, which may include the following steps:
[0247] 701: Process the real-time input voice data to obtain the first voice data of the first frame.
[0248] The voice data can be acquired in real time, or it can be acquired and transmitted in real time by a data acquisition device.
[0249] 702: Use a speech recognition model to extract the first speech feature of the first speech data and recognize the first text corresponding to the first speech feature.
[0250] 703: Output the first text.
[0251] 704: Accumulate multiple first speech features to obtain second speech data for the second frame number.
[0252] 705: Based on the second speech data, use a speech recognition model to recognize the corresponding second text.
[0253] 706: Update the output text using the second text.
[0254] In some embodiments, the method may further include:
[0255] Accumulate multiple first-speech data points to obtain third-speech data for the second frame number;
[0256] Based on the second speech data, recognizing the corresponding second text using the speech recognition model may include:
[0257] Based on the second and third speech data, the corresponding second text is identified using the speech recognition model.
[0258] In some embodiments, recognizing the corresponding second text using the speech recognition model based on the second speech data and the third speech data may include:
[0259] The second and third speech data are subjected to feature splicing and capacity reduction processing to obtain fourth speech data of a third capacity.
[0260] Based on the fourth speech data, the corresponding second text is identified using the speech recognition model.
[0261] In some embodiments, the method may further include:
[0262] By accumulating the first text corresponding to the multiple first speech features, a third text is obtained;
[0263] Based on the second speech data, the corresponding second text is identified using the speech recognition model, including:
[0264] Based on the second speech data and the third text, the corresponding second text is identified using the speech recognition model.
[0265] In some embodiments, segmenting the real-time target data to obtain the first audio data for the first frame number may include:
[0266] The real-time input voice data is segmented according to the chronological order to obtain the fifth voice data of the fourth frame.
[0267] Acquire the sixth voice data after the fifth voice data for a first predetermined number of frames, and the seventh voice data that was historically generated before the fifth voice data for a second predetermined number of frames;
[0268] The first voice data is composed of the fifth voice data, the sixth voice data, and the seventh voice data.
[0269] This embodiment and Figure 3 The difference in the illustrated embodiment is that the target data is specifically speech data, both the first data and the second data are speech data, and both the first feature and the second feature are speech features. Other identical or corresponding operations can be found in the preceding embodiments and will not be repeated here.
[0270] It should be noted that in speech recognition scenarios, for ease of speech recognition, the real-time input speech data is a speech signal. The first speech data can be a feature vector obtained by segmenting the speech signal into frames and performing feature extraction and transformation, consisting of feature vectors from several audio frames in the first frame. The first speech features extracted from the first speech data can be further processed from the feature vectors of the audio frames to obtain higher-level feature vectors, such as semantic feature vectors obtained by encoder encoding, etc. This application does not limit the feature processing operations in this regard.
[0271] In this context, accumulating multiple first speech features means concatenating multiple first speech features, and the second frame number can be the sum of the frame numbers of each of the multiple first speech features, etc.
[0272] The first text can consist of a single character or multiple characters. When the first text contains multiple characters, the first text can be output character by character or simultaneously.
[0273] Accordingly, the speech recognition model can be pre-trained as follows:
[0274] Determine the speech sample data and corresponding target text for the second frame;
[0275] The speech sample data is segmented to obtain multiple sub-sample data of the first frame number;
[0276] The speech recognition model is trained by using multiple subsample data as inputs and the target text as the model label.
[0277] In practical applications, the speech recognition method of this application embodiment can be applied to various scenarios, such as human-computer interaction, intelligent question answering, voice assistants, meetings, live streaming, and other scenarios where subtitle output is required. The technical solution of this application embodiment can be executed by the server, outputting the first text and updating the output result using the second text. This can be achieved by the server sending the first text to the corresponding client for display, or by the server sending the second text to the client, instructing the client to update its display result using the second text. Alternatively, it can be executed by the client.
[0278] Furthermore, in local all-in-one machine scenarios where voice recognition is required, the technology of this application embodiment can also be executed by the local all-in-one machine. The local all-in-one machine can realize recognition and output processing, and can achieve voice acquisition, recognition, and real-time output of recognized text without network connection. Local all-in-one machines are widely used in scenarios such as banks and courts to provide users with self-service.
[0279] Taking subtitle output as an example, this application also provides a speech recognition method, which may include:
[0280] Real-time collection of user voice data;
[0281] The voice data is processed to obtain the first voice data of the first frame number;
[0282] The speech recognition model is used to extract the first speech feature of the first speech data and to identify the first subtitle corresponding to the first speech feature.
[0283] Display the first subtitle;
[0284] Accumulate multiple first speech features to obtain second speech data for the second frame number;
[0285] Based on the second speech data, the corresponding second text is identified using the speech recognition model.
[0286] Update the displayed subtitles using the second subtitle.
[0287] The first and second texts, i.e., the subtitles, are displayed to allow the user to speak while the content of their speech is displayed in the form of subtitles.
[0288] Taking the display of subtitles corresponding to user voice data on a local all-in-one machine as an example, such as Figure 8 In a speech recognition scenario shown, the local all-in-one machine 801 can collect user speech data in real time. Using the recognition model 802 provided by the technical solution of this application, the first subtitle can be displayed in a timely manner, and the second subtitle can be used to update the first subtitle to ensure the accuracy of the display results and improve the recognition accuracy. This can provide users with low-latency, high-efficiency and high-accuracy recognition results.
[0289] Furthermore, the technical solutions of this application embodiment can also be applied to machine translation scenarios. The target data input in real time can also be voice data or text input in real time using the source language, which is used to convert and process the target voice into text, etc. For specific implementation operations, please refer to the corresponding embodiments above, which will not be repeated here.
[0290] Figure 9 This application provides a schematic diagram of the structure of a data processing apparatus according to one embodiment. The apparatus may include:
[0291] The first processing module 901 is used to segment the real-time input target data to obtain first data of a first capacity;
[0292] The first recognition module 902 is used to extract the first feature corresponding to the first data and recognize the first text corresponding to the first feature using a recognition model;
[0293] The first output module 903 is used to output the first text;
[0294] The second processing module 904 is used to accumulate multiple first features to obtain second data of a second capacity;
[0295] The second recognition module 905 is used to recognize the corresponding second text based on the second data using a recognition model.
[0296] The second output module 906 updates the already output text using the second text.
[0297] In some embodiments, the device may further include:
[0298] The third processing module accumulates multiple sets of first data to obtain a second set of third data.
[0299] The second recognition module is specifically used to recognize the corresponding second text based on the second and third data using a recognition model.
[0300] In some embodiments, the second recognition module is specifically used to perform feature splicing and capacity reduction processing on the second and third data to obtain fourth data of a third capacity; based on the fourth data, the corresponding second text is recognized using a recognition model.
[0301] In some embodiments, the device may further include:
[0302] The fourth processing module is used to accumulate the first text corresponding to multiple first features to obtain the third text;
[0303] The second recognition module is specifically used to recognize the corresponding second text based on the second data and the third text using a recognition model.
[0304] In some embodiments, the first processing module may be specifically used to segment the real-time input target data in chronological order to obtain the fifth data of the fourth capacity; obtain the sixth data of the first predetermined capacity after the fifth data, and the seventh data of the second predetermined capacity generated historically before the fifth data; and the first data is composed of the fifth data, the sixth data, and the seventh data.
[0305] In a speech recognition scenario, the first processing module can be specifically used to process the real-time input speech data to obtain the first speech data of the first frame number;
[0306] The first recognition module is specifically used to extract the first speech features of the first speech data using a speech recognition model and to recognize the first text corresponding to the first speech features, and then output the first text.
[0307] The second processing module is specifically used to accumulate multiple first speech features to obtain second speech data for the second frame number.
[0308] The second recognition module is specifically used to recognize the corresponding second text based on the second speech data using a speech recognition model.
[0309] In applications where text is displayed as subtitles, such as in conferences or live streaming, the first and second texts are specifically subtitles used to display on the interface.
[0310] Figure 9 The data processing device can perform Figure 3 The implementation principle and technical effects of the data processing method described in the illustrated embodiments will not be repeated here. The specific methods by which each module and unit of the data processing device in the above embodiments performs its operations have been described in detail in the embodiments related to the method, and will not be elaborated upon here.
[0311] In one possible design, Figure 9 The data processing apparatus of the illustrated embodiment can be implemented as a computing device, which in practical applications can be a server, such as... Figure 10 As shown, the computing device may include a storage component 1001 and a processing component 1002;
[0312] Storage component 1001 stores one or more computer instructions, wherein one or more computer instructions are invoked and executed by processing component 1002.
[0313] Processing component 1002 is used for:
[0314] The real-time input target data is segmented to obtain the first data of the first capacity;
[0315] The recognition model is used to extract the first feature corresponding to the first data and to identify the first text corresponding to the first feature.
[0316] Output the first text;
[0317] Accumulate multiple first features to obtain second data of second capacity;
[0318] Based on the second data, the corresponding second text is identified using a recognition model;
[0319] Update the output text using the second text.
[0320] Optionally, the processing component 1002 can be specifically implemented as follows: Figure 3 The data processing method shown.
[0321] Of course, computing devices may also include other components, such as input / output interfaces and communication components.
[0322] Input / output interfaces provide interfaces between processing components and peripheral interface modules, which can be output devices, input devices, etc.
[0323] The communication components are configured to facilitate wired or wireless communication between computing devices and other devices.
[0324] The computing device can be a physical device or an elastic computing host provided by a cloud computing platform. In this case, the computing device can refer to a cloud server, and the aforementioned processing components, storage components, etc., can be basic server resources rented or purchased from the cloud computing platform.
[0325] In yet another possible design, Figure 9 The data processing device in the illustrated embodiment can be implemented as an electronic device. In practical applications, this electronic device can be, for example, various mobile terminals or local all-in-one machines with voice recognition requirements, such as... Figure 11 As shown, the electronic device may include a storage component 1101, a display component 1102, and a processing component 1103;
[0326] Storage component 1101 stores one or more computer instructions, wherein one or more computer instructions are invoked and executed by processing component 1103.
[0327] Processing component 1103 is used for:
[0328] The real-time input target data is segmented to obtain the first data of the first capacity;
[0329] The recognition model is used to extract the first feature corresponding to the first data and to identify the first text corresponding to the first feature.
[0330] Display the first text in display component 1102;
[0331] Accumulate multiple first features to obtain second data of second capacity;
[0332] Based on the second data, the corresponding second text is identified using a recognition model;
[0333] The displayed text in component 1102 is updated using the second text.
[0334] In a speech recognition scenario, the processing component 1103 is specifically used for:
[0335] The real-time input voice data is processed to obtain the first voice data of the first frame.
[0336] The speech recognition model is used to extract the first speech feature of the first speech data and to identify the first text corresponding to the first speech feature.
[0337] The first text is displayed in the display component 1102;
[0338] Accumulate multiple first speech features to obtain second speech data for the second frame number;
[0339] Based on the second speech data, the corresponding second text is identified using the speech recognition model.
[0340] The displayed text in the display component 1102 is updated using the second text.
[0341] Of course, electronic devices may also include other components, such as input / output interfaces and communication components. Input / output interfaces provide an interface between processing components and peripheral interface modules, which can be output devices, input devices, etc. Communication components are configured to facilitate wired or wireless communication between the electronic device and other devices.
[0342] This application also provides a computer-readable storage medium storing a computer program, which, when executed by a computer, can perform the above-described functions. Figure 3 The data processing method of the embodiment shown.
[0343] Figure 12 This application provides a schematic diagram of the structure of another embodiment of a data processing apparatus, which may include:
[0344] Data acquisition module 1201 is used to determine the target sample data of the second capacity and the corresponding target text;
[0345] The data processing module 1202 is used to segment the target sample data to obtain multiple sub-sample data of the first capacity;
[0346] The model training module 1203 is used to train the recognition model by taking multiple subsample data as model inputs and the target text as model labels.
[0347] In some embodiments, the recognition model may include a first encoder and a first decoder, as well as a second encoder and a second decoder;
[0348] The module training module can be specifically used to obtain first encoded features using a first encoder and obtain corresponding first predicted text using a first decoder based on multiple subsample data; determine second sample data generated by accumulating the first encoded features corresponding to the multiple subsample data; obtain second encoded features using a second encoder and obtain corresponding second predicted text using a second decoder based on the second sample data; and train a recognition model based on the first predicted text, the second predicted text, and the target text.
[0349] In some embodiments, the recognition model further includes a first prediction module and a first attention mechanism module;
[0350] The model training module, based on multiple sub-sample data, uses a first encoder to obtain first encoded features and a first decoder to obtain the corresponding first predicted text. This can include: extracting the first encoded features corresponding to each of the multiple sub-sample data using the first encoder; predicting the number of text elements corresponding to the multiple first encoded features sequentially using the first prediction module; focusing on the currently participating first encoded feature using a first attention mechanism module based on the prediction results of the first prediction module; and obtaining the corresponding first predicted text using the first decoder based on the output results of the first attention mechanism module.
[0351] In some embodiments, the recognition model may further include a third encoder;
[0352] The data processing module is also used to obtain third encoded features by encoding the first predicted text corresponding to multiple subsample data using the third encoder.
[0353] The model training process, which involves obtaining the second encoded features based on the second sample data and the corresponding second predicted text using the second encoder and the second decoder, may include: obtaining the second encoded features based on the second sample data and the second encoder; and obtaining the corresponding second predicted text based on the second encoded features and the third encoded features using the second decoder.
[0354] In some embodiments, the recognition model may further include a scaling down module; model training to obtain second encoded features based on second sample data and obtaining the corresponding second predicted text using a second encoder and a second decoder may include: using the scaling down module to perform feature concatenation and scaling down on the second sample data and the target sample data to obtain the second sample features;
[0355] The second sample features are used as input data for the second encoder, and the second encoder is used to obtain the second encoded features.
[0356] Based on the second encoded features, the second decoder is used to identify the corresponding second predicted text.
[0357] In a speech recognition scenario, the data acquisition module specifically determines the speech sample data of the second frame and the corresponding target text.
[0358] In one possible design, Figure 12 The data processing apparatus of the illustrated embodiment can be implemented as a computing device, which in practical applications can be a server, such as... Figure 13 As shown, the computing device may include a storage component 1301 and a processing component 1302;
[0359] Storage component 1301 stores one or more computer instructions, wherein one or more computer instructions are invoked and executed by processing component 1302.
[0360] Processing component 1302 is used for:
[0361] Determine the target sample data and corresponding target text for the second capacity;
[0362] The target sample data is segmented to obtain multiple sub-sample data of the first size;
[0363] The recognition model is trained by using multiple subsample data as inputs and the target text as the model label.
[0364] Optionally, the processing component 1302 can specifically implement, as follows: Figure 4 The data processing method described above.
[0365] Of course, computing devices may also include other components, such as input / output interfaces and communication components.
[0366] In a speech recognition scenario, the processing component 1302 is specifically used for:
[0367] Determine the speech sample data and corresponding target text for the second frame;
[0368] The speech sample data is segmented to obtain multiple sub-sample data of the first frame number;
[0369] The speech recognition model is trained by using multiple subsample data as inputs and the target text as the model label.
[0370] Input / output interfaces provide interfaces between processing components and peripheral interface modules, which can be output devices, input devices, etc.
[0371] The communication components are configured to facilitate wired or wireless communication between computing devices and other devices.
[0372] The computing device can be a physical device or an elastic computing host provided by a cloud computing platform. In this case, the computing device can refer to a cloud server, and the aforementioned processing components, storage components, etc., can be basic server resources rented or purchased from the cloud computing platform.
[0373] This application also provides a computer-readable storage medium storing a computer program, which, when executed by a computer, can perform the above-described functions. Figure 4 The data processing method of the embodiment shown.
[0374] The processing components involved in the corresponding embodiments described above may include one or more processors to execute computer instructions to complete all or part of the steps in the methods described above. Alternatively, the processing components may be implemented as one or more application-specific integrated circuits (ASICs), digital signal processors (DSPs), digital signal processing devices (DSPDs), programmable logic devices (PLDs), field-programmable gate arrays (FPGAs), controllers, microcontrollers, microprocessors, or other electronic components to perform the methods described above.
[0375] Storage components are configured to store various types of data to support operations on the terminal. Storage components can be implemented by any type of volatile or non-volatile storage device or a combination thereof, such as static random access memory (SRAM), electrically erasable programmable read-only memory (EEPROM), erasable programmable read-only memory (EPROM), programmable read-only memory (PROM), read-only memory (ROM), magnetic storage, flash memory, magnetic disk, or optical disk.
[0376] The display component can be an electroluminescent (EL) element, a liquid crystal display or a microdisplay with a similar structure, or a retina-direct display or a similar laser scanning display.
[0377] Those skilled in the art will clearly understand that, for the sake of convenience and brevity, the specific working processes of the systems, devices, and units described above can be referred to the corresponding processes in the foregoing method embodiments, and will not be repeated here.
[0378] The device embodiments described above are merely illustrative. The units described as separate components may or may not be physically separate. The components shown as units may or may not be physical units; that is, they may be located in one place or distributed across multiple network units. Some or all of the modules can be selected to achieve the purpose of this embodiment according to actual needs. Those skilled in the art can understand and implement this without any creative effort.
[0379] Through the above description of the embodiments, those skilled in the art can clearly understand that each embodiment can be implemented by means of software plus necessary general-purpose hardware platforms, and of course, it can also be implemented by hardware. Based on this understanding, the above technical solutions, in essence or the part that contributes to the prior art, can be embodied in the form of a software product. This computer software product can be stored in a computer-readable storage medium, such as ROM / RAM, magnetic disk, optical disk, etc., and includes several instructions to cause a computer device (which may be a personal computer, server, or network device, etc.) to execute the methods described in the various embodiments or some parts of the embodiments.
[0380] Finally, it should be noted that the above embodiments are only used to illustrate the technical solutions of this application, and are not intended to limit them. Although this application has been described in detail with reference to the foregoing embodiments, those skilled in the art should understand that modifications can still be made to the technical solutions described in the foregoing embodiments, or equivalent substitutions can be made to some of the technical features. Such modifications or substitutions do not cause the essence of the corresponding technical solutions to deviate from the spirit and scope of the technical solutions of the embodiments of this application.
Claims
1. A speech recognition method, characterized in that, include: The real-time input voice data is segmented to obtain the first voice data of the first frame. The speech recognition model is used to extract the first speech feature of the first speech data and to identify the first text corresponding to the first speech feature. Output the first text; By concatenating and accumulating multiple first speech features, a second speech data with a second frame number is obtained, wherein the second frame number is the sum of the frame numbers of each of the multiple first speech features; Based on the second speech data, the corresponding second text is identified using the speech recognition model. Update the output text using the second text.
2. The method according to claim 1, characterized in that, Also includes: Accumulate multiple first-speech data points to obtain third-speech data for the second frame number; The step of recognizing the corresponding second text based on the second speech data using the speech recognition model includes: Based on the second and third speech data, the corresponding second text is identified using the speech recognition model.
3. The method according to claim 2, characterized in that, The step of recognizing the corresponding second text using the speech recognition model based on the second speech data and the third speech data includes: The second and third speech data are subjected to feature splicing and capacity reduction processing to obtain fourth speech data of a third capacity. Based on the fourth speech data, the corresponding second text is identified using the speech recognition model.
4. The method according to claim 1, characterized in that, Also includes: By accumulating the first text corresponding to the multiple first speech features, a third text is obtained; The step of recognizing the corresponding second text based on the second speech data using the speech recognition model includes: Based on the second speech data and the third text, the corresponding second text is identified using the speech recognition model.
5. The method according to claim 1, characterized in that, The step of segmenting the real-time target data to obtain the first audio data of the first frame number includes: The real-time input voice data is segmented according to the chronological order to obtain the fifth voice data of the fourth frame. Acquire the sixth voice data after the fifth voice data for a first predetermined number of frames, and the seventh voice data that was historically generated before the fifth voice data for a second predetermined number of frames; The first voice data is composed of the fifth voice data, the sixth voice data, and the seventh voice data.
6. A data processing method, characterized in that, include: The real-time input target data is segmented to obtain first data of a first capacity, wherein the target data is streaming input data; The recognition model is used to extract the first feature corresponding to the first data and to identify the first text corresponding to the first feature. Output the first text; The second data with a second capacity is obtained by concatenating and accumulating multiple first features, wherein the second capacity is the sum of the capacities of the multiple first features; Based on the second data, the corresponding second text is identified using the recognition model; Update the output text using the second text.
7. A data processing method, characterized in that, include: Determine the target sample data of the second capacity and the corresponding target text, wherein the target sample data is streamed input data; The target sample data is segmented to obtain multiple sub-sample data of a first capacity; The recognition model is trained by using multiple subsample data as model input and the target text as model label. The recognition model is used to extract the first feature corresponding to the first data of the first volume and to recognize the first text corresponding to the first feature. The first text is output, and the first data is obtained by segmenting the target data input in real time; the recognition model is also used to recognize the corresponding second text based on the second data with a second capacity, and the second text is used to update the output text. The second data is generated by concatenating multiple first features, and the second capacity is the sum of the capacities of the multiple first features.
8. A speech recognition method, characterized in that, include: Real-time collection of user voice data; The speech data is segmented to obtain the first speech data of the first frame number; The speech recognition model is used to extract the first speech feature of the first speech data and to identify the first subtitle corresponding to the first speech feature. Display the first subtitle; By concatenating and accumulating multiple first speech features, a second speech data with a second frame number is obtained, wherein the second frame number is the sum of the frame numbers of each of the multiple first speech features; Based on the second voice data, the corresponding second subtitle is identified using the speech recognition model; Update the displayed subtitles using the second subtitle.
9. An electronic device, characterized in that, This includes processing components, storage components, and display components; The storage component stores one or more computer instructions; the one or more computer instructions are invoked and executed by the processing component, the processing component being used to: The real-time input voice data is segmented to obtain the first voice data of the first frame. The speech recognition model is used to extract the first speech feature of the first speech data and to identify the first text corresponding to the first speech feature. The first text is displayed in the display component; By concatenating and accumulating multiple first speech features, a second speech data with a second frame number is obtained, wherein the second frame number is the sum of the frame numbers of each of the multiple first speech features; Based on the second speech data, the corresponding second text is identified using the speech recognition model. The displayed text in the display component is updated using the second text.
10. A computing device, characterized in that, This includes processing components and storage components; The storage component stores one or more computer instructions; the one or more computer instructions are invoked and executed by the processing component, the processing component being used to: Determine the target sample data and corresponding target text for the second capacity; The target sample data is segmented to obtain multiple sub-sample data of a first capacity, wherein the target sample data is stream-input data; The recognition model is trained by using multiple subsample data as model input and the target text as model label. The recognition model is used to extract the first feature corresponding to the first data of the first volume and to recognize the first text corresponding to the first feature. The first text is output, and the first data is obtained by segmenting the target data input in real time; the recognition model also uses the second data based on the second capacity to recognize the corresponding second text, the second text is used to update the output text, the second data is generated by concatenating multiple first features, and the second capacity is the sum of the capacities of the multiple first features.
Citation Information
Patent Citations
Speech recognition method and device thereof
CN111261166A
Training method and decoding method of streaming end-to-end speech recognition model
CN111415667A