Data processing method and device for in-vehicle environment and medium
By generating and processing datasets in an in-vehicle environment and utilizing speech synthesis and noise reduction technologies, the noise interference problem of speech recognition systems in in-vehicle environments has been solved, achieving higher recognition accuracy and reliability, and improving driving experience and safety.
Patent Information
- Application Number
- CN202211582308.2
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Filing Date
- 2022-12-09
- Publication Date
- 2026-02-27
- Estimated Expiration
- 2042-12-09
AI Technical Summary
In a vehicle environment, noise interference can reduce the accuracy and reliability of voice recognition systems, affecting the driving experience and driving safety.
By acquiring in-vehicle data and clean datasets, speech synthesis technology is used to generate a clean in-vehicle dataset in a noise-free environment and a second in-vehicle dataset in the in-vehicle environment. Combined with noise reduction processing, a dataset for a speech recognition model in the in-vehicle environment is constructed.
It improves the accuracy and reliability of voice recognition models in in-vehicle environments, enhancing the driving experience and road safety.
Smart Images

Figure CN116013300B_ABST
Abstract
Description
TECHNICAL FIELD
[0001] The present application mainly relates to the field of information processing technology, more particularly to a data processing method and device for a vehicle environment and a medium. BACKGROUND
[0002] Intelligent cabin aims to integrate various IT and artificial intelligence technologies to create a new integrated digital platform in the vehicle, providing intelligent experiences such as emotion recognition, age detection, left-behind object detection, seat belt detection, driving state detection, etc. for drivers, and promoting vehicle safety.
[0003] In practical applications, as a natural human-computer interaction method, the intelligent cabin can automatically recognize the voice instructions of the driver based on a voice recognition system, realize human-vehicle interaction, and solve the problems of touch interaction mode, such as easy distraction of the driver's attention, low operation efficiency, and safety hazards, etc., i.e. improve driving safety and cabin interaction convenience.
[0004] However, in the driving scenario, the acoustic environment of the closed cabin is complex, and there are various noises such as wind noise, engine noise, wheel noise, etc., which reduces the recognition accuracy and reliability of the voice recognition system, thereby affecting the intelligent cabin driving experience and driving safety. SUMMARY
[0005] To solve the above technical problems, the present application proposes the following technical solutions:
[0006] The present application proposes a data processing method for a vehicle environment, which comprises:
[0007] obtaining a first vehicle data set and a pure data set; wherein the first vehicle data set refers to data in a vehicle environment, and the pure data set refers to data in a noise-free environment and at least partially associated with the vehicle environment;
[0008] According to the first vehicle data set and the pure data set, a vehicle pure data set for a noise-free environment and a second vehicle data set for a vehicle environment are obtained by voice synthesis;
[0009] According to the pure data set and the vehicle noise data contained in the first vehicle data set, a third vehicle data set for a vehicle environment is obtained;
[0010] According to the first vehicle data set, the pure data set, the vehicle pure data set, the second vehicle data set, and the third vehicle data set, a data set for a voice recognition model in a vehicle environment is obtained.
[0011] Optionally, the first vehicle-mounted data set comprises first vehicle-mounted audio data and corresponding first vehicle-mounted text data, and the pure data set comprises pure audio data and corresponding text data; the vehicle-mounted pure data set for a noise-free environment and the second vehicle-mounted data set for a vehicle-mounted environment are obtained by voice synthesis according to the first vehicle-mounted data set and the pure data set, comprising:
[0012] A first voice synthesis model for a noise-free environment is trained according to the pure audio data and the corresponding text data included in the pure data set;
[0013] The vehicle-mounted pure data set for a noise-free environment is obtained according to the first vehicle-mounted text data and the first voice synthesis model;
[0014] The second vehicle-mounted data set for a vehicle-mounted environment is obtained according to the first vehicle-mounted data set, the first voice synthesis model, and the second vehicle-mounted text data associated with a vehicle-mounted environment included in the pure data set.
[0015] Optionally, the vehicle-mounted pure data set for a noise-free environment is obtained according to the first vehicle-mounted text data and the first voice synthesis model, comprising:
[0016] The first vehicle-mounted text data included in the first vehicle-mounted data set is obtained, and the first vehicle-mounted text data is subjected to voice synthesis processing according to the first voice synthesis model to obtain the vehicle-mounted pure data set for a noise-free environment.
[0017] Optionally, the second vehicle-mounted data set for a vehicle-mounted environment is obtained according to the first vehicle-mounted data set, the first voice synthesis model, and the second vehicle-mounted text data associated with a vehicle-mounted environment included in the pure data set, comprising:
[0018] The first voice synthesis model is subjected to parameter adjustment according to the first vehicle-mounted data set to obtain a second voice synthesis model for the vehicle-mounted environment;
[0019] The second vehicle-mounted text data associated with a vehicle-mounted environment is determined from the text data included in the pure data set; the second vehicle-mounted text data is subjected to voice synthesis processing according to the second voice synthesis model to obtain the second vehicle-mounted data set for a vehicle-mounted environment.
[0020] Optionally, the third vehicle-mounted data set is a noise-reduced vehicle-mounted data set, and the third vehicle-mounted data set for a vehicle-mounted environment is obtained according to the vehicle-mounted noise data included in the pure data set and the first vehicle-mounted data set, comprising:
[0021] Vehicle-mounted noise data is obtained from the first vehicle-mounted audio data included in the first vehicle-mounted data set;
[0022] mixing the in-vehicle noise data and the pure audio data contained in the pure data set to obtain second in-vehicle audio data;
[0023] respectively performing noise reduction processing on the first in-vehicle audio data and the second in-vehicle audio data to obtain corresponding first noise-reduced in-vehicle data set and second noise-reduced in-vehicle data set.
[0024] Optionally, the noise reduction processing on the first in-vehicle audio data and the second in-vehicle audio data to obtain corresponding first noise-reduced in-vehicle data set and second noise-reduced in-vehicle data set comprises:
[0025] training a deep noise reduction model for in-vehicle environment according to the pure data set and the second in-vehicle audio data;
[0026] respectively performing noise reduction processing on the first in-vehicle audio data and the second in-vehicle audio data according to the deep noise reduction model to obtain corresponding first noise-reduced in-vehicle data set and second noise-reduced in-vehicle data set.
[0027] Optionally, the obtaining of the data set for the speech recognition model in the in-vehicle environment according to the first in-vehicle data set, the pure data set, the in-vehicle pure data set, the second in-vehicle data set and the third in-vehicle data set comprises:
[0028] performing perturbation processing on at least the first in-vehicle data set, the pure data set, the in-vehicle pure data set, the second in-vehicle data set, the first noise-reduced in-vehicle data set and / or the second noise-reduced in-vehicle data set to obtain a data set for in-vehicle environment;
[0029] transmitting the data set to at least one speech recognition system, and training a speech recognition model by the speech recognition system according to the data set to obtain text data corresponding to the collected audio data in the in-vehicle environment by the speech recognition model.
[0030] The application further provides a data processing method for a specified environment, which comprises:
[0031] obtaining a first specified data set and a pure data set; wherein the first specified data set refers to data in a specified environment; and the pure data set refers to data in a noise-free environment and at least partially associated with the specified environment;
[0032] obtaining a specified pure data set for the noise-free environment and a second specified data set for the specified environment by speech synthesis according to the first specified data set and the pure data set;
[0033] obtaining a third specified data set for a specified environment according to the specified noise data contained in the first specified data set and the pure data set;
[0034] obtaining a data set for a speech recognition model in a specified environment according to the first specified data set, the pure data set, the specified pure data set, the second specified data set and the third specified data set.
[0035] The application also provides a data processing device for a vehicle environment, which comprises:
[0036] a data obtaining module, configured to obtain a first vehicle data set and a pure data set; wherein the first vehicle data set refers to data in a vehicle environment; and the pure data set refers to data in a noise-free environment and at least partially associated with the vehicle environment;
[0037] a second vehicle data set obtaining module, configured to obtain a vehicle pure data set for a noise-free environment and a second vehicle data set for a vehicle environment by a speech synthesis method according to the first vehicle data set and the pure data set;
[0038] a third vehicle data set obtaining module, configured to obtain a third vehicle data set for a vehicle environment according to the pure data set and vehicle noise data contained in the first vehicle audio data set;
[0039] a data set obtaining module, configured to obtain a data set for a speech recognition model in a vehicle environment by using the first vehicle data set, the pure data set, the vehicle pure data set, the second vehicle data set and the third vehicle data set.
[0040] The application also provides a computer readable storage medium, which stores a computer program loaded and executed by a processor to implement the data processing method.
[0041] Therefore, the application provides a data processing method, device and medium for a vehicle environment, after obtaining a first vehicle data set in a vehicle environment and a pure data set in a noise-free environment, a vehicle pure data set for a noise-free environment and a second vehicle data set for a vehicle environment can be obtained by a speech synthesis method according to the first vehicle data set and the pure data set; a third vehicle data set for a vehicle environment can be obtained according to the pure data set and vehicle noise data contained in the first vehicle data set; and then a data set for a speech recognition model in a vehicle environment can be obtained according to the first vehicle data set, the pure data set, the vehicle pure data set, the second vehicle data set and the third vehicle data set. BRIEF DESCRIPTION OF DRAWINGS
[0042] In order to more clearly illustrate the technical solutions in the embodiments of the present application or the prior art, the accompanying drawings needed to be used in the embodiments or prior art description will be briefly introduced as follows. Obviously, the accompanying drawings in the following description only only the embodiments of the present application, and for those skilled in the art, other drawings can also be obtained without creative labor on the basis of the provided drawings.
[0043] Figure 1 A flowchart of an optional example of a data processing method for a specified environment according to the present application;
[0044] Figure 2 A flowchart of an optional example of a data processing method for a vehicle-mounted environment according to the present application;
[0045] Figure 3 A flowchart of another optional example of a data processing method for a vehicle-mounted environment according to the present application;
[0046] Figure 4 A flowchart of another optional example of a data processing method for a vehicle-mounted environment according to the present application;
[0047] Figure 5 A flowchart of another optional example of a data processing method for a vehicle-mounted environment according to the present application;
[0048] Figure 6 A structural diagram of an optional example of a data processing device for a specified environment according to the present application;
[0049] Figure 7 A structural diagram of an optional example of a data processing device for a vehicle-mounted environment according to the present application;
[0050] Figure 8 A hardware structural diagram of an optional example of a computer device suitable for the data processing method according to the present application;
[0051] Figure 9 A hardware structural diagram of another optional example of a computer device suitable for the data processing method according to the present application. DETAILED DESCRIPTION
[0052] To solve the technical problems described in the background section, the sample data required for training the speech recognition model is pure data in a noise-free environment, which can accurately realize speech recognition in a noise-free environment. However, in a vehicle environment with special acoustic environment and content characteristics (such as place names, personal names, and other special vocabularies), using the speech recognition model to recognize the audio collected in the intelligent cockpit scene cannot guarantee the accuracy of the recognition result, affecting the driving experience and road safety.
[0053] To improve the above problems, it is proposed to obtain mixed pure data of vehicle data in a vehicle environment such as the above intelligent cockpit scene, to form training data for training a speech recognition model. However, the amount of vehicle data that can be collected in a vehicle environment is limited, making it difficult for the speech recognition model to learn vehicle acoustic scene information and content characteristics specific to the vehicle scene from a small amount of vehicle data. Therefore, the speech recognition model trained in this way cannot guarantee the accuracy of the recognition result in the vehicle environment, and cannot meet the driving needs in the vehicle environment such as the intelligent cockpit.
[0054] To this end, the present application hopes to use a small amount of vehicle data and a large amount of pure data to simulate the generation of a large amount of data for the vehicle environment (such as audio data containing vehicle content), thereby realizing the training and learning of a speech recognition model suitable for a vehicle environment such as an intelligent cockpit. Using the speech recognition model trained in this way, the audio data collected in the vehicle environment is recognized to accurately obtain the text recognition result, and the voice control of the corresponding object in the vehicle environment is realized, such as controlling the air conditioner, playing songs, making a phone call, navigation, chatting, etc.
[0055] The technical solutions in the embodiments of the present application will be described clearly and completely below with reference to the drawings in the embodiments of the present application. Obviously, the described embodiments are only a part of the embodiments of the present application, not all. Based on the embodiments in the present application, all other embodiments obtained by those skilled in the art without creative labor are within the scope of protection of the present application.
[0056] Referring to Figure 1 An optional example of a data processing method for a specified environment is provided in the present application. The method can be applied to computer devices such as terminal devices and / or servers with certain data processing capabilities, and the execution subject of the method can be determined according to actual conditions. For example, as shown in Figure 1 The method can include:
[0057] Step S11, obtaining a first specified data set and a pure data set;
[0058] In the embodiments of the present application, the first specified data set refers to data in a specified environment, which can include audio data and text data in the specified environment; and the pure data set refers to data in a noise-free environment and at least partially associated with the specified environment, which can include pure audio data and corresponding text data.
[0059] The audio data included in the first specified data set can be collected in a simulated specified environment, and the content thereof is textually annotated to obtain corresponding text data. The data included in the pure data set can be obtained from various open source corpora, and the data quantity is often large, so that a large amount of training data for the model in the specified environment can be simulated or constructed therefrom. The present application does not limit the data content and the obtaining method of the first specified data set and the pure data set.
[0060] In step S12, the specified pure data set for the noise-free environment and the second specified data set for the specified environment are obtained by voice synthesis based on the first specified data set and the pure data set.
[0061] Voice synthesis, also known as text-to-speech (TTS), can convert any input text into corresponding speech. Therefore, the present application can use voice synthesis technology to perform voice synthesis processing on the text data included in the first specified data set according to the emotional requirements of the synthesized voice in the specified environment, to generate second specified audio data containing the text data content in the specified environment, which can form the second specified data set for the specified environment in combination with the text data.
[0062] After that, considering the content characteristics in the specified environment, a voice synthesis model for the specified environment can also be obtained based on the first specified data set in the specified environment to perform voice synthesis on the text data associated with the specified environment in the pure data set, to obtain the specified pure data set for the noise-free environment. The implementation process is not described in detail.
[0063] In step S13, the third specified data set for the specified environment is obtained based on the specified noise data included in the first specified data set and the pure data set.
[0064] The present application considers the acoustic environment in the specified environment, and can simulate a large amount of specified audio data in the specified environment based on the specified noise data included in the first specified data set in combination with a large amount of pure audio data included in the pure data set, to form the third specified data set for the specified environment. In order to improve the accuracy of the recognition result in the specified environment, the directly obtained audio data and the simulated audio data can be denoised to form the third specified data set for the specified environment, and the implementation process of step S13 is not described in detail.
[0065] Step S14, obtaining the data set for the speech recognition model in the specified environment according to the first specified data set, the pure data set, the specified pure data set, the second specified data set, and the third specified data set.
[0066] In some embodiments, the application can be composed of the obtained original first specified data set and pure data set, and the specified pure data set, the second specified data set, and the third specified data set in the specified environment generated or constructed by the above-mentioned different ways to realize the training requirement of the speech recognition model in the specified environment, and the recognition accuracy of the obtained speech recognition model in the specified environment.
[0067] Optionally, in order to further increase the data amount of the training data, the application can also perform data enhancement on the above-mentioned data sets by various perturbation processing methods to obtain the data set for the speech recognition model in the specified environment, that is, using a small amount of data in the specified environment to obtain a large amount of data in the specified environment, and ensuring the recognition accuracy of the speech recognition model in the specified environment trained accordingly.
[0068] The above-mentioned specified environment will be taken as the vehicle-mounted environment as an example for description, and the data processing method in other environments is similar, which will not be described one by one.
[0069] Referring to Figure 2 , an optional example flowchart of a data processing method for the vehicle-mounted environment is provided, which can be executed by a computer device, such as Figure 2 , the method can include:
[0070] Step S21, obtaining a first vehicle-mounted data set and a pure data set;
[0071] In the embodiment of the application, the first vehicle-mounted data set can include an audio data set in the vehicle-mounted environment and a text data set corresponding to the content. The vehicle-mounted environment, i.e. the vehicle-mounted acoustic environment, can be the actual environment in the driving scene, which often has noise environments such as a windscreen wiper, wind, an engine, a wheel, background music, and an interference loudspeaker, and a reverberation background environment generated by speaking in a closed space. The user can speak various voice instructions in this environment, such as commands for controlling the air conditioner, such as "turn up the temperature (turn down the temperature)"; commands for making a phone call, such as "call Zhang XX (call Zhang XX)"; commands for controlling the media player, such as "play the song / movie of XX"; navigation commands, such as "navigate to a certain place (route to a certain place)"; and / or reading news, previous conversations of the user, and other content. The application does not limit the data content contained in the first vehicle-mounted data set, which can be determined as appropriate.
[0072] The pure data set refers to data in a noise-free environment and at least partially associated with the vehicle environment, and can include a pure audio data set in a noise-free environment and a corresponding text data set. The pure data contained in the pure data set can be obtained from one or more open source corpora, such as at least one of a conversational speech corpus, a reading speech corpus, a Chinese corpus set, and a free Chinese corpus. The application does not limit the method of obtaining the first vehicle data set and the pure data set.
[0073] It should be noted that according to the above description of the first vehicle data set and the pure data set, the amount of data that the computer device can obtain from the first vehicle data set is much smaller than the amount of data in the pure data set, but the numerical values of the data amounts of the two data sets are not limited.
[0074] Step S22, according to the first vehicle data set and the pure data set, a vehicle pure data set for a noise-free environment and a second vehicle data set for a vehicle environment are obtained by speech synthesis.
[0075] Since the amount of data in the first vehicle data set is usually small, in order to obtain the recognition accuracy and reliability of the speech recognition model suitable for the vehicle environment, the application proposes to combine the pure data set containing a large amount of data, and simulate the generation of the second vehicle data set for the vehicle environment and the vehicle pure data set containing a large amount of vehicle data for the noise-free environment by speech synthesis, to expand the data containing vehicle content in the vehicle environment.
[0076] The application can use speech synthesis technology to perform speech synthesis processing on the vehicle text data associated with the vehicle environment contained in the pure data set to obtain vehicle audio data with the vehicle text data content. Similarly, the text data contained in the first vehicle data set can also be processed by speech synthesis to obtain pure data containing the vehicle text content. The application does not describe in detail how to simulate the generation of audio data containing vehicle content by speech synthesis processing.
[0077] Step S23, according to the pure data set and the vehicle noise data contained in the first vehicle data set, a third vehicle data set for the vehicle environment is obtained.
[0078] In the embodiment of the application, in addition to the above-described implementation of expanding the vehicle data containing vehicle content, the application can also use the pure audio contained in the pure data set and the vehicle noise data contained in the first vehicle data set to simulate a large amount of vehicle audio data in the vehicle environment. For this purpose, the application can perform noise reduction processing on a large amount of vehicle audio data in the vehicle environment and audio data contained in the first vehicle data set to simulate relatively clean vehicle data in the vehicle environment, thereby forming a third vehicle data set in the vehicle environment.
[0079] Step S24, according to the first vehicle-mounted data set, the pure data set, the vehicle-mounted pure data set, the second vehicle-mounted data set and the third vehicle-mounted data set, a data set for a speech recognition model in a vehicle-mounted environment is obtained.
[0080] Based on the above analysis, the computer device directly obtains the first vehicle-mounted data set containing a small amount of original vehicle-mounted data, and the pure data set containing a large amount of data in a noise-free environment and at least partially associated with the vehicle-mounted environment. Then, according to the two data sets, a large amount of vehicle-mounted pure data containing vehicle-mounted content in a vehicle-mounted acoustic environment is synthesized by a speech synthesis method to obtain the second vehicle-mounted data set. Then, the reverberation characteristics in the vehicle-mounted acoustic environment can be simulated according to the vehicle-mounted pure data and the vehicle-mounted noise data contained in the first vehicle-mounted data set, and a third vehicle-mounted data set is obtained by a noise reduction processing method. Thus, a data set containing a large amount of data is obtained for training a speech recognition model in a vehicle-mounted environment to ensure the accuracy of speech recognition in a vehicle-mounted environment, according to the original first vehicle-mounted data set and the pure data set, and the simulated vehicle-mounted pure data set, the second vehicle-mounted data set and the third vehicle-mounted data set.
[0081] Referring to Figure 3 A flowchart of another optional example of a data processing method for a vehicle-mounted environment is proposed in the present application, which can describe an optional refinement of the data processing method for a vehicle-mounted environment proposed above, as shown in Figure 3 The method can include:
[0082] Step S31, obtaining a first vehicle-mounted data set and a pure data set;
[0083] In the embodiments of the present application, it can be known from the above analysis that the first vehicle-mounted data set contains first vehicle-mounted audio data and corresponding first vehicle-mounted text data, and the pure data set contains pure audio data and corresponding text data. The implementation method of step S31 can refer to but is not limited to the description of the corresponding part of the above embodiments.
[0084] Step S32, training a first speech synthesis model for a noise-free environment according to the pure audio data and the corresponding text data contained in the pure data set;
[0085] In actual application, the pure audio data and the corresponding text data contained in the pure data set can be selected to train the multi-speaker speech synthesis model, so as to obtain a speech synthesis model suitable for a noise-free environment, and the speech synthesis model is used to generate natural expression speech with different predefined emotion categories from text data. According to the need, the multi-person emotion speech synthesis technology can be used to convert the text data into speech with a specified emotion of a specified speaker. The training process of the speech synthesis model is not described in detail.
[0086] In step S33, the first vehicle text data contained in the first vehicle data set and the first speech synthesis model are used to obtain a vehicle pure data set for a noise-free environment.
[0087] After the first speech synthesis model is obtained, the first vehicle text data contained in the first vehicle data set can be input into the first speech synthesis model for speech synthesis processing, so as to convert the first vehicle text data into audio data of a preset speaker, that is, vehicle audio data with the first vehicle text data as the speaking content. Since the training process of the first speech synthesis model does not consider vehicle noise, the vehicle audio data synthesized according to the first speech synthesis model is pure data containing the first vehicle text data. The speech synthesis process is not described in detail.
[0088] According to the above method, each piece of first vehicle text data contained in the first vehicle data set is input into the first speech synthesis model, and corresponding vehicle pure audio data in a noise-free vehicle environment is output, that is, pure audio simulating the content of the first vehicle text data. The obtained pure audio data constitutes a vehicle pure data set.
[0089] In step S34, the first vehicle audio data and the corresponding first vehicle text data contained in the first vehicle data set, the first speech synthesis model, and the second vehicle text data associated with the vehicle environment contained in the pure data set are used to obtain a second vehicle data set for the vehicle environment.
[0090] In the embodiment of the present application, a small amount of first vehicle data set can be used to fine-tune the first speech synthesis model, and a second speech synthesis model suitable for the vehicle environment can be obtained. Then, the second vehicle text data associated with the vehicle environment contained in the pure data set is subjected to speech synthesis processing according to the second speech synthesis model, and second vehicle audio data containing the content of the second vehicle text data is obtained. Subsequently, a large amount of second vehicle audio data or the second vehicle audio data and the corresponding second vehicle text data can be used to constitute a second vehicle data set for the vehicle environment. The process of obtaining the second vehicle audio data is similar to the process of obtaining the vehicle audio data, and is not described in detail.
[0091] In combination with the foregoing description of the vehicle environment, the second vehicle text data associated with the vehicle environment included in the pure data set can include control instructions in various aspects such as controlling the air conditioner, playing songs, making phone calls, navigation, chatting, and a large number of feature words involved in these contents, such as contact persons, singer names, navigation destinations, and other words outside the vocabulary, and the content characteristics in the vehicle environment are not limited in the present application, but can be determined as appropriate.
[0092] In step S35, vehicle noise data is obtained from the first vehicle audio data included in the first vehicle data set.
[0093] In step S36, the vehicle noise data and the pure audio data included in the pure data set are mixed to obtain second vehicle audio data.
[0094] As described above in relation to the process of obtaining the first vehicle data set, each first vehicle audio data included therein is audio data in the vehicle environment, which can include noise such as a wiper, wind, engine, wheels, background music, and interfering speakers inside and outside the vehicle, as well as reverberation noise generated in a closed vehicle, in addition to the audio of the user speaking.
[0095] Therefore, in order to simulate a large amount of vehicle data in the vehicle environment, the present application can extract the vehicle noise data included in each piece of first vehicle audio data included in the first vehicle data set, and then determine the pure audio data of the first speech synthesis model used for training as user speaking audio data, and mix it with the vehicle noise data to obtain second vehicle audio data for the vehicle environment. The present application does not limit the implementation method of mixing different audios, such as direct superimposition mixing.
[0096] In step S37, the first vehicle audio data and the second vehicle audio data are respectively denoised to obtain corresponding first denoised vehicle data set and second denoised vehicle data set.
[0097] In the application of speech recognition in the vehicle environment, the audio collected in the vehicle environment is usually denoised before speech recognition to improve the accuracy and efficiency of speech recognition. Based on this, after a large amount of second vehicle audio data in the vehicle environment is simulated according to the method described above, the second vehicle audio data constructed / simulated and the first vehicle audio data directly obtained can be denoised to obtain corresponding second denoised vehicle audio data and first denoised vehicle audio data.
[0098] Afterwards, the obtained large amount of second denoised in-vehicle audio data can be used to form a second denoised in-vehicle dataset in the in-vehicle environment, or the large amount of second denoised in-vehicle audio data and the text data corresponding to the pure audio data in the pure dataset can be used to form a second denoised in-vehicle dataset in the in-vehicle environment; similarly, the obtained first denoised in-vehicle audio data can be used to form a first denoised in-vehicle dataset in the in-vehicle environment; or the first denoised in-vehicle audio data and the corresponding first in-vehicle text data can be used to form a first denoised in-vehicle dataset in the in-vehicle environment.
[0099] In step S38, at least the first in-vehicle dataset, the pure dataset, the in-vehicle pure dataset, the second in-vehicle dataset, the first denoised in-vehicle dataset, and / or the second denoised in-vehicle dataset is subjected to a perturbation process to obtain a dataset for the in-vehicle environment.
[0100] In order to further expand the amount of data required for the dataset for the in-vehicle environment, each simulated in-vehicle dataset of the in-vehicle data, i.e., the in-vehicle pure dataset, the second in-vehicle dataset, the first denoised in-vehicle dataset, and the second denoised in-vehicle dataset, can be subjected to a random speed perturbation, a reverberation perturbation, and other perturbation processing methods to achieve data augmentation of at least one of these simulated in-vehicle datasets and the originally obtained first in-vehicle dataset and pure dataset, thereby further expanding the amount of data for the dataset for the in-vehicle environment. The implementation process of the data perturbation processing method will not be described in detail.
[0101] Optionally, the data contained in each simulated in-vehicle dataset obtained above and the originally obtained first in-vehicle dataset and pure dataset can be subjected to a perturbation process to obtain a large amount of perturbed data, and these data or these data and the data contained in each dataset before perturbation can be used to form a dataset for the in-vehicle environment.
[0102] In step S39, the dataset is transmitted to at least one speech recognition system, and the speech recognition system trains a speech recognition model according to the dataset to obtain text data corresponding to the collected audio data in the in-vehicle environment through the speech recognition model.
[0103] After the computer device obtains the dataset containing a large amount of data for the in-vehicle environment according to the above method, it sends the dataset as a training dataset of a speech recognition model to each speech recognition system of the computer device for training and learning to obtain a corresponding speech recognition model; or sends the dataset to a speech recognition system in another computer device for speech recognition model training. The implementation process of the speech recognition model training will not be described in detail.
[0104] In order to improve the accuracy and reliability of speech recognition, the application proposes to fuse the recognition results of multiple speech recognition models to obtain the text data of the to-be-recognized audio, so that the obtained data set containing a large amount of vehicle-mounted data in the vehicle-mounted environment can be transmitted to multiple speech recognition systems for model training to obtain multiple speech recognition models, so as to accurately obtain the text data corresponding to the collected audio data in the vehicle-mounted environment through the multiple speech recognition models, and improve the driving experience and driving safety.
[0105] Referring to Figure 4 For another optional example of the data processing method for the vehicle-mounted environment proposed in the application, the embodiment can further describe the acquisition process of the vehicle-mounted pure data set, the second vehicle-mounted data set and the third vehicle-mounted data set in the data processing method for the vehicle-mounted environment proposed above. For other execution steps of the method, refer to the description of the corresponding part of the above and below embodiments, and the embodiment will not be described in detail. As shown in Figure 4 The method can include:
[0106] Step S41, training a first speech synthesis model for a noise-free environment according to the pure audio data and the corresponding text data contained in the pure data set;
[0107] Step S42, obtaining first vehicle-mounted text data contained in the first vehicle-mounted data set, and performing speech synthesis processing on the first vehicle-mounted text data according to the first speech synthesis model to obtain a vehicle-mounted pure data set for a noise-free environment;
[0108] In combination with the above embodiment for the related description of the first vehicle-mounted data set and the pure data set, referring to the flowchart shown in Figure 5 The embodiment can obtain part of the pure audio data and the corresponding text data from the pure data set that meet the training requirements of the speech synthesis model, and then train and learn the initial speech synthesis model according to the part of the pure audio data and the corresponding text data to obtain the first speech synthesis model for the noise-free environment.
[0109] The initial speech synthesis model can be a neural network model based on deep learning, which usually includes a preprocessing network for extracting pronunciation and prosody and other linguistic features from text data, an acoustic model for generating acoustic features according to linguistic features, and a vocoder for synthesizing audio data according to acoustic features. The application does not describe the network structure of the three parts, which can be determined according to the processing requirements.
[0110] Based on the above analysis, this application can use the aforementioned first speech synthesis model to perform speech synthesis processing on each directly obtained first in-vehicle text data to obtain in-vehicle clean audio data containing the corresponding first in-vehicle text data content, thereby constructing an in-vehicle clean dataset for noise-free environments. Regarding this speech synthesis processing, the composition structure and function of the first speech synthesis model can be determined with reference to the above description.
[0111] Step S43: Based on the first in-vehicle audio data and the first in-vehicle text data contained in the first in-vehicle dataset, adjust the parameters of the first speech synthesis model to obtain a second speech synthesis model for the in-vehicle environment.
[0112] Step S44: Determine the second in-vehicle text data associated with the in-vehicle environment from the text data contained in the clean dataset;
[0113] Step S45: Based on the second speech synthesis model, perform speech synthesis processing on the second vehicle text data to obtain a second vehicle dataset for the vehicle environment.
[0114] In order to synthesize a large amount of data in the vehicle acoustic environment, such as Figure 5 As shown, in this embodiment, the first vehicle audio data and the first vehicle text data contained in the first vehicle dataset can be used to train and learn the first speech synthesis model, adjust its model parameters, and obtain a second speech synthesis model for the vehicle environment. The implementation process is similar to the training process of the first speech synthesis model, and will not be described in detail in this application.
[0115] In addition, this application can select text data related to the vehicle environment from the large amount of text data contained in the clean dataset and record it as the second vehicle text data. The second speech synthesis model is used to perform speech synthesis processing on the second vehicle text data to obtain the second vehicle audio data containing the content of the second vehicle text data, and thus constitute the second vehicle dataset for the vehicle environment.
[0116] Step S46: Obtain vehicle noise data from the first vehicle audio data contained in the first vehicle dataset;
[0117] Step S47: Mix the vehicle noise data and the clean audio data contained in the clean dataset to obtain the second vehicle audio data.
[0118] Step S48: Based on the clean audio data and the second in-vehicle audio data, a deep noise reduction model for the in-vehicle environment is trained.
[0119] In the embodiments of the present application, for the acoustic environment in the vehicle environment, vehicle noise data can be extracted from the first vehicle audio data in the vehicle environment for noise adding and reverberation operation, such as mixing the vehicle noise data with the pure audio data to simulate a large number of second vehicle audio data in the vehicle acoustic environment, and then determining the second vehicle audio data containing the same text data content and the pure audio data as a pair of positive and negative sample data, training and learning the neural network denoising algorithm based on deep learning to obtain a deep denoising model for the vehicle environment, i.e., a DNS (Deep Noise Suppression) model. The training and implementation process of the model is not described in detail.
[0120] In step S49, the first vehicle audio data and the second vehicle audio data are respectively processed by the deep denoising model to obtain the corresponding first denoising vehicle data set and the second denoising vehicle data set.
[0121] After obtaining the deep denoising model in the vehicle environment, each first vehicle audio data contained in the directly obtained first vehicle data set can be sequentially input into the deep denoising model to obtain relatively pure first denoising vehicle audio data, thereby forming the first denoising vehicle data set. Similarly, a large number of second vehicle audio data generated by simulation can also be sequentially input into the deep denoising model to obtain second denoising vehicle audio data, thereby forming the second denoising vehicle data set. The denoising process of the audio data can be determined in combination with the structure and operation principle of the deep denoising model, which is not described in detail herein.
[0122] In some embodiments, the language model LM can also be trained and obtained according to the first vehicle data set, so that the language model LM is used to select data with similar content from the pure data set as new training data, which is added to the data set for the vehicle environment, or used to implement the screening of vehicle text data, etc. The data processing requirements can be determined.
[0123] Based on the above analysis, in addition to the original first vehicle data set and the pure data set obtained, the corresponding audio data and / or text data can be processed by the speech synthesis TTS model and the deep denoising DNS model obtained by training to construct four new data sets, i.e., the vehicle pure data set, the second vehicle data set, the first denoising vehicle data set, and the second denoising vehicle data set. Then, the data amount of each data set can be further expanded through the perturbation processing method for data enhancement to obtain a data set containing a large amount of data for the speech recognition model in the vehicle environment. Then, the speech recognition model in the vehicle environment can be trained and obtained based on the data set, and the speech recognition model in the vehicle environment can be used for speech recognition in the vehicle environment. Figure 5As shown, the dataset can be transmitted to multiple speech recognition systems for training and learning, resulting in multiple speech recognition models. These models can then be integrated to achieve speech recognition in an in-vehicle environment, thereby improving the accuracy and reliability of the recognition results.
[0124] In some embodiments, after training multiple speech recognition models according to the method described above, such as Figure 5 As shown, an in-vehicle test dataset can be obtained from the in-vehicle environment, enabling the testing of the recognition accuracy and reliability of multiple speech recognition models. The acquisition process for this in-vehicle test dataset can involve collecting audio data continuously for 20 hours (but not limited to this, depending on the actual situation) in an in-vehicle environment. The audio data collected continuously for 10 hours can be used as the first in-vehicle audio data, and the audio data collected for the other 10 hours as the in-vehicle test audio data. This, combined with corresponding in-vehicle test text data, constitutes the in-vehicle test dataset, but it is not limited to this method of acquiring in-vehicle audio data.
[0125] Based on the above description of in-vehicle audio data, the collected audio data may include, but is not limited to, audio of different types of user commands such as air conditioning, telephone calls, music, points of interest, or others. This application does not limit the content of the aforementioned audio data and may determine it as appropriate.
[0126] Subsequently, the in-vehicle test audio data contained in the in-vehicle test dataset is transmitted to each speech recognition system. The speech recognition models of these systems perform text recognition to obtain the corresponding test text data. The obtained test text data are then fused to obtain the test recognition results of the in-vehicle test audio data, which is used to determine whether to adjust the parameters of the speech recognition model. It is evident that this application employs a multi-model content fusion method, which, compared to the recognition performance of a single speech recognition model, ensures high efficiency and high accuracy recognition results under different in-vehicle environments.
[0127] Optionally, the testing process for the speech recognition model described above can be performed by an in-vehicle terminal with multiple trained speech recognition models, or the in-vehicle terminal can connect to the network and send the collected in-vehicle test audio data to the server, and the server can test the multiple trained speech recognition models according to the above method. The implementation process is not detailed in this application.
[0128] Reference Figure 6 This is a schematic diagram of an optional example of a data processing apparatus for a specified environment proposed in this application, as shown below. Figure 6 As shown, the device may include:
[0129] The data obtaining module 61 is configured to obtain a first designated data set and a pure data set; wherein the first designated data set refers to data in a designated environment; and the pure data set refers to data in a noise-free environment and at least partially associated with the designated environment.
[0130] The second designated data set obtaining module 62 is configured to obtain a designated pure data set for a noise-free environment and a second designated data set for the designated environment by means of voice synthesis based on the first designated data set and the pure data set.
[0131] The third designated data set obtaining module 63 is configured to obtain a third designated data set for the designated environment based on the pure data set and designated noise data contained in the first designated data set.
[0132] The data set obtaining module 64 is configured to obtain a data set for a voice recognition model in the designated environment based on the first designated data set, the pure data set, the designated pure data set, the second designated data set and the third designated data set.
[0133] Referring to Figure 7 , a structural schematic diagram of an optional example of a data processing device for a vehicle-mounted environment is shown in FIG. 1. Figure 7 As shown in the figure, the device can include:
[0134] The data obtaining module 71 is configured to obtain a first vehicle-mounted data set and a pure data set; wherein the first vehicle-mounted data set refers to data in a vehicle-mounted environment; and the pure data set refers to data in a noise-free environment and at least partially associated with the vehicle-mounted environment.
[0135] The second vehicle-mounted data set obtaining module 72 is configured to obtain a vehicle-mounted pure data set for a noise-free environment and a second vehicle-mounted data set for the vehicle-mounted environment by means of voice synthesis based on the first vehicle-mounted data set and the pure data set.
[0136] The third vehicle-mounted data set obtaining module 73 is configured to obtain a third vehicle-mounted data set for the vehicle-mounted environment based on the pure data set and vehicle-mounted noise data contained in the first vehicle-mounted audio data set.
[0137] The data set obtaining module 74 is configured to obtain a data set for a voice recognition model in the vehicle-mounted environment based on the first vehicle-mounted data set, the pure data set, the vehicle-mounted pure data set, the second vehicle-mounted data set and the third vehicle-mounted data set.
[0138] In some embodiments, the first in-vehicle data set contains first in-vehicle audio data and corresponding first in-vehicle text data, and the pure data set contains pure audio data and corresponding text data, based on which the second in-vehicle data set obtaining module 72 can include:
[0139] a first speech synthesis model training unit configured to train a first speech synthesis model for a noise-free environment according to the pure audio data and the corresponding text data contained in the pure data set;
[0140] an in-vehicle pure data set obtaining unit configured to obtain an in-vehicle pure data set for a noise-free environment according to the first in-vehicle text data and the first speech synthesis model;
[0141] a second in-vehicle data set obtaining unit configured to obtain a second in-vehicle data set for a vehicle environment according to the first in-vehicle data set, the first speech synthesis model, and second in-vehicle text data associated with the vehicle environment contained in the pure data set.
[0142] Optionally, the in-vehicle pure data set obtaining unit can include:
[0143] a first in-vehicle text data obtaining unit configured to obtain the first in-vehicle text data contained in the first in-vehicle data set;
[0144] a first speech synthesis processing unit configured to perform speech synthesis processing on the first in-vehicle text data according to the first speech synthesis model to obtain the in-vehicle pure data set for a noise-free environment.
[0145] Optionally, the second in-vehicle data set obtaining unit can include:
[0146] a second speech synthesis model obtaining unit configured to adjust parameters of the first speech synthesis model according to the first in-vehicle data set to obtain a second speech synthesis model for the vehicle environment;
[0147] a second in-vehicle text data determining unit configured to determine the second in-vehicle text data associated with the vehicle environment from the text data contained in the pure data set;
[0148] a second speech synthesis processing unit configured to perform speech synthesis processing on the second in-vehicle text data according to the second speech synthesis model to obtain the second in-vehicle data set for the vehicle environment.
[0149] In yet some embodiments, the third in-vehicle data set obtaining module 73 can include:
[0150] an in-vehicle noise data obtaining unit configured to obtain in-vehicle noise data from the first in-vehicle audio data contained in the first in-vehicle data set;
[0151] a second in-vehicle audio data obtaining unit, configured to perform mixing processing on the in-vehicle noise data and pure audio data contained in the pure data set to obtain second in-vehicle audio data;
[0152] a noise reduction processing unit, configured to perform noise reduction processing on the first in-vehicle audio data and the second in-vehicle audio data respectively to obtain corresponding first noise-reduced in-vehicle data set and second noise-reduced in-vehicle data set.
[0153] Optionally, the noise reduction processing unit can include:
[0154] a deep noise reduction model training unit, configured to train a deep noise reduction model for in-vehicle environment according to the pure data set and the second in-vehicle audio data;
[0155] a noise-reduced in-vehicle data set obtaining unit, configured to perform noise reduction processing on the first in-vehicle audio data and the second in-vehicle audio data respectively according to the deep noise reduction model to obtain corresponding first noise-reduced in-vehicle data set and second noise-reduced in-vehicle data set.
[0156] Optionally, the data set obtaining module 74 can include:
[0157] a perturbation processing unit, configured to perform perturbation processing on at least the first in-vehicle data set, the pure data set, the in-vehicle pure data set, the second in-vehicle data set, the first noise-reduced in-vehicle data set and / or the second noise-reduced in-vehicle data set to obtain a data set for in-vehicle environment.
[0158] a data set transmission unit, configured to transmit the data set to at least one speech recognition system, and train a speech recognition model by the speech recognition system according to the data set to obtain text data corresponding to the collected audio data in in-vehicle environment by the speech recognition model.
[0159] It should be noted that the various modules, units, etc. in the above-mentioned device embodiments can be stored in the memory as program modules, and the processor executes the above-mentioned program modules stored in the memory to realize the corresponding functions. For the functions realized by each program module and its combination, and the technical effects achieved, reference can be made to the descriptions of the corresponding parts of the above-mentioned method embodiments, and the present embodiment will not be described again.
[0160] The present application also provides a computer readable storage medium, which can store computer readable instructions, the computer readable instructions can be called and loaded by the processor to realize each step of the data processing method described in the above-mentioned corresponding method embodiments.
[0161] Reference Figure 8For an optional example of the hardware structure of the computer device applicable to the data processing method proposed in this application, the computer device can be a server, such as a standalone physical server, a server cluster composed of multiple physical servers, or a cloud server capable of implementing cloud computing. In some embodiments, the computer device can also be a terminal device with certain data processing capability, such as a vehicle-mounted terminal, a smartphone, etc. This application takes the computer device as a server for example. As shown in Figure 8 The computer device can include at least one memory 81 and at least one processor 82, wherein:
[0162] The memory 81 can be used to store the program of the data processing method described in the above method embodiments (such as the data processing method in a specified environment or the data processing method in a vehicle-mounted environment); the processor 82 can load and execute the program stored in the memory to implement the steps of the data processing method described in the above corresponding method embodiments. The specific implementation process can refer to the description of the corresponding part of the above embodiments, and will not be repeated here.
[0163] In the embodiments of this application, the memory 81 can include a high-speed random access memory, and can also include a non-volatile memory, such as at least one magnetic disk storage device or other volatile solid-state storage device. The processor 82 can be a central processing unit (CPU), an application-specific integrated circuit (ASIC), a digital signal processor (DSP), a dedicated integrated circuit (ASIC), a ready-to-program gate array (FPGA), or other programmable logic devices, etc. This application does not limit the structure and model of the above-mentioned memory 81 and processor 82, which can be flexibly adjusted according to actual needs.
[0164] It should be understood that, Figure 8 The structure of the computer device shown in the above-mentioned figures does not constitute a limitation on the computer device in the embodiments of this application. In actual applications, the computer device can include more components than Figure 8 shown, or combine some components, such as Figure 9 shown, in the case of a terminal device, it can also include at least one input component such as a touch sensing unit for sensing touch events on a touch display panel, a keyboard, a mouse, a camera, a sound pickup, etc.; at least one output component such as a display, a speaker, a vibration mechanism, a lamp, etc.; an antenna; a sensor module; a power supply module, etc. Figure 9 The listed input components and output components are not shown, and the hardware structure can be determined according to the type of the terminal device and its functional requirements, which will not be listed one by one in this application.
[0165] Finally, it is to be understood that the singular forms "a", "an", and "the" include plural referents unless the context clearly dictates otherwise. Typically, terms "comprising" and "containing" are synonymous with "including" and "comprising" and are inclusive or open-ended and do not exclude additional, unrecited elements or method steps. Elements defined by followings "a" or "an" do not exclude the existence of additional identical elements in the process, method, article, or apparatus including the element.
[0166] In the description of the present application, unless otherwise stated, " / " means or, for example, A / B can mean A or B; "and / or" in this document merely describes associated objects in association, which means that there can be three relationships, for example, A and / or B, which means that there are three cases of A alone, A and B together, and B alone. In addition, in the description of the embodiments of the present application, "multiple" means two or more than two.
[0167] The terms such as "first", "second" and the like used in the present application are only for the purpose of description, to distinguish one operation, unit or module from another operation, unit or module, and do not necessarily require or imply any such actual relationship or order between the units, operations or modules. And it cannot be understood as indicating or implying relative importance or implicitly indicating the number of indicated technical features, so the features with "first", "second" can explicitly or implicitly include one or more features.
[0168] In addition, each embodiment in the present specification is described in a progressive or parallel manner, and each embodiment focuses on the difference from other embodiments, and the same or similar parts between each embodiment can be referred to each other. For the device and computer device disclosed by the embodiments, since it corresponds to the method disclosed by the embodiments, the description is relatively simple, and the related parts can be referred to the method part.
[0169] The above description of the disclosed embodiments enables a person skilled in the art to implement or use the present application. Various modifications to the embodiments will be apparent to those skilled in the art, and the general principles defined herein can be implemented in other embodiments without departing from the spirit or scope of the present application. Therefore, the present application will not be limited to the embodiments shown herein, but will conform to the widest scope consistent with the principles and novel features disclosed herein.
Claims
1. A method for data processing in a vehicle environment, the method comprising: obtaining a first vehicle dataset and a clean dataset, wherein the first vehicle dataset refers to data in a vehicle environment, and the clean dataset refers to data in a noise-free environment and at least partially associated with the vehicle environment; obtaining a clean vehicle dataset for a noise-free environment and a second vehicle dataset for the vehicle environment by voice synthesis based on the first vehicle dataset and the clean dataset, wherein the second vehicle dataset is obtained by voice synthesis processing of text data associated with the vehicle environment included in the clean dataset; obtaining a third vehicle dataset for the vehicle environment based on vehicle noise data included in the first vehicle dataset and the clean dataset; obtaining a dataset for a voice recognition model in the vehicle environment based on the first vehicle dataset, the clean dataset, the clean vehicle dataset, the second vehicle dataset, and the third vehicle dataset. 2.The method of claim 1, wherein the first vehicle dataset includes first vehicle audio data and corresponding first vehicle text data, and the clean dataset includes clean audio data and corresponding text data; and the obtaining the clean vehicle dataset for a noise-free environment and the second vehicle dataset for the vehicle environment by voice synthesis based on the first vehicle dataset and the clean dataset comprises: training a first voice synthesis model for a noise-free environment based on clean audio data and corresponding text data included in the clean dataset; obtaining the clean vehicle dataset for a noise-free environment based on the first vehicle text data and the first voice synthesis model; obtaining the second vehicle dataset for the vehicle environment based on the first vehicle dataset, the first voice synthesis model, and second vehicle text data associated with the vehicle environment included in the clean dataset. 3.The method of claim 2, wherein the obtaining the second vehicle dataset for the vehicle environment based on the first vehicle dataset, the first voice synthesis model, and the second vehicle text data associated with the vehicle environment included in the clean dataset comprises: adjusting the first voice synthesis model based on the first vehicle dataset to obtain a second voice synthesis model for the vehicle environment; determining the second vehicle text data associated with the vehicle environment from text data included in the clean dataset; and obtaining the second vehicle dataset for the vehicle environment by voice synthesis processing of the second vehicle text data based on the second voice synthesis model. 4.The method of any one of claims 1-3, wherein the third vehicle dataset is a denoised vehicle dataset, and the obtaining the third vehicle dataset for the vehicle environment based on the clean dataset and the vehicle noise data included in the first vehicle dataset comprises: obtaining vehicle noise data from the first vehicle audio data included in the first vehicle dataset. mixing the in-vehicle noise data and the pure audio data contained in the pure data set to obtain second in-vehicle audio data; respectively performing noise reduction processing on the first in-vehicle audio data and the second in-vehicle audio data to obtain corresponding first noise-reduced in-vehicle data set and second noise-reduced in-vehicle data set.
5. The method of claim 4, wherein the respective noise reduction processing on the first in-vehicle audio data and the second in-vehicle audio data to obtain corresponding first noise-reduced in-vehicle data set and second noise-reduced in-vehicle data set comprises: training a deep noise reduction model for in-vehicle environment according to the pure data set and the second in-vehicle audio data; respectively performing noise reduction processing on the first in-vehicle audio data and the second in-vehicle audio data according to the deep noise reduction model to obtain corresponding first noise-reduced in-vehicle data set and second noise-reduced in-vehicle data set.
6. The method of claim 4, wherein the obtaining a data set for a speech recognition model in an in-vehicle environment according to the first in-vehicle data set, the pure data set, the in-vehicle pure data set, the second in-vehicle data set, and the third in-vehicle data set comprises: performing perturbation processing on at least the first in-vehicle data set, the pure data set, the in-vehicle pure data set, the second in-vehicle data set, the first noise-reduced in-vehicle data set, and / or the second noise-reduced in-vehicle data set to obtain a data set for an in-vehicle environment; transmitting the data set to at least one speech recognition system, and training a speech recognition model by the speech recognition system according to the data set to obtain text data corresponding to the collected audio data in an in-vehicle environment by the speech recognition model.
7. A data processing method for a specified environment, the method comprising: obtaining a first specified data set and a pure data set; wherein the first specified data set refers to data in a specified environment; and the pure data set refers to data in a noise-free environment and at least partially associated with the specified environment; obtaining a specified pure data set for a noise-free environment and a second specified data set for the specified environment by speech synthesis according to the first specified data set and the pure data set; wherein the second specified data set is obtained by speech synthesis processing on text data associated with the specified environment contained in the pure data set; obtaining a third specified data set for the specified environment according to specified noise data contained in the pure data set and the first specified data set; obtaining a data set for a speech recognition model in the specified environment according to the first specified data set, the pure data set, the specified pure data set, the second specified data set, and the third specified data set.
8. A data processing apparatus for an in-vehicle environment, the apparatus comprising: a data obtaining module configured to obtain a first in-vehicle data set and a pure data set; wherein the first in-vehicle data set refers to data in an in-vehicle environment; and the pure data set refers to data in a noise-free environment and at least partially associated with the in-vehicle environment. a second vehicle-mounted data set obtaining module, configured to obtain, by means of voice synthesis, a vehicle-mounted clean data set for a noise-free environment and a second vehicle-mounted data set for a vehicle-mounted environment according to the first vehicle-mounted data set and the clean data set, wherein the second vehicle-mounted data set is obtained by performing voice synthesis processing on text data associated with the vehicle-mounted environment and contained in the clean data set; a third vehicle-mounted data set obtaining module, configured to obtain a third vehicle-mounted data set for the vehicle-mounted environment according to the clean data set and vehicle-mounted noise data contained in the first vehicle-mounted audio data set; a data set obtaining module, configured to obtain a data set for a voice recognition model in the vehicle-mounted environment by using the first vehicle-mounted data set, the clean data set, the vehicle-mounted clean data set, the second vehicle-mounted data set, and the third vehicle-mounted data set. 9.A computer readable storage medium, having a computer program stored thereon, wherein the computer program is loaded and executed by a processor to implement the data processing method according to any one of claims 1 to 7.
Citation Information
Patent Citations
Processing method and device for voice recognition in vehicle and electronic equipment
CN108022591A
Vehicle-mounted voice recognition method and system
CN110459234A