A data generation and vehicle-mounted speech recognition method, device and electronic equipment
By acquiring full audio recordings of English alphabet pronunciations from subjects of different genders, performing audio splicing and data enhancement, training data for English proper nouns of in-vehicle functional modules is generated. This solves the problem of low accuracy in recognizing English proper nouns in in-vehicle voice recognition systems and enables accurate recognition of newly added in-vehicle functional modules.
Patent Information
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Filing Date
- 2022-12-30
- Publication Date
- 2026-04-07
AI Technical Summary
Existing in-vehicle voice recognition systems have low accuracy in recognizing English proper nouns, especially when adding new in-vehicle function modules, which lacks relevant data and makes recognition difficult.
By acquiring full audio recordings of English alphabet pronunciations from subjects of different genders, audio splicing and data enhancement are performed to generate training data of English proper nouns for in-vehicle functional modules, which are then used to train the in-vehicle speech recognition model.
This improves the accuracy of the in-vehicle voice recognition system in recognizing English proper nouns, ensuring accurate recognition of newly added in-vehicle functional modules.
Smart Images

Figure CN116206598B_ABST
Abstract
Description
TECHNICAL FIELD
[0001] The present application relates to the technical field of speech recognition, and in particular to a data generation and vehicle-mounted speech recognition method and device and electronic equipment. BACKGROUND
[0002] There are many vehicle-mounted function modules in the vehicle-mounted field, some of which have corresponding English proper nouns. By recognizing the English proper nouns input by the user, the corresponding vehicle-mounted function module can be recognized and activated. For example, the user can input "open GPS" by voice, and the corresponding navigation module can be activated after voice recognition.
[0003] The existing training of vehicle-mounted speech recognition models is mostly based on the English proper nouns of existing vehicle-mounted function modules. There is not much English proper noun data for training of letter pronunciation. If a new corresponding proper noun is generated based on a vehicle-mounted function module, the speech recognition system lacks relevant data, which causes the speech recognition system to be unable to recognize the English proper nouns of the newly added vehicle-mounted function module or to have a low recognition accuracy. Therefore, it is urgent to propose a data generation method to improve the recognition accuracy of the speech recognition system for English proper nouns of letter pronunciation. SUMMARY
[0004] Therefore, the technical problem to be solved by the present application is to overcome the low recognition accuracy of the existing speech recognition system for English proper nouns, thereby providing a data generation and vehicle-mounted speech recognition method, device and electronic equipment.
[0005] According to a first aspect, the present application discloses a data generation method, which comprises: obtaining a plurality of recorded audio of pronunciation of full-quantity English letters corresponding to different gender subjects; performing audio splicing based on the recorded audio of pronunciation of full-quantity English letters corresponding to each gender subject to obtain English proper noun training data of a vehicle-mounted function module corresponding to different gender subjects. The recorded audio of pronunciation of full-quantity English letters of different gender subjects and the corresponding English proper noun training data of the vehicle-mounted function module are used for training of a vehicle-mounted speech recognition model.
[0006] Optionally, the audio splicing based on the pronunciation recording audio of the full amount of English letters corresponding to each gender subject comprises: audio splicing of the pronunciation recording audio of the full amount of English letters corresponding to each subject in the same gender subject to obtain first spliced audio data corresponding to the English proper noun of the vehicle-mounted function module; audio splicing of the pronunciation recording audio of the full amount of English letters corresponding to multiple subjects in the same gender subject to obtain second spliced audio data corresponding to the English proper noun of the vehicle-mounted function module; adding background noise to the first spliced audio data and the second spliced audio data; and generating the English proper noun training data of the vehicle-mounted function module corresponding to different gender subjects based on the first spliced audio data and the second spliced audio data after adding background noise.
[0007] Optionally, the method further comprises: performing data enhancement processing on the pronunciation recording audio of the full amount of English letters of the different gender subjects and the English proper noun training data corresponding to the vehicle-mounted function module.
[0008] Optionally, the number of pronunciation recording audio samples of the full amount of English letters corresponding to different gender subjects is the same.
[0009] According to a second aspect, the embodiments of the present application also disclose a vehicle-mounted voice recognition method, which comprises: responding to a receiving operation of voice data; when the voice data is received, recognizing the voice data by using a vehicle-mounted voice recognition model, the vehicle-mounted voice recognition model being trained by using the data generated by the data generation method of the first aspect or any optional implementation manner of the first aspect; and determining whether a corresponding vehicle-mounted function module needs to be enabled according to a recognition result.
[0010] According to a third aspect, the embodiments of the present application also disclose a data generation device, which comprises: a pronunciation recording audio acquisition module, configured to acquire pronunciation recording audio of multiple full amounts of English letters corresponding to different gender subjects; and an audio splicing module, configured to perform audio splicing based on the pronunciation recording audio of the full amount of English letters corresponding to each gender subject to obtain English proper noun training data of a vehicle-mounted function module corresponding to different gender subjects, the pronunciation recording audio of the full amount of English letters of the different gender subjects and the English proper noun training data corresponding to the vehicle-mounted function module being used for training of a vehicle-mounted voice recognition model.
[0011] Optionally, the audio splicing module comprises: a first audio splicing module configured to splice the audio of pronunciation of all the English letters corresponding to each subject in the same gender group to obtain first spliced audio data corresponding to the English proper noun of the vehicle-mounted function module; a second audio splicing module configured to splice the audio of pronunciation of all the English letters corresponding to multiple subjects in the same gender group to obtain second spliced audio data corresponding to the English proper noun of the vehicle-mounted function module; a noise adding module configured to add background noise to the first spliced audio data and the second spliced audio data; and a noun data generation module configured to generate the English proper noun training data of the vehicle-mounted function module corresponding to different gender groups based on the first spliced audio data and the second spliced audio data after adding the background noise.
[0012] According to a fourth aspect, the embodiments of the present application also disclose a vehicle-mounted voice recognition device, the device further comprising: a data receiving module configured to respond to a receiving operation of voice data; a voice data recognition module configured to recognize the voice data by using a vehicle-mounted voice recognition model when the voice data is received, the vehicle-mounted voice recognition model being trained by data generated by the data generation method according to the first aspect or any optional embodiment of the first aspect; and a function module enabling module configured to determine whether a corresponding vehicle-mounted function module needs to be enabled according to a recognition result.
[0013] According to a fifth aspect, the embodiments of the present application also disclose an electronic device, comprising: at least one processor; and a memory connected with the at least one processor in communication; wherein the memory stores instructions executable by the at least one processor, and the instructions are executed by the at least one processor to enable the at least one processor to perform the steps of the data generation method according to the first aspect or any optional embodiment of the first aspect, or the steps of the vehicle-mounted voice recognition method according to the second aspect.
[0014] According to a sixth aspect, the embodiments of the present application also disclose a computer-readable storage medium, which stores a computer program, and the computer program is executed by a processor to implement the steps of the data generation method according to the first aspect or any optional embodiment of the first aspect, or the steps of the vehicle-mounted voice recognition method according to the second aspect.
[0015] The technical scheme of the present application has the following advantages:
[0016] The data generation method provided by the application, by obtaining the pronunciation recording audio of all letters corresponding to different gender subjects, based on the pronunciation recording audio of all letters corresponding to each gender subject, the audio splicing is performed to obtain the English proper noun training data of the vehicle-mounted function module corresponding to different gender subjects, the vehicle-mounted speech recognition model is trained, the data types and quantity for training the vehicle-mounted speech recognition model are increased, and the recognition accuracy of the vehicle-mounted speech recognition system integrated with the vehicle-mounted speech recognition model for the pronunciation of English proper nouns of letters is improved. BRIEF DESCRIPTION OF DRAWINGS
[0017] In order to more clearly illustrate the technical solutions in the specific embodiments of the present application or the prior art, the drawings needed to be used in the specific embodiments or prior art description will be briefly introduced. Obviously, the drawings in the following description are some embodiments of the present application, and other drawings can be obtained by those skilled in the art without creative labor.
[0018] Figure 1 A flow chart of a specific example of the data generation method in the embodiments of the present application;
[0019] Figure 2 A flow chart of a specific example of the vehicle-mounted speech recognition method in the embodiments of the present application;
[0020] Figure 3 A principle block diagram of a specific example of the data generation device in the embodiments of the present application;
[0021] Figure 4 A principle block diagram of a specific example of the vehicle-mounted speech recognition device in the embodiments of the present application;
[0022] Figure 5 A specific example of the electronic device in the embodiments of the present application. DETAILED DESCRIPTION
[0023] The technical solutions of the present application will be described clearly and completely in combination with the drawings. Obviously, the described embodiments are some embodiments of the present application, not all embodiments. Based on the embodiments in the present application, all other embodiments obtained by those skilled in the art without creative labor are within the protection scope of the present application.
[0024] In the description of the present application, it should be noted that the terms "center", "upper", "lower", "left", "right", "vertical", "horizontal", "inner", "outer" and the like indicate the orientation or positional relationship based on the orientation or positional relationship shown in the drawings, and are only for the convenience of describing the present application and simplifying the description, and do not indicate or imply that the devices or elements referred to must have a particular orientation, be constructed and operated in a particular orientation, and therefore cannot be understood as a limitation on the present application. In addition, the terms "first", "second", "third" are only for descriptive purposes and cannot be understood as indicating or implying relative importance.
[0025] In the description of the present application, it should be noted that unless otherwise explicitly specified and limited, the terms "mounting", "connecting", "connecting" should be understood in a broad sense, for example, it can be fixedly connected, or it can be detachably connected, or integrally connected; it can be mechanically connected, or it can be electrically connected; it can be directly connected, or it can be indirectly connected through an intermediate medium; it can be the communication between two elements, it can be wireless connection, or it can be wired connection. For those skilled in the art, the specific meaning of the above terms in the present application can be understood according to the specific circumstances.
[0026] In addition, the technical features involved in the different embodiments of the application described below can be combined with each other as long as they do not conflict with each other.
[0027] The embodiments of the present application disclose a data generation method, as shown in the method, the method comprises the following steps: Figure 1 As shown in the method, the method comprises the following steps:
[0028] Step 101, obtaining a plurality of pronunciation recording audios of full English letters corresponding to different gender subjects, for example, in the embodiments of the present application, pronunciation recording audios of 26 English letters "a-z" corresponding to male, female adult gender subjects and child category subjects are obtained, and pronunciation recording audios of letter pronunciation units composed of multiple letters, such as abc, gps, etc. to obtain full English letter audios, which are only examples. After obtaining the recording audios, the audios can be pre-processed, such as uniform volume and vad mute suppression processing, which are only examples, to ensure that the obtained audios are in a uniform standard format for subsequent processing.
[0029] Step 102, based on the pronunciation recording audios of full English letters corresponding to each gender subject, audio splicing is performed to obtain English proper noun training data of vehicle-mounted function modules corresponding to different gender subjects, and the pronunciation recording audios of full English letters of the different gender subjects and the corresponding English proper noun training data of the vehicle-mounted function modules are used for training of the vehicle-mounted voice recognition model.
[0030] Exemplarily, in the embodiments of the present application, according to some existing English proper nouns corresponding to vehicle-mounted function modules, such as gps, etc., the audio splicing processing between multiple letters in the pronunciation recording audio of male category subjects or female category subjects or child category is performed, such as splicing the pronunciation recording audio of the letters g, p and s in the pronunciation recording audio of a certain subject of the male category to obtain the gps English proper noun training data, corresponding to the gps navigation module in the vehicle-mounted field, based on the above steps, the English proper noun training data of the vehicle-mounted function modules corresponding to different gender subjects can be obtained, and all the English proper noun training data of the vehicle-mounted function modules after splicing and the pronunciation recording audio of all the English letters of different gender subjects can be used to train the vehicle-mounted voice recognition model. If a new corresponding English proper noun is set according to a certain newly added vehicle-mounted function module in the vehicle-mounted system, since the pronunciation audio of all the English letters is used to train the vehicle-mounted voice recognition model, the new English proper noun received by the voice recognition system can be accurately recognized. The English proper noun training data obtained after splicing in the embodiments of the present application is used to train the vehicle-mounted voice recognition model, which is only an example, and can also be used in practice in other fields according to actual conditions.
[0031] The data generation method provided by the present application obtains the pronunciation recording audio of all the English letters corresponding to different gender subjects, performs audio splicing based on the pronunciation recording audio of all the English letters corresponding to each gender subject to obtain the English proper noun training data of the vehicle-mounted function modules corresponding to different gender subjects, and trains the vehicle-mounted voice recognition model, thereby increasing the types and quantity of data used to train the vehicle-mounted voice recognition model and improving the recognition accuracy of the vehicle-mounted voice recognition system integrated with the vehicle-mounted voice recognition model for the pronunciation of English proper nouns.
[0032] As an optional embodiment of the present application, the audio splicing based on the pronunciation recording audio of all the English letters corresponding to each gender subject comprises: performing audio splicing on the pronunciation recording audio of all the English letters corresponding to each subject in the same gender subject to obtain first spliced audio data corresponding to the English proper nouns of the vehicle-mounted function modules; performing audio splicing on the pronunciation recording audio of all the English letters corresponding to multiple subjects in the same gender subject to obtain second spliced audio data corresponding to the English proper nouns of the vehicle-mounted function modules; adding background noise to the first spliced audio data and the second spliced audio data; and generating the English proper noun training data of the vehicle-mounted function modules corresponding to different gender subjects based on the first spliced audio data and the second spliced audio data after adding the background noise.
[0033] Exemplarily, in the embodiment of the present application, the pronunciation recording audio of the full amount of English letters corresponding to each gender subject is spliced, which can be the pronunciation recording audio of the full amount of English letters corresponding to each subject in each gender subject, such as only the pronunciation recording audio "a-z" of 26 letters of any subject, i.e. a speaker, is spliced to obtain the English proper noun data, or the pronunciation recording audio of multiple subjects in the same gender subject is spliced, such as the pronunciation of multiple letters "a-z" in the pronunciation recording audio corresponding to two male subjects in the male category is mixed to obtain the English proper noun audio splicing data of the vehicle function module, wherein various types of vehicle background noise can be added during the pronunciation recording audio splicing process, such as environmental noise, noise generated by friction between the car and the air during driving, etc., which are only examples, or the vehicle background noise can be added after the pronunciation recording audio splicing is completed, which is determined by the actual situation. The embodiment of the present application adds vehicle background noise to the spliced audio data for data enhancement processing, which improves the accuracy of English proper noun recognition.
[0034] As an optional embodiment of the present application, the method further comprises: performing data enhancement processing on the pronunciation recording audio of the full amount of English letters of the different gender subjects and the corresponding English proper noun training data of the vehicle function module.
[0035] Exemplarily, in the embodiment of the present application, the pronunciation recording audio of the full amount of English letters of the different gender subjects and the corresponding English proper noun training data of the vehicle function module obtained by the above optional embodiment are subjected to data enhancement, such as speed disturbance, 1.5 times speed processing, spectral enhancement or spectral replacement, etc. The actual situation can be determined by itself. The embodiment of the present application uses the English proper noun training data and the pronunciation recording audio of the full amount of English letters of the different gender subjects after data enhancement to train the vehicle voice recognition model, which improves the accuracy of the vehicle voice recognition model in recognizing English proper nouns.
[0036] As an optional embodiment of the present application, the number of pronunciation recording audio samples of the full amount of English letters corresponding to different gender subjects is the same.
[0037] Exemplarily, in the embodiment of the application, the number of pronunciation recording audio samples of the full amount of English letters corresponding to the male and female subjects and the child category subjects is the same, for example, the number of pronunciation recording audio samples corresponding to the male subject is set to 100, and the number of pronunciation recording audio samples corresponding to the female subject and the child category subject is also set to 100. Because the audio of the male, female, and child is different, setting the number of recording audio samples of the male, female, and child category to be the same can ensure that the accuracy of recognizing the proper nouns input by the male user, recognizing the proper nouns input by the female user, and recognizing the proper nouns input by the child category user is the same when training the vehicle-mounted voice recognition model, thereby improving the accuracy of recognizing English proper nouns.
[0038] The embodiment of the application further discloses a vehicle-mounted voice recognition method, as shown in the method comprises the following steps: Figure 2
[0039] Step 201, in response to the receiving operation of the voice data; exemplarily, the voice recognition system in the embodiment of the application receives the voice data in real time after being enabled.
[0040] Step 202, when the voice data is received, the voice data is recognized by using the vehicle-mounted voice recognition model, and the vehicle-mounted voice recognition model is trained by using the data generated by the data generation method in the above embodiment; exemplarily, after receiving the voice data input by the vehicle-mounted user, the voice data is recognized by using the integrated vehicle-mounted voice recognition model, and the way of model training based on the data generated by the data generation method in the above embodiment is not limited in the embodiment of the application, and those skilled in the art can train the vehicle-mounted voice recognition model that meets the recognition requirements based on the generated data according to actual needs.
[0041] Step 203, determining whether the corresponding vehicle-mounted function module needs to be enabled according to the recognition result; exemplarily, the voice data is recognized by using the vehicle-mounted voice recognition model based on the above step to obtain the recognition result, and it is determined whether there is a corresponding vehicle-mounted function module in the vehicle-mounted system according to the recognition result, if there is, it is determined that the corresponding vehicle-mounted function module in the vehicle-mounted system needs to be enabled, so that the enabled vehicle-mounted function module performs the corresponding function.
[0042] The vehicle-mounted voice recognition method provided by the application receives the voice data input by the user, recognizes the voice data by using the vehicle-mounted voice recognition model, and determines whether the corresponding vehicle-mounted function module needs to be enabled according to the recognition result, thereby improving the intelligent degree and accuracy of the voice recognition mode of the voice recognition system.
[0043] The embodiment of the application further discloses a data generation device, as shown in the data generation device comprises the following steps: Figure 3 As shown, the device comprises: a recording audio acquisition module 301, configured to acquire recording audio of pronunciation of full-quantity English letters corresponding to different gender subjects; and an audio splicing module 302, configured to perform audio splicing based on the recording audio of pronunciation of full-quantity English letters corresponding to each gender subject, to obtain English proper noun training data of a vehicle-mounted function module corresponding to different gender subjects, wherein the recording audio of pronunciation of full-quantity English letters of the different gender subjects and the English proper noun training data of the vehicle-mounted function module corresponding thereto are used to train a vehicle-mounted voice recognition model.
[0044] The data generation device provided by the application increases the types and quantity of data used to train the vehicle-mounted voice recognition model, and improves the recognition accuracy of the vehicle-mounted voice recognition system integrated with the vehicle-mounted voice recognition model for pronunciation of English proper nouns of letters.
[0045] As an optional embodiment of the application, the audio splicing module comprises: a first audio splicing module, configured to perform audio splicing on the recording audio of pronunciation of full-quantity English letters corresponding to each subject in the same gender subject, to obtain first spliced audio data corresponding to the English proper noun of the vehicle-mounted function module; a second audio splicing module, configured to perform audio splicing on the recording audio of pronunciation of full-quantity English letters corresponding to multiple subjects in the same gender subject, to obtain second spliced audio data corresponding to the English proper noun of the vehicle-mounted function module; a noise adding module, configured to add background noise to the first spliced audio data and the second spliced audio data; and a noun data generation module, configured to generate English proper noun training data of the vehicle-mounted function module corresponding to different gender subjects based on the first spliced audio data and the second spliced audio data after adding background noise.
[0046] As an optional embodiment of the application, the device further comprises a data enhancement module, configured to perform data enhancement processing on the recording audio of pronunciation of full-quantity English letters of the different gender subjects and the English proper noun training data corresponding thereto.
[0047] As an optional embodiment of the application, the number of samples of the recording audio of pronunciation of full-quantity English letters corresponding to different gender subjects is the same.
[0048] The application further provides a vehicle-mounted voice recognition device, as shown in Figure 4As shown, the device comprises: a data receiving module 401, configured to respond to a receiving operation of voice data; a voice data identification module 402, configured to identify the voice data by using a vehicle-mounted voice recognition model when the voice data is received, wherein the vehicle-mounted voice recognition model is trained by the data generated by the data generation method in the above embodiment; and a function module enabling module 403, configured to determine whether a corresponding vehicle-mounted function module needs to be enabled according to the identification result.
[0049] The vehicle-mounted voice recognition device provided by the application receives voice data input by a user, identifies the voice data by using a vehicle-mounted voice recognition model, and determines whether a corresponding vehicle-mounted function module needs to be enabled according to the identification result, thereby improving the intelligent degree and accuracy of the voice recognition mode of the voice recognition system.
[0050] The embodiment of the application further provides an electronic device, which can be used to implement the above method. Figure 5 As shown, the electronic device can comprise a processor 501 and a memory 502, wherein the processor 501 and the memory 502 can be connected by a bus or other means, Figure 5 For example, the connection by the bus.
[0051] The processor 501 can be a central processing unit (CPU). The processor 501 can also be other general-purpose processors, digital signal processors (DSP), application specific integrated circuits (ASIC), field-programmable gate arrays (FPGA) or other programmable logic devices, discrete gate or transistor logic devices, discrete hardware components, or combinations thereof.
[0052] The memory 502 is a non-transitory computer readable storage medium, which can be used to store non-transitory software programs, non-transitory computer executable programs and modules, such as the program instructions / modules corresponding to the data generation method or the vehicle-mounted voice recognition method in the embodiment of the application. The processor 501 performs various functional applications and data processing of the processor by running the non-transitory software programs, instructions and modules stored in the memory 502, that is, implements the data generation method or the vehicle-mounted voice recognition method in the above method embodiment.
[0053] The memory 502 can include a program storage area and a data storage area, where the program storage area can store an operating system, at least one application required by a function, and the data storage area can store data created by the processor 501 and the like. In addition, the memory 502 can include a high-speed random access memory, and can also include a non-transitory memory such as at least one disk storage device, a flash memory device, or other non-transitory solid-state memory device. In some embodiments, the memory 502 can optionally include a memory disposed remotely with respect to the processor 501, which can be connected to the processor 501 through a network. Examples of the above network include, but are not limited to, the Internet, an intranet, a local area network, a mobile communication network, and combinations thereof.
[0054] The one or more modules are stored in the memory 502 and, when executed by the processor 501, perform the data generation method as described in the embodiments, or the vehicle-mounted voice recognition method as described in the embodiments. Figure 1 The one or more modules are stored in the memory 502 and, when executed by the processor 501, perform the data generation method as described in the embodiments, or the vehicle-mounted voice recognition method as described in the embodiments. Figure 2 The one or more modules are stored in the memory 502 and, when executed by the processor 501, perform the data generation method as described in the embodiments, or the vehicle-mounted voice recognition method as described in the embodiments.
[0055] The above electronic device specific details can be understood with reference to the embodiments shown in the accompanying drawings or the corresponding related descriptions and effects in the embodiments, which will not be described here again. Figure 1 The above electronic device specific details can be understood with reference to the embodiments shown in the accompanying drawings or the corresponding related descriptions and effects in the embodiments, which will not be described here again. Figure 2 The above electronic device specific details can be understood with reference to the embodiments shown in the accompanying drawings or the corresponding related descriptions and effects in the embodiments, which will not be described here again.
[0056] Those skilled in the art can understand that all or part of the processes in the above-mentioned embodiments can be completed by a computer program instructing related hardware, and the program can be stored in a computer readable storage medium. When the program is executed, it can include the processes of the above-mentioned embodiments. The storage medium can be a disk, an optical disk, a read-only memory (ROM), a random access memory (RAM), a flash memory, a hard disk drive (HDD) or a solid state drive (SSD), etc. The storage medium can also include a combination of the above-mentioned types of memories.
[0057] Although the embodiments of the present application are described in conjunction with the accompanying drawings, various modifications and changes can be made by those skilled in the art without departing from the spirit and scope of the present application, and such modifications and changes fall within the scope defined.
Claims
1. A data generation method, characterized in that, The method includes: Obtain audio recordings of the pronunciation of multiple full English letters corresponding to subjects of different genders; Audio recordings of the pronunciation of all English letters corresponding to each gender are spliced together to obtain training data of English proper nouns for vehicle functional modules corresponding to different genders. The audio recordings of the pronunciation of all English letters for different genders and the training data of English proper nouns for corresponding vehicle functional modules are used to train the vehicle speech recognition model.
2. The data generation method according to claim 1, characterized in that, The audio splicing based on the pronunciation recordings of the full set of English letters corresponding to each gender includes: The audio recordings of the pronunciation of all English letters corresponding to each subject within the same gender are spliced together to obtain the first spliced audio data corresponding to the English proper nouns of the vehicle function modules. Audio recordings of the pronunciation of all English letters corresponding to multiple subjects within the same gender are spliced together to obtain the second spliced audio data corresponding to the English proper nouns of the vehicle function modules; Background noise is added to the first and second spliced audio data; Based on the first and second spliced audio data after adding background noise, training data of English proper nouns for vehicle function modules corresponding to different gender subjects are generated.
3. The data generation method according to claim 1, characterized in that, The method further includes: Data augmentation processing was performed on the audio recordings of the pronunciation of all English letters for the subjects of different genders, as well as the training data of the English proper nouns for the corresponding vehicle function modules.
4. The data generation method according to claim 1, characterized in that, The number of audio samples recording the pronunciation of all English letters is the same for subjects of different genders.
5. A vehicle-mounted voice recognition method, characterized in that, The method includes: Respond to voice data reception operations; When voice data is received, the in-vehicle voice recognition model is used to recognize the voice data. The in-vehicle voice recognition model is trained based on the data generated by the data generation method according to any one of claims 1-4. The identification results will determine whether the corresponding vehicle function modules need to be activated.
6. A data generation apparatus, characterized in that, The device includes: The audio recording acquisition module is used to acquire the pronunciation recordings of multiple full English letters corresponding to subjects of different genders; The audio splicing module is used to splice audio based on the pronunciation recordings of all English letters corresponding to each gender subject, to obtain training data of English proper nouns for vehicle functional modules corresponding to different gender subjects. The audio recordings of the pronunciation of all English letters for different gender subjects and the corresponding training data of English proper nouns for vehicle functional modules are used to train the vehicle speech recognition model.
7. The data generation apparatus according to claim 6, characterized in that, The audio splicing module includes: The first audio splicing module is used to splice the audio recordings of the pronunciation of all English letters corresponding to each subject within the same gender subject, and obtain the first spliced audio data corresponding to the English proper nouns of the vehicle function module. The second audio splicing module is used to splice the audio recordings of the pronunciation of all English letters corresponding to multiple subjects within the same gender to obtain the second spliced audio data corresponding to the English proper nouns of the vehicle function modules. A noise addition module is used to add background noise to the first spliced audio data and the second spliced audio data; The noun data generation module is used to generate English proper noun training data for vehicle function modules corresponding to different genders, based on the first and second spliced audio data after adding background noise.
8. A vehicle-mounted voice recognition device, characterized in that, The device includes: The data receiving module is used to respond to voice data receiving operations; A voice data recognition module is used to recognize voice data when it is received using an in-vehicle voice recognition model, wherein the in-vehicle voice recognition model is trained based on data generated by the data generation method according to any one of claims 1-4. The function module activation module is used to determine whether the corresponding vehicle function module needs to be activated based on the recognition results.
9. An electronic device, characterized in that, include: At least one processor; And a memory communicatively connected to the at least one processor; wherein the memory stores instructions executable by the at least one processor, the instructions being executed by the at least one processor to cause the at least one processor to perform the steps of the data generation method as described in any one of claims 1-4, or the steps of the vehicle-mounted voice recognition method as described in claim 5.
10. A computer-readable storage medium having a computer program stored thereon, characterized in that, When the computer program is executed by the processor, it implements the steps of the data generation method as described in any one of claims 1-4, or the steps of the vehicle-mounted voice recognition method as described in claim 5.
Citation Information
Patent Citations
Voice recognition data expansion method and system
CN111354346A
Speech synthesis using deep neural networks
US8527276B1