A vehicle-mounted speech recognition method, device, equipment and storage medium
By setting up sound zones in the vehicle, using microphones to collect audio signals, and performing feature processing and neural network recognition, the accuracy problem of multi-speaker speech recognition in the vehicle was solved, and efficient recognition of mixed speech was achieved.
Patent Information
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- CHINA FAW CO LTD
- Filing Date
- 2023-06-02
- Publication Date
- 2026-07-28
AI Technical Summary
Existing technologies cannot guarantee good speech recognition accuracy in in-vehicle multi-speaker speech recognition tasks, especially in complex acoustic environments, where they cannot effectively recognize mixed speech.
By setting multiple sound zones in the vehicle, multiple raw audio signals are collected using the vehicle microphone, signal processing is performed to obtain mixed sound zone features, which are then input into a preset sound zone coding recognition neural network and a speech recognition network to identify the coding features of each sound zone and determine the target speech recognition result.
It improves the accuracy and efficiency of in-vehicle mixed speech recognition, and can accurately recognize the speech of multiple speakers in complex acoustic environments.
Smart Images

Figure CN116580713B_ABST
Abstract
Description
Technical Field
[0001] The present invention relates to the field of autonomous driving technology, and in particular to an in-vehicle voice recognition method, device, equipment and storage medium. Background Technology
[0002] Currently, speech recognition technology has made significant progress in quiet environments, demonstrating high accuracy and stability in single-speaker speech recognition tasks. However, it remains challenging in multi-speaker speech recognition tasks facing complex acoustic environments, and the results have not yet reached a satisfactory level. For example, existing technologies achieve high accuracy in recognizing single speech within a vehicle, but they cannot guarantee good accuracy when recognizing mixed speech in situations involving multiple speakers within the vehicle. Summary of the Invention
[0003] This invention provides an in-vehicle voice recognition method, apparatus, device, and storage medium, which can improve the accuracy and efficiency of recognizing mixed voices in vehicles.
[0004] In a first aspect, embodiments of the present invention provide an in-vehicle voice recognition method, the method comprising:
[0005] Acquire multiple raw audio signals collected by vehicle-mounted microphones in each sound zone of the target vehicle, and perform signal processing on the multiple raw audio signals to obtain mixed sound zone characteristics;
[0006] The mixed vocal region features are input into a preset vocal region encoding and recognition neural network to obtain the encoding features of each vocal region.
[0007] The encoded features of each vocal region are input into a preset speech recognition network to obtain the speech recognition text content of each vocal region, and the target speech recognition result is determined based on the speech recognition text content of each vocal region.
[0008] Secondly, embodiments of the present invention provide an in-vehicle voice recognition device, the device comprising:
[0009] The voice signal acquisition module is used to acquire multiple raw audio signals collected by the vehicle microphones in each sound zone of the target vehicle, and to perform signal processing on the multiple raw audio signals to obtain mixed sound zone features.
[0010] The vocal region coding feature determination module is used to input the mixed vocal region features into a preset vocal region coding recognition neural network to obtain the coding features of each vocal region.
[0011] The speech recognition result determination module is used to input the encoded features of each voice region into a preset speech recognition network to obtain the speech recognition text content of each voice region, and determine the target speech recognition result based on the speech recognition text content of each voice region.
[0012] Thirdly, embodiments of the present invention provide a computer device, the computer device comprising:
[0013] One or more processors;
[0014] Memory, used to store one or more programs;
[0015] When the one or more programs are executed by the one or more processors, the one or more processors implement the in-vehicle voice recognition method described in any embodiment.
[0016] Fourthly, embodiments of the present invention provide a computer-readable storage medium having a computer program stored thereon, which, when executed by a processor, implements the vehicle-mounted voice recognition method described in any embodiment.
[0017] The technical solution provided by this invention involves acquiring multiple raw audio signals collected by in-vehicle microphones in various audio zones of a target vehicle, processing these raw audio signals to obtain mixed audio zone features, inputting these mixed audio zone features into a preset audio zone encoding and recognition neural network to obtain encoding features for each audio zone, inputting these encoding features into a preset speech recognition network to obtain speech recognition text content for each audio zone, and determining the target speech recognition result based on the speech recognition text content for each audio zone. This technical solution solves the problem of inaccurate and inefficient recognition of mixed in-vehicle speech in existing technologies, improving the accuracy and efficiency of mixed in-vehicle speech recognition. Attached Figure Description
[0018] Figure 1 This is a flowchart of an in-vehicle voice recognition method provided by an embodiment of the present invention;
[0019] Figure 2 This is a flowchart of another vehicle-mounted voice recognition method provided in an embodiment of the present invention;
[0020] Figure 3 This is a schematic diagram of a vehicle-mounted microphone acquiring multiple raw audio signals according to an embodiment of the present invention;
[0021] Figure 4 This is a flowchart of a method for performing in-vehicle voice recognition according to an embodiment of the present invention;
[0022] Figure 5 This is a schematic diagram of the structure of an in-vehicle voice recognition device provided in an embodiment of the present invention;
[0023] Figure 6 This is a schematic diagram of the structure of a computer device provided in an embodiment of the present invention. Detailed Implementation
[0024] To make the objectives, technical solutions, and advantages of the embodiments of the present invention clearer, the technical solutions of the embodiments of the present invention will be clearly and completely described below with reference to the accompanying drawings. Obviously, the described embodiments are only some embodiments of the present invention, not all embodiments. Based on the embodiments of the present invention, all other embodiments obtained by those skilled in the art without creative effort are within the scope of protection of the present invention.
[0025] Figure 1 This is a flowchart of an in-vehicle voice recognition method provided by an embodiment of the present invention. The embodiment of the present invention can be applied to scenarios where mixed voices in a vehicle are recognized. The method can be executed by an in-vehicle voice recognition device, which can be implemented by software and / or hardware.
[0026] like Figure 1 As shown, the in-vehicle voice recognition method includes the following steps:
[0027] S110. Acquire multiple original audio signals collected by the vehicle-mounted microphones in each sound zone of the target vehicle, and perform signal processing on the multiple original audio signals to obtain mixed sound zone characteristics.
[0028] The target vehicle can be any vehicle requiring in-vehicle voice recognition. The audio zone can be a pre-defined area for receiving audio signals. Specifically, multiple audio zones can be set up within the target vehicle, and an in-vehicle microphone can be installed in each zone to receive the audio signal from the speaker. The multiple raw audio signals can be the audio signals collected by the various in-vehicle microphones within the target vehicle. Specifically, in-vehicle microphones can be installed in different areas within the target vehicle, and the audio signals collected by each microphone can be used as multiple raw audio signals. The audio signal can be the voice signal emitted by the speaker within the target vehicle.
[0029] Mixed-range characteristics can describe the attribute features of the raw audio signals collected by the vehicle-mounted microphone in each range. Specifically, signal processing can be performed on multiple raw audio signals to obtain mixed-range characteristics. For example, a Fast Fourier Transform can be performed on multiple raw audio signals to obtain corresponding time-spectrum graphs, then the time-spectrum graphs can be superimposed and averaged, and the result of the superposition and averaging can be used as the mixed-range characteristics.
[0030] S120. Input the mixed vocal region features into a preset vocal region coding and recognition neural network to obtain the coding features of each vocal region.
[0031] The preset encoding recognition neural network can be a pre-defined neural network that identifies the encoded features of mixed vocal ranges. Since the preset encoding recognition neural network needs to identify the encoded features of mixed vocal ranges, it needs to have a good ability to analyze these features. Specifically, it can be obtained by training a model based on preset training samples of mixed vocal range features.
[0032] Furthermore, the vocal region coding features can be the coding feature data of audio signals in each vocal region. Specifically, mixed vocal region features can be input into a preset vocal region coding recognition neural network to obtain the coding features of each vocal region. These coding feature data, through encoding, can establish an intrinsic connection between the audio signal and the text content, facilitating subsequent recognition of the audio signal to obtain the corresponding text content.
[0033] S130. Input the coding features of each voice region into a preset speech recognition network to obtain the speech recognition text content of each voice region, and determine the target speech recognition result based on the speech recognition text content of each voice region.
[0034] The preset speech recognition network can be a neural network that recognizes speech text content corresponding to preset encoding features of each voice region. The speech text content for each voice region can be the text obtained by performing speech recognition on the encoding features of each voice region.
[0035] Specifically, the encoded features of each vocal region can be input into a preset speech recognition network to obtain the speech recognition text content for each vocal region. Furthermore, when performing speech recognition on the encoded features of each vocal region, the preset speech recognition network can also output the confidence level of the speech recognition text content for each vocal region, which facilitates the identification of text with higher recognition accuracy from the speech recognition text content of each vocal region in subsequent processes.
[0036] The target speech recognition result can be the final speech recognition result determined after performing speech recognition on multiple raw audio signals. Specifically, the target speech recognition result can be determined based on the recognition accuracy of the speech recognition text content in each voice region. For example, the confidence level of the speech recognition text content in each voice region can be compared with a preset confidence reference threshold. The speech recognition text content in the voice region with a confidence level greater than the preset confidence reference threshold can be used as the target speech recognition result. When the confidence level of the speech recognition text content in a voice region is greater than the preset confidence reference threshold, it indicates that the speech recognition text content in that voice region has high recognition accuracy. Therefore, speech recognition text content in the voice region with a confidence level greater than the preset confidence reference threshold can be used as the target speech recognition result to improve the accuracy of speech recognition.
[0037] In one alternative implementation, after obtaining the target speech recognition result, the target speech recognition result can also be displayed on the vehicle's in-vehicle display screen to improve the visibility of the target speech recognition result.
[0038] The technical solution provided by this invention acquires multiple raw audio signals collected by in-vehicle microphones in each audio region of the target vehicle, and performs signal processing on these raw audio signals to obtain mixed audio region features. These mixed audio region features are then input into a preset audio region encoding and recognition neural network to obtain the encoding features for each audio region. Finally, these encoding features are input into a preset speech recognition network to obtain the speech recognition text content for each audio region, and the target speech recognition result is determined based on the speech recognition text content for each audio region. This technical solution solves the problem of inaccurate and inefficient recognition of mixed in-vehicle speech in existing technologies, and can improve the accuracy and efficiency of mixed in-vehicle speech recognition.
[0039] Figure 2 This is a flowchart of another in-vehicle voice recognition method provided by an embodiment of the present invention. This embodiment is applicable to scenarios involving the recognition of mixed voice within a vehicle. Based on the above embodiments, this embodiment further explains how to perform signal processing on the multiple original audio signals to obtain mixed voice region features, and how to determine the target voice recognition result based on the voice recognition text content of each voice region. This device can be implemented by software and / or hardware and integrated into a computer device with application development capabilities.
[0040] like Figure 2 As shown, the in-vehicle voice recognition method includes the following steps:
[0041] S210. Acquire multiple original audio signals collected by the vehicle-mounted microphones in each sound zone of the target vehicle, and perform fast Fourier transform on the multiple original audio signals to obtain the corresponding time-frequency spectrum.
[0042] The target vehicle can be any vehicle requiring in-vehicle voice recognition. The audio zone can be a pre-defined area for receiving audio signals. Specifically, multiple audio zones can be set up within the target vehicle, and an in-vehicle microphone can be installed in each zone to receive the audio signal from the speaker. The multiple raw audio signals can be the audio signals collected by the various in-vehicle microphones within the target vehicle. Specifically, in-vehicle microphones can be installed in different areas within the target vehicle, and the audio signals collected by each microphone can be used as multiple raw audio signals. The audio signal can be the voice signal emitted by the speaker within the target vehicle.
[0043] For example, Figure 3 This is a schematic diagram of a vehicle-mounted microphone acquiring multiple raw audio signals according to an embodiment of the present invention, as shown below. Figure 3As shown, microphones are set up in four areas (four sound zones) in the target vehicle to collect voice signals. There are two voice-producing objects, speaker A and speaker D, in the figure. Each microphone can collect the audio signals emitted by the two voice-producing objects, and finally obtain four raw audio signals. Each raw audio signal includes the audio signals of the two voice-producing objects, speaker A and speaker D.
[0044] The Fast Fourier Transform (FFT) is a data processing method that rapidly calculates the Discrete Fourier Transform (DFT) or inverse DFT of a sequence. The FFT can transform a signal from its original domain (usually time or space) to its frequency domain representation, or vice versa. A time-spectrum graph can be a graph showing the frequency and amplitude changes of multiple original audio signals over time. Specifically, after acquiring multiple original audio signals, a FFT can be performed on them to transform them from the time domain to the frequency domain, obtaining the corresponding time-spectrum graph.
[0045] S220. The time-spectrum diagrams are superimposed and the mean is calculated to obtain a mixed average time-spectrum diagram, which is used as the feature of the mixed sound region.
[0046] The mixed average time-spectrum graph can be a time-spectrum graph obtained by averaging the values in each of the individual time-spectrum graphs. Specifically, the individual time-spectrum graphs can be superimposed and averaged to obtain the mixed average time-spectrum graph. The mixed sound region features can be attribute features describing the original audio signals collected by the vehicle-mounted microphone in each sound region. Specifically, the mixed average time-spectrum graph calculated above can be used as the mixed sound region features.
[0047] S230. Input the mixed vocal region features into a preset vocal region coding and recognition neural network to obtain the coding features of each vocal region.
[0048] The preset encoding recognition neural network can be a preset neural network that identifies the encoded features of mixed vocal ranges. Specifically, since the preset encoding recognition neural network needs to identify the encoded features of multi-range audio signals, it needs to have a good ability to analyze the features of mixed vocal ranges.
[0049] The preset voice region coding recognition neural network can be obtained through pre-training. For example, the training process of the preset voice region coding recognition neural network includes: taking the mixed average time-spectrum diagram of multiple audio sample signals collected by the vehicle microphones of each voice region in the target vehicle when a voice emits speech as a first-class sample, and taking the time-spectrum diagram of the audio sample signal collected by the vehicle microphone in the voice region of the voice emitter and the corresponding number of pure noise audio time-spectrum diagrams as a first-class sample label; training the model based on the first-class samples and the first-class sample labels to obtain the basic voice region coding recognition neural network; taking the mixed average time-spectrum diagram of multiple audio sample signals collected by the vehicle microphones of each voice region in the target vehicle when multiple voice emitters speech as a second-class sample, and taking the time-spectrum diagram of the audio sample signal collected by the vehicle microphones in the voice regions of multiple voice emitters and the corresponding number of pure noise audio time-spectrum diagrams as a second-class sample label; training the model based on the second-class samples and the second-class sample labels to obtain the preset voice region coding recognition neural network.
[0050] One type of sample can be a mixed average time-spectrum map sample obtained based on a single speaker. The pure noise audio time-spectrum map can be the time-spectrum map of noise sound when no spoken speech is received. Since there are situations where the vehicle microphone does not receive audio sample signals when a single person is speaking, the time-spectrum map of noise sound when no spoken speech is received by the vehicle microphone can also be used as a sample label. Using the pure noise audio time-spectrum map as the background time-spectrum map can improve the noise resistance in the speech recognition process and improve the accuracy of the pre-trained preset voice region encoding recognition neural network in voice region encoding recognition. Specifically, the number of pure noise audio time-spectrum maps can be determined based on the number of vehicle microphones in each voice region and the number of speakers. The number of pure noise audio time-spectrum maps can be equal to the number of vehicle microphones in each voice region minus the number of speakers. For example, if there is a vehicle microphone in each of the four audio zones of the target vehicle, and the speaker is a single person speaking in audio zone 1, then there are three pure noise audio time spectrum diagrams. That is, the three pure noise audio time spectrum diagrams are pure noise audio time spectrum diagrams for audio zones 2, 3, and 4.
[0051] A basic vocal range coding recognition neural network can be obtained by training on sample data from a single vocal range. Specifically, the time-spectrum graphs of audio sample signals collected by a vehicle microphone in the vocal range of the vocal range, along with the corresponding number of pure noise audio time-spectrum graphs, can be used as one type of sample label. The model is then trained based on this type of sample and its label to obtain the basic vocal range coding recognition neural network. This network exhibits good accuracy in recognizing the coding features of the mixed average time-spectrum graph of a single vocal range, but it lacks the ability to effectively recognize the coding features of the mixed average time-spectrum graph of multiple vocal ranges.
[0052] The second type of samples can be mixed average time-spectrum samples obtained from multiple speaking objects. Furthermore, when there are multiple speaking objects, the number of pure noise audio time-spectrum samples will be reduced compared to the first type of samples. For example, when there is a vehicle microphone in each of the four audio regions of the target vehicle, and there are three speaking objects emitting speech in regions 1, 2, and 3 respectively, the number of pure noise audio time-spectrum samples will become only one, namely, the pure noise audio time-spectrum sample in region 4. After obtaining the second type of samples, the time-spectrum samples of the audio sample signals collected by the vehicle microphones in the multiple speaking object regions, along with the corresponding number of pure noise audio time-spectrum samples, can be used as second-type sample labels. Based on the second type of samples and their labels, a basic audio region coding recognition neural network is trained to obtain the preset audio region coding recognition neural network. Training a basic vocal range coding recognition neural network based on two types of samples and their labels can enable the trained neural network to achieve good recognition accuracy when identifying the coding features of the mixed average time-spectrum of multiple vocal objects.
[0053] In one optional implementation, the process of training the neural network based on training samples and sample labels includes: inputting training samples into the neural network to be trained, obtaining output results, then calculating the similarity between the output results and the sample labels, and updating the weights of the pre-trained model based on the similarity through backpropagation to obtain the trained neural network.
[0054] In one alternative implementation, the model can be trained based on the ResNet50 neural network to obtain a preset vocal region encoding recognition neural network.
[0055] Furthermore, the vocal region coding features can be the coding feature data of audio signals in each vocal region. Specifically, mixed vocal region features can be input into a preset vocal region coding recognition neural network to obtain the coding features of each vocal region. These coding feature data, through encoding, can establish an intrinsic connection between the audio signal and the text content, facilitating subsequent recognition of the audio signal to obtain the corresponding text content.
[0056] S240. Input the encoding features of each vocal region into a preset speech recognition network to obtain the speech recognition text content of each vocal region.
[0057] The preset speech recognition network can be a neural network that recognizes speech text content corresponding to preset phonic region coding features. The preset phonic region coding recognition neural network can be obtained through pre-training. Specifically, training samples for the preset speech recognition network can be obtained based on the training samples of the preset phonic region coding recognition neural network and the preset speech recognition neural network itself. Then, the model is trained based on the obtained training samples to obtain the preset speech recognition network. For example, the training process of the preset speech recognition network includes: inputting first-class and second-class samples into the preset phonic region coding recognition neural network to obtain the corresponding phonic region coding features of each sample; using the phonic region coding features of each sample as model training samples for the speech recognition network, and using the text content of the multiple original audio signals corresponding to the first-class or second-class samples as sample labels to train the initial speech recognition model to obtain the preset speech recognition network.
[0058] In one alternative implementation, a preset speech recognition network can be obtained by training the model based on the Transformer-CTC neural network.
[0059] The text content for speech recognition in each vocal region can be obtained by performing speech recognition on the encoded features of each vocal region. Specifically, the encoded features of each vocal region can be input into a preset speech recognition network to obtain the speech recognition text content for each vocal region. Furthermore, when performing speech recognition on the encoded features of each vocal region, the preset speech recognition network can also output the confidence level of the speech recognition text content for each vocal region, which facilitates the identification of text with higher recognition accuracy from the speech recognition text content of each vocal region in subsequent processes.
[0060] S250. Compare the recognition confidence of the speech recognition text content of each voice region with the preset confidence reference threshold.
[0061] Among them, recognition confidence can be a parameter describing the credibility of the speech recognition text content. After recognizing the encoded features of each input voice region, the preset speech recognition network can obtain the speech recognition text content of each voice region and the recognition confidence of each voice region speech recognition text content. For example, the preset speech recognition network can output a recognized text L = {l1, l2, l3} and the confidence of each word in the text, that is, the confidence vector P = {p1, p2, p3}.
[0062] The preset confidence reference threshold can be a reference threshold for judging whether the recognition confidence is reliable. After obtaining the recognition confidence of the speech recognition text content in each audio region, the recognition confidence of the speech recognition text content in each audio region can be compared with the preset confidence reference threshold, so as to determine whether the speech recognition text content in each audio region is reliable.
[0063] S260. Use the speech recognition text content of the audio region with a recognition confidence greater than the preset confidence reference threshold as the target speech recognition result.
[0064] Among them, the target speech recognition result can be the speech recognition result finally determined after speech recognition of multiple original audio signals.
[0065] Specifically, after respectively comparing the recognition confidence of the speech recognition text content in each audio region with the preset confidence reference threshold, the speech recognition text content of the audio region with a recognition confidence greater than the preset confidence reference threshold can be used as the target speech recognition result. When the recognition confidence of the speech recognition text content in an audio region is greater than the preset confidence reference threshold, it indicates that the speech recognition text content in this audio region has a high recognition accuracy. Therefore, the speech recognition text content of the audio region with a recognition confidence greater than the preset confidence reference threshold can be used as the target speech recognition result to improve the accuracy of speech recognition.
[0066] Specifically, the recognition confidence of each character in the speech recognition text content of the audio region can be averaged to obtain an average confidence, and then the average confidence is compared with the preset confidence reference threshold, and the speech recognition text content of the audio region corresponding to the average confidence greater than the preset confidence reference threshold is used as the target speech recognition result.
[0067] For example, the speech recognition text content of audio region A is L = {"The weather is nice today"}, and the recognition confidence P = {0.81, 0.82, 0.83, 0.84, 0.85, 0.86}.
[0068] The speech recognition text content of audio region B is L = {"Today"}, and the recognition confidence P = {0.2}.
[0069] The speech recognition text content of audio region C is L = {"Not"}, and the recognition confidence P = {0.1}.
[0070] The speech recognition text content of audio region D is L = {"It's really quite nice"}, and the recognition confidence P = {0.91, 0.92, 0.93, 0.94, 0.95}.
[0071] For the obtained recognition text and confidence vector, the average value of the confidence vector of each text can be calculated. For example:
[0072] The mean confidence vector Pa for L = {The weather is nice today} in register A is 0.835.
[0073] The mean of the confidence vector Pa for register B is 0.2.
[0074] The mean of the L = {not} confidence vector for register C is Pa = 0.1.
[0075] The mean confidence vector Pa for the L = {Really pretty good} in the D register is 0.93.
[0076] Because the average confidence level Pa of the recognition results for voice region A is greater than the preset confidence reference threshold of 0.75, and the same applies to voice region D, the final target speech recognition results are: Speaker A: The weather is nice today; Speaker D: It's really quite nice.
[0077] In one alternative implementation, after obtaining the target speech recognition result, the target speech recognition result can also be displayed on the vehicle's in-vehicle display screen to improve the visibility of the target speech recognition result.
[0078] For example, Figure 4 This is a flowchart of a method for performing in-vehicle voice recognition according to an embodiment of the present invention, as shown below. Figure 4 As shown, the method for in-vehicle voice recognition includes: inputting original audio 1, original audio 2, original audio 3, and original audio 4 into a sound region encoding module, which identifies the sound region features of each original audio to obtain corresponding sound region feature 1, sound region feature 2, sound region feature 3, and sound region feature 4; then, inputting sound region feature 1, sound region feature 2, sound region feature 3, and sound region feature 4 into a voice recognition module, which identifies the sound region features to obtain corresponding recognition result 1, recognition result 2, recognition result 3, and recognition result 4; further, inputting recognition result 1, recognition result 2, recognition result 3, and recognition result 4 into a text determination module, which determines the target recognition result based on the confidence level of each recognition result; finally, displaying the target recognition result on the in-vehicle display screen of the target vehicle to complete the in-vehicle voice recognition process.
[0079] The technical solution provided by this invention involves acquiring multiple raw audio signals collected by in-vehicle microphones in each audio region of a target vehicle, performing Fast Fourier Transform on each raw audio signal to obtain corresponding time-spectrum maps, superimposing the time-spectrum maps and calculating their average to obtain a mixed average time-spectrum map, which serves as the mixed audio region feature. The mixed audio region feature is then input into a preset audio region encoding and recognition neural network to obtain the encoding features for each audio region. These encoding features are further input into a preset speech recognition network to obtain the speech recognition text content for each audio region. The recognition confidence of the speech recognition text content for each audio region is compared with a preset confidence reference threshold. The speech recognition text content for the audio region with a recognition confidence greater than the preset confidence reference threshold is taken as the target speech recognition result. This technical solution solves the problem of inaccurate and efficient recognition of mixed in-vehicle speech in existing technologies, improving the accuracy and efficiency of mixed in-vehicle speech recognition.
[0080] Figure 5 This is a schematic diagram of the structure of an in-vehicle voice recognition device provided in an embodiment of the present invention. The embodiment of the present invention can be applied to scenarios of recognizing mixed voices in a vehicle. The device can be implemented by software and / or hardware and integrated into a computer device with application development capabilities.
[0081] like Figure 5 As shown, the vehicle-mounted voice recognition device includes: a voice signal acquisition module 310, a voice region coding feature determination module 320, and a voice recognition result determination module 330.
[0082] The speech signal acquisition module 310 is used to acquire multiple original audio signals collected by the vehicle microphones in each sound zone of the target vehicle, and to perform signal processing on the multiple original audio signals to obtain mixed sound zone features; the sound zone coding feature determination module 320 is used to input the mixed sound zone features into a preset sound zone coding recognition neural network to obtain the sound zone coding features; the speech recognition result determination module 330 is used to input the sound zone coding features into a preset speech recognition network to obtain the speech recognition text content of each sound zone, and to determine the target speech recognition result based on the speech recognition text content of each sound zone.
[0083] The technical solution provided by this invention involves acquiring multiple raw audio signals collected by in-vehicle microphones in various audio zones of a target vehicle, processing these raw audio signals to obtain mixed audio zone features, inputting these mixed audio zone features into a preset audio zone encoding and recognition neural network to obtain encoding features for each audio zone, inputting these encoding features into a preset speech recognition network to obtain speech recognition text content for each audio zone, and determining the target speech recognition result based on the speech recognition text content for each audio zone. This technical solution solves the problem of inaccurate and inefficient recognition of mixed in-vehicle speech in existing technologies, improving the accuracy and efficiency of mixed in-vehicle speech recognition.
[0084] In one optional implementation, the speech signal acquisition module 310 is specifically used to: perform fast Fourier transform on the multiple original audio signals respectively to obtain corresponding time-spectrum diagrams; superimpose the time-spectrum diagrams and calculate the mean to obtain a mixed average time-spectrum diagram, which serves as the mixed voice region feature.
[0085] In one optional implementation, the speech recognition result determination module 320 is specifically used to: compare the recognition confidence of the speech recognition text content of each voice region with a preset confidence reference threshold; and take the speech recognition text content of the voice region with a recognition confidence greater than the preset confidence reference threshold as the target speech recognition result.
[0086] In an optional embodiment, the vehicle-mounted voice recognition device further includes: a preset voice region encoding recognition neural network training module, configured to: use the mixed average time-spectrum diagram of multiple audio sample signals collected by vehicle-mounted microphones in each voice region of the target vehicle when a voice emits speech as a first-class sample, and use the time-spectrum diagram of the audio sample signal collected by the vehicle-mounted microphone in the voice region of the voice emitter and the corresponding number of pure noise audio time-spectrum diagrams as a first-class sample label; perform model training based on the first-class samples and the first-class sample labels to obtain a basic voice region encoding recognition neural network; use the mixed average time-spectrum diagram of multiple audio sample signals collected by vehicle-mounted microphones in each voice region of the target vehicle when multiple voice emitters speech as a second-class sample, and use the time-spectrum diagram of the audio sample signal collected by the vehicle-mounted microphones in the voice regions of the multiple voice emitters and the corresponding number of pure noise audio time-spectrum diagrams as a second-class sample label; perform model training based on the second-class samples and the second-class sample labels to obtain the preset voice region encoding recognition neural network.
[0087] In one optional embodiment, the vehicle-mounted voice recognition device further includes: a preset voice recognition network training module, configured to: input the first-class samples and the second-class samples into the preset voice region encoding recognition neural network respectively to obtain corresponding voice region encoding features for each sample; use the voice region encoding features for each sample as model training samples for the voice recognition network, and use the text content of the multiple original audio signals corresponding to the first-class samples or the second-class samples as sample labels to train the initial voice recognition model to obtain the preset voice recognition network.
[0088] In one alternative implementation, the preset speech recognition network is a model trained based on the Transformer-CTC model.
[0089] In one optional embodiment, the in-vehicle voice recognition device further includes a voice recognition result display module, used to display the target voice recognition result on the in-vehicle display screen of the target vehicle.
[0090] The vehicle-mounted voice recognition device provided in the embodiments of the present invention can execute the vehicle-mounted voice recognition method provided in any embodiment of the present invention, and has the corresponding functional modules and beneficial effects of the method.
[0091] Figure 6 This is a schematic diagram of the structure of a computer device provided in an embodiment of the present invention. Figure 6 A block diagram of an exemplary computer device 12 suitable for implementing embodiments of the present invention is shown. Figure 6 The computer device 12 shown is merely an example and should not be construed as limiting the functionality and scope of the embodiments of the present invention. The computer device 12 can be any terminal device with computing capabilities and can be integrated into an in-vehicle voice recognition device.
[0092] like Figure 6 As shown, the computer device 12 is represented in the form of a general-purpose computing device. The components of the computer device 12 may include, but are not limited to: one or more processors or processing units 16, system memory 28, and a bus 18 connecting different system components (including system memory 28 and processing unit 16).
[0093] Bus 18 can be one or more of several bus architectures, including a memory bus or memory controller, a peripheral bus, a graphics acceleration port, a processor, or a local bus using any of the various bus architectures. For example, these architectures include, but are not limited to, the Industry Standard Architecture (ISA) bus, the Micro Channel Architecture (MAC) bus, the Enhanced ISA bus, the Video Electronics Standards Association (VESA) local bus, and the Peripheral Component Interconnect (PCI) bus.
[0094] Computer device 12 typically includes a variety of computer system readable media. These media can be any available media that can be accessed by computer device 12, including volatile and non-volatile media, removable and non-removable media.
[0095] System memory 28 may include computer system readable media in the form of volatile memory, such as random access memory (RAM) 30 and / or cache 32. Computer device 12 may further include other removable / non-removable, volatile / non-volatile computer system storage media. By way of example only, storage system 34 may be used to read and write non-removable, non-volatile magnetic media (…). Figure 6 Not shown; usually referred to as a "hard drive"). Although Figure 6 Not shown, a disk drive for reading and writing to a removable non-volatile disk (e.g., a "floppy disk") and an optical disk drive for reading and writing to a removable non-volatile optical disk (e.g., a CD-ROM, DVD-ROM, or other optical media) may be provided. In these cases, each drive may be connected to bus 18 via one or more data media interfaces. System memory 28 may include at least one program product having a set (e.g., at least one) of program modules configured to perform the functions of the embodiments of the present invention.
[0096] A program / utility 40 having a set (at least one) of program modules 42 may be stored, for example, in system memory 28. Such program modules 42 include, but are not limited to, an operating system, one or more application programs, other program modules, and program data. Each or some combination of these examples may include an implementation of a network environment. Program modules 42 typically perform the functions and / or methods described in the embodiments of the present invention.
[0097] Computer device 12 can also communicate with one or more external devices 14 (e.g., keyboard, pointing device, display 24, etc.), and with one or more devices that enable a user to interact with the computer device 12, and / or with any device that enables the computer device 12 to communicate with one or more other computing devices (e.g., network card, modem, etc.). This communication can be performed via input / output (I / O) interface 22. Furthermore, computer device 12 can also communicate with one or more networks (e.g., local area network (LAN), wide area network (WAN), and / or public networks, such as the Internet) via network adapter 20. As shown, network adapter 20 communicates with other modules of computer device 12 via bus 18. It should be understood that, although... Figure 6As not shown, it can be used in conjunction with computer device 12 with other hardware and / or software modules, including but not limited to: microcode, device drivers, redundant processing units, external disk drive arrays, RAID systems, tape drives, and data backup storage systems.
[0098] Processing unit 16 executes various functional applications and data processing by running programs stored in system memory 28, such as implementing the in-vehicle voice recognition method provided in this embodiment, which includes:
[0099] Acquire multiple raw audio signals collected by vehicle-mounted microphones in each sound zone of the target vehicle, and perform signal processing on the multiple raw audio signals to obtain mixed sound zone characteristics;
[0100] The mixed vocal region features are input into a preset vocal region encoding and recognition neural network to obtain the encoding features of each vocal region.
[0101] The encoded features of each vocal region are input into a preset speech recognition network to obtain the speech recognition text content of each vocal region, and the target speech recognition result is determined based on the speech recognition text content of each vocal region.
[0102] This embodiment provides a computer-readable storage medium storing a computer program that, when executed by a processor, implements the in-vehicle voice recognition method as provided in any embodiment of the present invention, including:
[0103] Acquire multiple raw audio signals collected by vehicle-mounted microphones in each sound zone of the target vehicle, and perform signal processing on the multiple raw audio signals to obtain mixed sound zone characteristics;
[0104] The mixed vocal region features are input into a preset vocal region encoding and recognition neural network to obtain the encoding features of each vocal region.
[0105] The encoded features of each vocal region are input into a preset speech recognition network to obtain the speech recognition text content of each vocal region, and the target speech recognition result is determined based on the speech recognition text content of each vocal region.
[0106] The computer storage medium of this invention can be any combination of one or more computer-readable media. A computer-readable medium can be a computer-readable signal medium or a computer-readable storage medium. For example, a computer-readable storage medium can be, but is not limited to, an electrical, magnetic, optical, electromagnetic, infrared, or semiconductor system, apparatus, or device, or any combination thereof. More specific examples (a non-exhaustive list) of computer-readable storage media include: an electrical connection having one or more wires, a portable computer disk, a hard disk, random access memory (RAM), read-only memory (ROM), erasable programmable read-only memory (EPROM or flash memory), optical fiber, portable compact disk read-only memory (CD-ROM), optical storage device, magnetic storage device, or any suitable combination thereof. In this document, a computer-readable storage medium can be any tangible medium that contains or stores a program that can be used by or in conjunction with an instruction execution system, apparatus, or device.
[0107] Computer-readable signal media may include data signals propagated in baseband or as part of a carrier wave, carrying computer-readable program code. Such propagated data signals may take various forms, including but not limited to electromagnetic signals, optical signals, or any suitable combination thereof. Computer-readable signal media may also be any computer-readable medium other than computer-readable storage media, capable of sending, propagating, or transmitting programs for use by or in connection with an instruction execution system, apparatus, or device.
[0108] Program code contained on a computer-readable medium may be transmitted using any suitable medium, including but not limited to: wireless, wire, optical fiber, RF, etc., or any suitable combination thereof.
[0109] Computer program code for performing the operations of this invention can be written in one or more programming languages or a combination thereof. Programming languages include object-oriented programming languages such as Java, Smalltalk, and C++, as well as conventional procedural programming languages such as C or similar languages. The program code can be executed entirely on the user's computer, partially on the user's computer, as a standalone software package, partially on the user's computer and partially on a remote computer, or entirely on a remote computer or server. In cases involving remote computers, the remote computer can be connected to the user's computer via any type of network, including a local area network (LAN) or a wide area network (WAN), or it can be connected to an external computer (e.g., via the Internet using an Internet service provider).
[0110] Those skilled in the art will understand that the modules or steps of the present invention described above can be implemented using general-purpose computing devices. They can be centralized on a single computing device or distributed across a network of multiple computing devices. Optionally, they can be implemented using computer-executable program code, thereby allowing them to be stored in a storage device for execution by a computing device, or they can be fabricated as separate integrated circuit modules, or multiple modules or steps can be fabricated as a single integrated circuit module. Thus, the present invention is not limited to any particular combination of hardware and software.
[0111] Note that the above description is merely a preferred embodiment of the present invention and the technical principles employed. Those skilled in the art will understand that the present invention is not limited to the specific embodiments described herein, and various obvious changes, readjustments, and substitutions can be made without departing from the scope of protection of the present invention. Therefore, although the present invention has been described in detail through the above embodiments, the present invention is not limited to the above embodiments, and may include many other equivalent embodiments without departing from the concept of the present invention, the scope of which is determined by the scope of the appended claims.
Claims
1. A vehicle-mounted voice recognition method, characterized in that, include: Multiple raw audio signals collected by vehicle-mounted microphones in each sound zone of the target vehicle are acquired, and fast Fourier transforms are performed on the multiple raw audio signals to obtain the corresponding time-spectrum diagrams; wherein, the time-spectrum diagram is a graph of the frequency and amplitude of the multiple raw audio signals changing over time; The time-spectrum diagrams are superimposed and their average values are calculated to obtain a mixed average time-spectrum diagram, which serves as the feature of the mixed sound region. The mixed vocal region features are input into a preset vocal region encoding and recognition neural network to obtain the vocal region encoding features; wherein, the vocal region encoding features are the encoding feature data of the audio signal of each vocal region; The encoding features of each vocal region are input into a preset speech recognition network to obtain the speech recognition text content of each vocal region, and the target speech recognition result is determined based on the speech recognition text content of each vocal region. The training process of the preset vocal range encoding and recognition neural network includes: The mixed average time-frequency spectrum of multiple audio sample signals collected by the vehicle-mounted microphones in each sound zone of the target vehicle when a voice emits a voice is used as one type of sample, and the time-frequency spectrum of the audio sample signal collected by the vehicle-mounted microphone in the sound zone of the voice emitter and the corresponding number of pure noise audio time-frequency spectrums are used as one type of sample label. Based on the aforementioned sample type and its label, a basic vocal region coding recognition neural network is obtained through model training. The mixed average time-spectrum diagram of multiple audio sample signals collected by the vehicle microphones in each sound zone of the target vehicle when multiple voices emit voices is used as the second type of sample, and the time-spectrum diagram of the audio sample signals collected by the vehicle microphones in the sound zones of the multiple voices and the corresponding number of pure noise audio time-spectrum diagrams are used as the second type of sample labels. The basic vocal region encoding and recognition neural network is trained based on the two types of samples and the labels of the two types of samples to obtain the preset vocal region encoding and recognition neural network.
2. The method according to claim 1, characterized in that, The determination of the target speech recognition result based on the speech recognition text content of each vocal region includes: The recognition confidence of the speech recognition text content in each voice region is compared with the preset confidence reference threshold. The speech recognition text content of the vocal region with a confidence level greater than the preset confidence reference threshold is taken as the target speech recognition result.
3. The method according to claim 1, characterized in that, The training process of the preset speech recognition network includes: The first type of sample and the second type of sample are respectively input into the preset sound region coding recognition neural network to obtain the corresponding sound region coding features of each sample. The audio region coding features of each sample are used as training samples for the speech recognition network model, and the text content of the multiple original audio signals corresponding to the first-class sample or the second-class sample is used as sample labels to train the initial speech recognition model, thereby obtaining the preset speech recognition network.
4. The method according to claim 3, characterized in that, The preset speech recognition network is a model trained based on the Transformer-CTC model.
5. The method according to claim 1, characterized in that, The method further includes: The target speech recognition result is displayed on the vehicle's in-vehicle display screen.
6. A vehicle-mounted voice recognition device, characterized in that, include: The voice signal acquisition module is used to acquire multiple raw audio signals collected by vehicle-mounted microphones in each sound zone of the target vehicle, and to perform fast Fourier transform on the multiple raw audio signals to obtain the corresponding time-spectrum diagrams; the time-spectrum diagrams are superimposed and averaged to obtain a mixed average time-spectrum diagram, which serves as the mixed sound zone feature; wherein, the time-spectrum diagram is a graph of the frequency and amplitude of the multiple raw audio signals changing over time; The pitch region coding feature determination module is used to input the mixed pitch region features into a preset pitch region coding recognition neural network to obtain the pitch region coding features; wherein, the pitch region coding features are the coding feature data of the audio signal of each pitch region; The speech recognition result determination module is used to input the coding features of each voice region into a preset speech recognition network to obtain the speech recognition text content of each voice region, and determine the target speech recognition result based on the speech recognition text content of each voice region; The vehicle-mounted voice recognition device further includes: a preset voice region encoding recognition neural network training module, used to take the mixed average time-spectrum map of multiple audio sample signals collected by the vehicle-mounted microphones of each voice region in the target vehicle when a voice emits speech as a first-class sample, and the time-spectrum map of the audio sample signal collected by the vehicle-mounted microphone in the voice region of the voice emitter and the corresponding number of pure noise audio time-spectrum maps as a first-class sample label; to perform model training based on the first-class samples and the first-class sample labels to obtain a basic voice region encoding recognition neural network; to take the mixed average time-spectrum map of multiple audio sample signals collected by the vehicle-mounted microphones of each voice region in the target vehicle when multiple voice emitters speech as a second-class sample, and the time-spectrum map of the audio sample signal collected by the vehicle-mounted microphones in the voice regions of the multiple voice emitters and the corresponding number of pure noise audio time-spectrum maps as a second-class sample label; to perform model training based on the second-class samples and the second-class sample labels to obtain the preset voice region encoding recognition neural network.
7. A computer device, characterized in that, The computer device includes: One or more processors; Memory, used to store one or more programs; When the one or more programs are executed by the one or more processors, the one or more processors implement the in-vehicle voice recognition method as described in any one of claims 1-5.
8. A computer-readable storage medium having a computer program stored thereon, characterized in that, When the program is executed by the processor, it implements the in-vehicle voice recognition method as described in any one of claims 1-5.