Voice recognition method and device and motor train unit cab voice interaction system

By adopting offline voice recognition methods in the EMU, combined with voice wake-up and broadcasting processes, the implementation problems and network security requirements of voice recognition in the EMU are solved, and the convenience and security of voice interaction are achieved.

CN120015017APending Publication Date: 2025-05-16CRRC DALIAN R & D CO LTD
View PDF 0 Cites 0 Cited by

Patent Information

Application Number
CN202510036838.4
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-01-09
Publication Date
2025-05-16

AI Technical Summary

Technical Problem

The prior art cannot realize voice recognition in EMUs, and online voice recognition products cannot meet the network security requirements of EMUs. Offline version of voice recognition products requires functional customization.

Method used

An offline speech recognition method is adopted to obtain vocal signals, convert them into text strings, and match them with voice command words, and use editing distance algorithms and deep neural networks to identify them, combining voice wake-up and broadcasting processes to realize voice interaction functions.

Benefits of technology

It realizes voice recognition in the EMU without an Internet connection, meets the network security requirements of the EMU, and improves the driver's operational convenience and driving safety through voice interaction.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120015017A_ABST
    Figure CN120015017A_ABST
Patent Text Reader

Abstract

The invention discloses a voice recognition method. The method comprises the following steps: acquiring a human voice signal to be recognized; converting the human voice signal to be recognized into a text character string; the method comprises the following steps of: matching a text character string with a voice command word, calculating the text character string and an entry affiliated set in a command word data dictionary one by one by adopting an editing distance algorithm based on a command word dictionary rule to obtain a similarity list, and comparing a maximum value in the similarity list with a set matching threshold value, the voice interaction system for the driver's cab of the motor train unit comprises a sound pickup used for collecting the sound of the driver in the driver's cab of the motor train unit, a voice recognition device used for recognizing the collected sound of the driver in the driver's cab of the motor train unit, and a loudspeaker used for receiving the sound of the driver in the driver's cab of the motor train unit. And the broadcasting device is used for broadcasting the voice command recognized by the voice recognition device.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The invention belongs to the field of fully automated products and relates to a speech recognition method, a device and a speech interaction system in a motor vehicle driver's cab. Background Art

[0002] With the continuous updating and iteration of voice interaction technology, the application of voice interaction products in the field of rail transit is gradually increasing, and voice interaction technology brings a series of conveniences and improvements to locomotive drivers. Drivers can control various auxiliary functions of the vehicle that do not affect driving through voice commands, such as air conditioning mode design, start or stop of lights, wipers and other equipment, etc., which improves the convenience of operation and reduces the workload; voice broadcasts can remind drivers of key driving information, signal changes or potential safety hazards in real time, enhancing driving safety; in an emergency, drivers can quickly send out help signals or trigger emergency measures through voice to improve emergency response speed; through voice interaction technology, drivers can reduce the operation of physical control panels, reduce interference and distraction caused by manual operation, and focus more on driving and monitoring driving safety; in addition, voice interaction can provide real-time operation guidance and suggestions during driver training and use, providing convenience for drivers to deal with emergencies.

[0003] According to the requirements of intelligent design of the driver's cab of Fuxing EMU, the new generation of centralized power EMU requires the driver's cab to have voice interaction function. In response to the requirements, a voice interaction device suitable for EMU is designed. The interactive device is guided by command control and information broadcasting, and some interactive items are extracted from the driver's cab console, microcomputer display screen, and safety monitoring system display screen. It is realized through voice interaction, which optimizes the driver's driving experience and improves the intelligence and integration of EMU.

[0004] Technical solution of the prior art: Currently common voice recognition products are all implemented online. After the sound pickup device collects the audio information, it will convert it into digital information and transmit it to the remote server via the Internet. The server performs calculations and feeds the results back to the voice recognition product for display.

[0005] Disadvantages of the existing technology: Currently, there is no device for voice recognition installed in the EMU models in operation;

[0006] Existing online voice recognition products cannot meet the network security requirements of EMUs, and offline voice recognition products need to be customized according to the actual situation of the EMUs. Summary of the invention

[0007] In order to solve the above problem, the technical solution adopted by the present invention is: a speech recognition method, comprising the following steps:

[0008] Acquire a human voice signal to be recognized;

[0009] Convert the human voice signal to be recognized into a text string;

[0010] The text string is matched with the voice command word. Based on the command word dictionary rules, the edit distance algorithm is used to calculate the text string and the entry attachment set in the command word data dictionary one by one to obtain a similarity list.

[0011] The maximum value in the similarity list is then compared with the set matching threshold. When the comparison result is greater than the matching threshold, it is determined that valid action instruction information is recognized.

[0012] Furthermore: the process of converting the human voice signal to be recognized into a text string is: using a speech recognition decoder, based on a customized offline acoustic model and language model, obtained through deep neural network calculation.

[0013] Furthermore: it also includes a voice wake-up process before the voice recognition method, and the voice wake-up process is as follows: after the wake-up engine determines that the audio input successfully matches the resource file wake-up word, the voice recognition engine is activated.

[0014] Furthermore: it also includes a voice broadcast process after the voice recognition process, and the voice broadcast process converts the recognized effective action instruction information into text through a voice synthesis engine relying on a linear prediction coding model.

[0015] A speech recognition device, comprising:

[0016] Voice wake-up module: used to activate the voice recognition engine after the wake-up engine determines that the audio input matches the wake-up word in the resource file successfully;

[0017] Speech recognition module: used to obtain the human voice signal to be recognized;

[0018] Convert the human voice signal to be recognized into a text string;

[0019] The text string is matched with the voice command word. Based on the command word dictionary rules, the edit distance algorithm is used to calculate the text string and the entry attachment set in the command word data dictionary one by one to obtain a similarity list.

[0020] Then compare the maximum value in the similarity list with the set matching threshold. When the comparison result is greater than the matching threshold, it is determined that valid action instruction information is recognized;

[0021] The voice broadcast module is used to convert the recognized effective action command information into sound through the speech synthesis engine relying on the linear prediction coding model.

[0022] An interactive method for a speech recognition method in a motor vehicle driver's cab comprises the following steps:

[0023] Get the ID of the recognized valid action command information;

[0024] Determine the category attribute carried by the ID of the valid action instruction information;

[0025] When the category attribute is a control instruction, the instruction content in the ID is sent to the EMU network control bus in the form of trdp protocol, and the EMU actuator acts according to the instruction; when the category attribute is interface interaction, the target interface information in the ID is transmitted to the interface jump module, and then the driver jumps to the corresponding interface;

[0026] When the category attribute is a status query, the query item points contained in the ID are transmitted to the data layer, the query content is found from the data layer after polling, and the result is fed back to the interface in voice or text form.

[0027] A voice interaction device for a motor vehicle driver's cab, comprising:

[0028] Acquisition module: used to obtain the ID of the recognized valid action instruction information;

[0029] Judgment module: used to judge the category attribute carried by the ID of valid action instruction information;

[0030] When the category attribute is a control instruction, the instruction content in the ID is sent to the EMU network control bus in the form of trdp protocol, and the EMU actuator acts according to the instruction; when the category attribute is interface interaction, the target interface ID is transmitted to the interface jump module, and then the driver jumps to the corresponding interface;

[0031] When the category attribute is a status query, the query ID is transmitted to the data layer, and after polling, the query result is fed back to the interface in the form of voice or text.

[0032] A voice interaction system for a motor vehicle driver's cab, comprising:

[0033] Pickup: used to collect the voice of the driver in the EMU cab, and is arranged in front of the driver's seat in the EMU cab;

[0034] Voice recognition device: used to recognize the voice of the driver in the EMU driver's cab collected, and judge the type of the recognized voice command, and feed back the judged voice command to the corresponding interactive interface,

[0035] A loudspeaker is used to announce the voice commands recognized by the voice recognition device.

[0036] Furthermore, the microphone array is provided with a spacing of 20 cm.

[0037] The present invention provides a voice interaction device for the EMU driver's cab, which improves the intelligence level of the EMU driver's cab; reduces the complexity of driver operations, reduces the need for physical operations, improves driving concentration, provides a way of emergency response in emergency situations, and provides voice operation guidance and suggestions, which makes it convenient for the driver to deal with various situations and greatly improves the driver's driving experience.

[0038] The voice device is built into the intelligent integrated display screen, saving equipment and wiring costs, and saving software development costs such as protocol analysis and data forwarding. BRIEF DESCRIPTION OF THE DRAWINGS

[0039] In order to more clearly illustrate the embodiments of the present invention or the technical solutions in the prior art, the drawings required for use in the embodiments or the description of the prior art will be briefly introduced below. Obviously, the drawings described below are some embodiments of the present invention. For ordinary technicians in this field, other drawings can be obtained based on these drawings without paying creative labor.

[0040] Figure 1 is a flow chart of a speech recognition method;

[0041] Figure 2 It is a flow chart of voice broadcast;

[0042] Figure 3 It is a voice interaction system in the driver's cab of an EMU;

[0043] Figure 4 This is a diagram of the voice interaction process in the driver's cab of an EMU in Example 1. DETAILED DESCRIPTION

[0044] It should be noted that, in the absence of conflict, the embodiments of the present invention and the features in the embodiments may be combined with each other. The present invention will be described in detail below with reference to the accompanying drawings and in combination with the embodiments.

[0045] In order to make the purpose, technical solutions and advantages of the embodiments of the present invention clearer, the technical solutions in the embodiments of the present invention will be clearly and completely described below in conjunction with the drawings in the embodiments of the present invention. Obviously, the described embodiments are only part of the embodiments of the present invention, not all of the embodiments. The following description of at least one exemplary embodiment is actually only illustrative and is by no means intended to limit the present invention and its application or use. Based on the embodiments of the present invention, all other embodiments obtained by ordinary technicians in this field without creative work are within the scope of protection of the present invention.

[0046] Figure 1is a flow chart of a speech recognition method;

[0047] A speech recognition method comprises the following steps:

[0048] S1: Obtain a human voice signal to be recognized;

[0049] S2: Convert the human voice signal to be recognized into a text string;

[0050] The text string is matched with the voice command word. Based on the command word dictionary rules, the Levenshtein distance algorithm is used to calculate the text string and the entry attachment set in the command word data dictionary one by one to obtain a similarity list. The similarity list is an intermediate value in the matching algorithm. The ultimate value is the maximum value in the list.

[0051] S3: The maximum value in the similarity list is then compared with the set matching threshold. When the comparison result is greater than the matching threshold, it is determined that valid action instruction information is recognized.

[0052] The steps S1 / S2 / S3 are performed sequentially;

[0053] Each element in the command word dictionary is assigned a unique ID (carrying attributes such as category, text, action, and remarks), and considering the needs of fuzzy speech recognition, an attached generalization set is added to each command word;

[0054] The matching threshold range is 0-100; considering the accuracy requirements of the EMU speech recognition usage scenario, a higher matching threshold is designed. The matching threshold used in this application is 95;

[0055] The process of converting the human voice signal to be recognized into a text string is as follows: the MFCC-based speech recognition engine extracts features from the speech input in units of words and sentences, and sends the extracted results to the speech recognition decoder, which is used to obtain the text string through deep neural network (DNN) calculation based on a customized offline acoustic model and language model.

[0056] Figure 2 It is a flow chart of voice broadcast;

[0057] The method also includes setting a voice wake-up process before the voice recognition method, and the voice wake-up process is as follows: after the wake-up engine determines that the audio input successfully matches the resource file wake-up word, the voice recognition engine is activated.

[0058] The method also includes a voice broadcast process after the voice recognition process, wherein the voice broadcast process converts the recognized effective action instruction information into sound through a voice synthesis engine relying on a linear predictive coding model (PLC).

[0059] A speech recognition device, comprising:

[0060] Voice wake-up module: used to activate the voice recognition engine after the wake-up engine determines that the audio input matches the wake-up word in the resource file successfully; after power-on, initialize the wake-up recognition engine, receive the audio data collected and converted by the microphone in real time, and start voice wake-up; the operation object of the recognition engine is the voice input;

[0061] Speech recognition module: used to obtain the human voice signal to be recognized;

[0062] Convert the human voice signal to be recognized into a text string;

[0063] The text string is matched with the voice command word. Based on the command word dictionary rules, the edit distance algorithm is used to calculate the text string and the entry attachment set in the command word data dictionary one by one to obtain a similarity list.

[0064] Then compare the maximum value in the similarity list with the set matching threshold. When the comparison result is greater than the matching threshold, it is determined that valid action instruction information is recognized;

[0065] The voice broadcast module is used to convert the recognized effective action command information into sound through the speech synthesis engine relying on the linear prediction coding model.

[0066] The application scenarios of voice recognition in EMUs can be divided into three categories: issuing control commands, interface interaction, and status query.

[0067] An interactive method for a speech recognition method in a motor vehicle driver's cab comprises the following steps:

[0068] Get the ID of the recognized valid action command information;

[0069] Determine the category attribute carried by the ID of the valid action instruction information;

[0070] When the category attribute is a control instruction, the instruction content in the ID is sent to the EMU network control bus in the form of trdp protocol, and the EMU actuator acts according to the instruction; when the category attribute is interface interaction, the target interface information in the ID is transmitted to the interface jump module, and then the driver jumps to the corresponding interface;

[0071] When the category attribute is a status query, the query item points contained in the ID are transmitted to the data layer, the query content is found from the data layer after polling, and the result is fed back to the interface in voice or text form.

[0072] A voice interaction device for a motor vehicle driver's cab, comprising:

[0073] Acquisition module: used to obtain the ID of the recognized valid action instruction information;

[0074] Judgment module: used to judge the category attribute carried by the ID of valid action instruction information;

[0075] When the category attribute is a control instruction, the instruction content in the ID is sent to the EMU network control bus in the form of trdp protocol, and the EMU actuator acts according to the instruction; when the category attribute is interface interaction, the target interface ID is transmitted to the interface jump module, and then the driver jumps to the corresponding interface;

[0076] When the category attribute is a status query, the query ID is transmitted to the data layer, and after polling, the query result is fed back to the interface in the form of voice or text.

[0077] Figure 3 It is a voice interaction system in the driver's cab of an EMU;

[0078] A voice interaction system for a motor vehicle driver's cab, comprising:

[0079] Pickup: used to collect the voice of the driver in the EMU cab, and is arranged in front of the driver's seat in the EMU cab to facilitate the pickup of the driver;

[0080] Voice recognition device: used to recognize the voice of the driver in the EMU driver's cab collected, and judge the type of the recognized voice command, and feed back the judged voice command to the corresponding interactive interface,

[0081] The voice recognition device is integrated into the intelligent integrated display screen in the EMU driver's cab, connected to the microphone through a USB interface, performs logical processing on the voice information collected and converted by the microphone, and outputs the results to the application software (intelligent integrated display screen interface software) and the speaker;

[0082] The speaker is used to broadcast the voice commands recognized by the voice recognition device. The speaker is externally connected to the side of the driver's cab and connected to the voice module through a DB9 interface.

[0083] The microphone is used to collect sound elements in the working scene, filter and reduce noise source information, extract key human voices, convert human voice signals into electrical signals, and transmit them to the voice module through the USB interface. The microphone array is used to collect 6 channels of original audio to ensure the range and quality of the sound. The sampling depth of the microphone is 16bit, the sampling rate is 16KHz, the audio format is PCM, and the signal-to-noise ratio is greater than 20dB, which meets the index requirements of industry specifications.

[0084] Figure 4 This is a diagram of the voice interaction process in the driver's cab of an EMU in Example 1.

[0085] Embodiment 1: The functional requirements and performance indicators of the voice interaction device in the EMU driver's cab are as follows.

[0086] Voice recognition wake-up word: "Hello, Fuxing"

[0087] Wake-up recognition feedback announcement: "Hello, Fuxing is here"

[0088] Sleeping announcement: "Fuxing train takes a break, available at any time"

[0089] Wake-up hold time: 15s. If a command word is recognized during the hold time, the timer will be reset to 15s, and the sleep announcement will be made after 15s.

[0090] The recognition rate of voice recognition command words is not less than 90% (ambient noise is less than 72dB);

[0091] The average response time of speech recognition is no more than 500ms;

[0092] The command word capacity is not less than 500;

[0093] Voice wake-up time is less than 1s;

[0094] Voice wake-up rate: average > 95%.

[0095] The voice interaction process in the EMU driver's cab is shown in the figure above. The user says "Hello, Fuxing" to the microphone. After the voice module recognizes the wake-up word, it broadcasts the wake-up feedback "Fuxing is here". At this time, the voice module enters the monitoring state for 15 seconds. During the monitoring period, all voices will be collected and analyzed. If the analysis result is a keyword instruction set or a generalized word of the keyword instruction set, the voice module will pass the command signal mapped by the keyword command to the train main control device to realize the voice control function.

[0096] Finally, it should be noted that the above embodiments are only used to illustrate the technical solutions of the present invention, rather than to limit it. Although the present invention has been described in detail with reference to the aforementioned embodiments, those skilled in the art should understand that they can still modify the technical solutions described in the aforementioned embodiments, or replace some or all of the technical features therein with equivalents. However, these modifications or replacements do not cause the essence of the corresponding technical solutions to deviate from the scope of the technical solutions of the embodiments of the present invention.

Claims

1. A speech recognition method, characterized in that: The following steps are involved: Acquire a human voice signal to be recognized; Convert the human voice signal to be recognized into a text string; The text string is matched with the voice command word. Based on the command word dictionary rules, the edit distance algorithm is used to calculate the text string and the entry attachment set in the command word data dictionary one by one to obtain a similarity list. The maximum value in the similarity list is then compared with the set matching threshold. When the comparison result is greater than the matching threshold, it is determined that valid action instruction information is recognized.

2. A speech recognition method according to claim 1, characterized in that: The process of converting the human voice signal to be recognized into a text string is: using a speech recognition decoder, based on a customized offline acoustic model and language model, and obtained through deep neural network calculation.

3. A speech recognition method according to claim 1, characterized in that: It also includes a voice wake-up process set before the voice recognition method, and the voice wake-up process is as follows: after the wake-up engine determines that the audio input successfully matches the resource file wake-up word, the voice recognition engine is activated.

4. A speech recognition method according to claim 1, characterized in that: It also includes a voice broadcast process after the voice recognition process, in which the voice broadcast process converts the recognized effective action instruction information into text through a voice synthesis engine relying on a linear prediction coding model.

5. A speech recognition device, characterized in that: include: Voice wake-up module: used to activate the voice recognition engine after the wake-up engine determines that the audio input matches the wake-up word in the resource file successfully; Speech recognition module: used to obtain the human voice signal to be recognized; Convert the human voice signal to be recognized into a text string; The text string is matched with the voice command word. Based on the command word dictionary rules, the edit distance algorithm is used to calculate the text string and the entry attachment set in the command word data dictionary one by one to obtain a similarity list. Then compare the maximum value in the similarity list with the set matching threshold. When the comparison result is greater than the matching threshold, it is determined that valid action instruction information is recognized; The voice broadcast module is used to convert the recognized effective action command information into sound through the speech synthesis engine relying on the linear prediction coding model.

6. The interactive method for the method for speech recognition in the driver's cab of a train set according to any one of claims 1 to 4, comprising the following steps: Get the ID of the recognized valid action command information; Determine the category attribute carried by the ID of the valid action instruction information; When the category attribute is a control command, the command content in the ID is sent to the EMU network control bus in the form of trdp protocol, and the EMU actuator acts according to the command; When the category attribute is interface interaction, the target interface information in the ID is transmitted to the interface jump module, and then the driver jumps to the corresponding interface; When the category attribute is a status query, the query item points contained in the ID are transmitted to the data layer, the query content is found from the data layer after polling, and the result is fed back to the interface in voice or text form.

7. A voice interaction device for a motor vehicle driver's cab, characterized in that: include: Acquisition module: used to obtain the ID of the recognized valid action instruction information; Judgment module: used to judge the category attribute carried by the ID of valid action instruction information; When the category attribute is a control command, the command content in the ID is sent to the EMU network control bus in the form of trdp protocol, and the EMU actuator acts according to the command; When the category attribute is interface interaction, the target interface ID is transmitted to the interface jump module, and then the driver jumps to the corresponding interface; When the category attribute is a status query, the query ID is transmitted to the data layer, and after polling, the query result is fed back to the interface in the form of voice or text.

8. A voice interaction system for a motor vehicle driver's cab, characterized in that: include: Pickup: used to collect the voice of the driver in the EMU cab, and is arranged in front of the driver's seat in the EMU cab; Voice recognition device: used to recognize the voice of the driver in the EMU driver's cab collected, and judge the type of the recognized voice command, and feed back the judged voice command to the corresponding interactive interface, A loudspeaker is used to announce the voice commands recognized by the voice recognition device.

9. The EMU driver's cab voice interaction system according to claim 1, characterized in that: The pickup adopts a microphone array with a spacing of 20 cm.