System for realizing intelligent cabin communication function for hearing disorder passengers
By combining gesture recognition and voice recognition technology in the intelligent cockpit system, the communication problem between passengers and drivers with hearing impairments is solved, and accurate and convenient barrier-free communication is achieved.
Patent Information
- Application Number
- CN202510109045.0
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-01-23
- Publication Date
- 2025-05-13
AI Technical Summary
In online car-hailing scenarios, hearing-impaired passengers find it difficult to communicate effectively with the driver, resulting in insufficient communication of demand.
An intelligent cockpit system is designed, using the first domain controller to realize the gesture recognition algorithm, convert gestures into text through the convolutional neural network, and implement the voice recognition algorithm through the second domain controller to communicate through the MQTT protocol to ensure that drivers and passengers can communicate without barriers.
It realizes effective communication for passengers with hearing impairments, improves the accuracy and convenience of communication, and meets a wide range of application needs.
Smart Images

Figure CN119993155A_ABST
Abstract
Description
Technical Field
[0001] The present invention relates to the field of smart cockpits, and in particular to the application of network and gesture recognition in the field of smart cockpits, and specifically refers to a system for realizing smart cockpit communication functions for hearing-impaired passengers. Background Art
[0002] With the trend of digital transformation of automobiles, the cockpits in cars tend to be intelligent and technological. Taxis and online ride-hailing cars also respond to the trend of intelligence and are equipped with corresponding intelligent facilities: for example, they can collect and recognize the facial information of passengers or drivers, and generate customized services through preset configurations, such as seat adjustment, window settings, temperature settings, etc.; in terms of voice, you can also use custom wake-up words to wake up the robot for voice operations, such as turning on and off the air conditioner, turning on music, etc., to achieve interaction with users and satisfy the driving experience of drivers and passengers. Some online ride-hailing cars are equipped with intelligent cockpit systems such as one chip and one screen (a combination of a domain controller and a large central control screen), two chips and three screens (a domain controller controls the large central control screen, and an additional domain controller controls two rear screens) and one chip and three screens (a domain controller controls three screens), so that passengers and drivers have convenient and operable screens respectively, realizing intelligent communication and exchange.
[0003] However, in the scenario of online car-hailing, for passengers with hearing impairments, there is a problem of difficulty in communication between the driver and the passenger. If the passenger wants to express his or her corresponding emotions and needs at the time, such as being in a hurry, or hoping to close the driver's side window to remind the driver that he or she has taken the wrong road, etc., due to physiological barriers, it will be difficult for the two to reach an agreement and the needs cannot be fully conveyed. Therefore, it is necessary to invent an intelligent cockpit system that can be used by passengers with hearing impairments to solve the problem of communication between drivers and passengers in online car-hailing. Summary of the invention
[0004] The purpose of the present invention is to overcome the shortcomings of the above-mentioned prior art and provide a system for realizing intelligent cockpit communication functions for hearing-impaired passengers, which has high accuracy, simple operation and a wide range of applications.
[0005] In order to achieve the above-mentioned purpose, the system of the present invention for realizing the intelligent cockpit communication function for hearing-impaired passengers is as follows:
[0006] The system for realizing intelligent cockpit communication function for hearing-impaired passengers has the following main features: the system includes a first domain controller and a second domain controller, the first domain controller is placed in the back row and connected to two passenger screens behind the front seats, the second domain controller is deployed below the central control screen and connected to the central control screen and the instrument screen, and the first domain controller and the second domain controller communicate with each other via MQTT; the first domain controller is used to deploy convolutional neural networks and can implement gesture recognition algorithms, the second domain controller implements instrument and Android dual systems based on virtualization, and can implement speech recognition algorithms.
[0007] Preferably, the second domain controller includes a virtualization module, and also includes an Android system and an instrument system, and the Android system and the instrument system are respectively connected to the virtualization module.
[0008] Preferably, the first domain controller includes a HandCollect module, a first WordManager module, a HandToWordModel module, and an MQTTServer module. The HandCollect module is an application layer application for collecting gesture information. The first WordManager module is connected to the HandCollect module, the HandToWordModel module is connected to the first WordManager module, the input end of the MQTTServer module is connected to the first WordManager module, and the output end of the MQTTServer module is connected to the passenger screen and also connected to the second domain controller.
[0009] Preferably, the second domain controller includes an MQTTClient module, a second WordManager module, a SoundToWordModel module, and a ShowToDriver module.
[0010] The input end of the MQTTClient module is connected to the MQTTServer module of the first domain controller, the second WordManager module is connected to the MQTTClient module, the SoundToWordModel module is connected to the second WordManager module, the input end of the ShowToDriver module is connected to the MQTTClient module, and the output end of the ShowToDriver module is connected to the central control screen and the instrument screen.
[0011] Preferably, the first domain controller implements a gesture recognition algorithm through a GoogleNET convolutional neural network, and the GoogleNET convolutional neural network is composed of 5 GoogleNET single modules.
[0012] Preferably, the GoogleNET single module includes a concatenator, multiple convolution kernels, a pooling layer and a preprocessing layer, the multiple convolution kernels are connected to the concatenator, the pooling layer is connected to the convolution kernel, and the preprocessing layer is connected to the convolution kernel or the pooling layer.
[0013] Preferably, the multiple convolution kernels include 1×1 convolution kernels, 3×3 convolution kernels and 5×5 convolution kernels, the pooling layer includes a 3×3 pooling layer, one 3×3 convolution kernel, one 5×5 convolution kernel and two 1×1 convolution kernels are connected to the series concatenator, the input end of one 1×1 convolution kernel is connected to the preprocessing layer, the input end of another 1×1 convolution kernel is connected to the 3×3 pooling layer and the input end of the 3×3 pooling layer is connected to the preprocessing layer, the input ends of the 3×3 convolution kernel and the 5×5 convolution kernel are respectively connected to a 1×1 convolution kernel, and the input end of the 1×1 convolution kernel is connected to the preprocessing layer.
[0014] Preferably, the GoogleNET convolutional neural network is built using the Tensorflow framework.
[0015] Preferably, the activation function used in each layer of the GoogleNET convolutional neural network is a ReLU activation function.
[0016] Preferably, the GoogleNET convolutional neural network converts the input passenger fixed gesture into text output, and training the GoogleNET convolutional neural network includes the following steps:
[0017] (1) Collect a large amount of training data from hearing-impaired people by collecting questionnaires;
[0018] (2) The collected training data and the corresponding Chinese information form the training set of the GoogleNet convolutional neural network;
[0019] (3) Perform reinforcement learning training on the GoogleNet convolutional neural network.
[0020] The system of the present invention is used to realize the intelligent cockpit communication function for hearing-impaired passengers. The two-core four-screen cockpit domain architecture provides sufficient computing resources for face recognition, separates the computing power of the rear screen and the central control screen, and the computing power of the instrument, and separates the computing power of voice recognition and image recognition, which reduces the computing power burden of the domain controller and is more versatile. The central control and instrument are implemented based on virtualization on MTK2715, with one core and two systems, and more lightweight deployment. The present invention pioneered the link from gesture to text using the GoogleNET neural network, which has a wide range of applications. BRIEF DESCRIPTION OF THE DRAWINGS
[0021] Figure 1 This is a workflow diagram of the gesture recognition algorithm of the system for realizing the intelligent cockpit communication function for hearing-impaired passengers of the present invention.
[0022] Figure 2 This is a single-module architecture diagram of GoogleNet of the system for realizing intelligent cockpit communication function for hearing-impaired passengers of the present invention.
[0023] Figure 3 This is a diagram of the GoogleNet convolutional neural network architecture of the system of the present invention for realizing intelligent cockpit communication functions for hearing-impaired passengers.
[0024] Figure 4 This is an architectural diagram of the first domain controller and the second domain controller of the system for realizing the intelligent cockpit communication function for hearing-impaired passengers of the present invention.
[0025] Figure 5 This is a cockpit domain architecture diagram of the system of the present invention that realizes the intelligent cockpit communication function for hearing-impaired passengers.
[0026] Figure 6 This is a virtualized architecture diagram of the system for realizing intelligent cockpit communication functions for hearing-impaired passengers of the present invention. DETAILED DESCRIPTION
[0027] In order to more clearly describe the technical content of the present invention, further description is given below in conjunction with specific embodiments.
[0028] The system of the present invention for realizing intelligent cockpit communication function for hearing-impaired passengers includes a first domain controller and a second domain controller. The first domain controller is placed in the back row and connected to two passenger screens behind the front seats. The second domain controller is deployed below the central control screen and connected to the central control screen and the instrument screen. The first domain controller and the second domain controller communicate with each other via MQTT. The first domain controller is used to deploy a convolutional neural network and can realize a gesture recognition algorithm. The second domain controller realizes the instrument and Android dual systems based on virtualization and can realize a speech recognition algorithm.
[0029] As a preferred implementation manner of the present invention, the second domain controller includes a virtualization module, and also includes an Android system and an instrument system, and the Android system and the instrument system are respectively connected to the virtualization module.
[0030] As a preferred embodiment of the present invention, the first domain controller includes a HandCollect module, a first WordManager module, a HandToWordModel module, and an MQTTServer module. The HandCollect module is an application layer application for collecting gesture information. The first WordManager module is connected to the HandCollect module, the HandToWordModel module is connected to the first WordManager module, the input end of the MQTTServer module is connected to the first WordManager module, and the output end of the MQTTServer module is connected to the passenger screen and also connected to the second domain controller.
[0031] As a preferred embodiment of the present invention, the second domain controller includes an MQTTClient module, a second WordManager module, a SoundToWordModel module, and a ShowToDriver module.
[0032] The input end of the MQTTClient module is connected to the MQTTServer module of the first domain controller, the second WordManager module is connected to the MQTTClient module, the SoundToWordModel module is connected to the second WordManager module, the input end of the ShowToDriver module is connected to the MQTTClient module, and the output end of the ShowToDriver module is connected to the central control screen and the instrument screen.
[0033] As a preferred embodiment of the present invention, the first domain controller implements a gesture recognition algorithm through a GoogleNET convolutional neural network, and the GoogleNET convolutional neural network is composed of 5 GoogleNET single modules.
[0034] As a preferred embodiment of the present invention, the GoogleNET single module includes a concatenator, multiple convolution kernels, a pooling layer and a preprocessing layer, the multiple convolution kernels are connected to the concatenator, the pooling layer is connected to the convolution kernel, and the preprocessing layer is connected to the convolution kernel or the pooling layer.
[0035] As a preferred embodiment of the present invention, the multiple convolution kernels include a 1×1 convolution kernel, a 3×3 convolution kernel and a 5×5 convolution kernel, the pooling layer includes a 3×3 pooling layer, one of the 3×3 convolution kernels, one of the 5×5 convolution kernels and two of the 1×1 convolution kernels are connected to the series concatenator, the input end of one of the 1×1 convolution kernels is connected to the preprocessing layer, the input end of another 1×1 convolution kernel is connected to the 3×3 pooling layer and the input end of the 3×3 pooling layer is connected to the preprocessing layer, the input ends of the 3×3 convolution kernel and the 5×5 convolution kernel are respectively connected to a 1×1 convolution kernel, and the input end of the 1×1 convolution kernel is connected to the preprocessing layer.
[0036] As a preferred embodiment of the present invention, the GoogleNET convolutional neural network is built using the Tensorflow framework.
[0037] As a preferred embodiment of the present invention, the activation function used in each layer of the GoogleNET convolutional neural network is a ReLU activation function.
[0038] As a preferred embodiment of the present invention, the GoogleNET convolutional neural network converts the input passenger fixed gesture into text output, and training the GoogleNET convolutional neural network includes the following steps:
[0039] (1) Collect a large amount of training data from hearing-impaired people by collecting questionnaires;
[0040] (2) The collected training data and the corresponding Chinese information form the training set of the GoogleNet convolutional neural network;
[0041] (3) Perform reinforcement learning training on the GoogleNet convolutional neural network.
[0042] In a specific embodiment of the present invention, a software system suitable for a smart cockpit is designed. On the passenger screen, the sign language information and facial information of the hearing-impaired passenger can be comprehensively translated into text that the driver can understand, and projected onto the instrument screen and the central control screen. The driver then responds to the corresponding voice according to the needs, and the system translates the voice into text information and outputs it to the passenger screen. The smart cockpit system solves the problem of timely communication between the driver and the hearing-impaired passenger.
[0043] In order to achieve the adaptive effect of gestures, a convolutional neural network is constructed based on computer vision, from the input of fixed gestures of passengers to the output of text. The GoogleNet open source model is used as the basis of the algorithm module. The convolutional neural network includes input layer, convolution layer, activation layer, pooling layer and output layer. The network adds two auxiliary softmax layers for forward gradient conduction (auxiliary classifier), modularizes the neural network, encapsulates the pooling layer, convolution layer and activation layer, and uses Inception V1 to open up the neural network link from fixed gestures to text.
[0044] In order to complete the speech-to-text link, it is also necessary to add the Transformer speech model and the ChatGLM-6B bilingual dialogue language model. Both are open source models. After combining them, the speech recognition system of the intelligent cockpit domain controller can be realized.
[0045] In terms of sample data collection, for the GoogleNet convolutional neural network, a questionnaire survey was used to collect data. Ten thousand main gestures of hearing-impaired people when riding a car were collected: turn left ahead, turn right ahead, U-turn ahead, pull over ahead, go the wrong way, drive according to the navigation, go faster, go slower, open the windows (front, back, left, and right), the air conditioner is too cold (hot), in a hurry, how long will it take to arrive. The above gestures and the corresponding Chinese information form a training set, and GoogleNet is trained for reinforcement learning. For the Transformer speech model, 30 hours of open source data from Tsinghua University is used for basic training, and 10,000 voices commonly used by drivers are collected for intensive training, achieving a recognition accuracy of 95%. For the language model, no training is required and it can be used directly.
[0046] In the cockpit domain of the automotive electronic and electrical architecture, two MTK2715 domain controllers and two screens are used to implement a two-core four-screen design architecture: an MTK2715 domain controller (hereinafter referred to as domain controller A) is placed in the back row, responsible for deploying convolutional neural networks and linking the two screens behind the front seats. The two screens serve as entertainment screens for rear passengers. In addition to providing conventional entertainment functions and in-vehicle functions, they are also responsible for collecting gesture image information of passengers on the bus, and transmitting it to the front MTK2715 (hereinafter referred to as domain controller B) as algorithm input through the internal local area network via MQTT; domain controller B is deployed below the central control screen, and implements the instrument and Android dual systems based on virtualization, and deploys a voice recognition algorithm. The two domain controllers communicate with each other through MQTT to achieve communication between domain controllers.
[0047] The present invention aims to build a complete intelligent cockpit domain, which consists of three parts: the training of gesture recognition algorithm, the integration of Android platform based on virtualization, and the design of the overall architecture of the cockpit domain.
[0048] For the gesture recognition algorithm, we first need to collect 10,000 training data. The data is collected by designing a questionnaire. We found 10,000 people with hearing impairments and collected their gestures for the following instructions: turn left ahead, turn right ahead, U-turn ahead, pull over ahead, go the wrong way, follow the navigation, go faster, go slower, open the window (front, back, left, right), the air conditioner is too cold (hot), and in a hurry, and how long it will take to reach these instructions. We also took pictures of the gestures to form the training set of GoogleNet.
[0049] The GoogleNET convolutional neural network is built using the Tensorflow framework. This neural network combines and modularizes the pooling layer and the convolution layer, and uses a 1×1 convolution kernel for network optimization, which reduces the input parameters and reduces the overfitting of the algorithm. Figure 2 shown.
[0050] Figure 2 yes Figure 3 The single module structure Figure 3 It is composed of five single modules in sequence. A single module is divided into two layers. The first layer is two 1×1 convolution kernels, a 3×3 convolution kernel, and a 5×5 convolution kernel. After the parameters are input, they are calculated in this layer and then go to the next layer, which is composed of two 1×1 convolution kernels and a pooling layer. The connection relationship is the calculation.
[0051] The overall neural network structure consists of 5 single-module structures. After the image data is input, the first single-module consists of a 7×7 convolution layer with a stride of 2 and a 3×3 pooling layer with a stride of 2. After calculation, it is normalized (LocalRespNorm) and input to the second block (64 1×1 convolution layers and a 3×3 convolution layer with a stride of 1). After normalization, it enters the maximum pooling layer (3×3 with a stride of 2). The third module consists of two parts, each with four branches. For the first part: the first branch uses 64 1×1 convolution kernels for operation; the second branch uses 96 1×1 convolution kernels for operation, and then performs 128 3×3 convolutions; the third branch uses 16 1×1 convolution kernels, and then performs 32 5×5 convolution layers; the pooling layer uses a 3×3 kernel, and then performs 32 1×1 convolutions. After the four branches are completed, the output results of the four branches are connected in parallel and input to the second part of the third module. For the second part, four branches are also used, and the structure is similar to the first part, but different convolution scales are used: the first branch is 128 1×1 convolution kernels; the second branch is 128 1×1 convolution kernels, and then 192 3×3 convolutions; the third branch is 32 1×1 convolution kernels, and then 96 5×5 convolutions; the fourth branch is a pooling layer, using a 3×3 kernel, and then 64 1×1 convolutions. Connect the four results and input them to the fourth module. The structure of the fourth and fifth modules is similar to the third module, only the number of convolution kernels is different. The number of convolution kernels of all modules is now sorted out as shown in the following table:
[0052]
[0053]
[0054] The activation function used in each layer is the ReLU activation function.
[0055] Since the middle layer of the neural network also has strong recognition capabilities, the features of the two layers are extracted in the fourth module and output to the model results. After the training is completed, the three models are fused to avoid overfitting and improve the accuracy of the model. In terms of deployment, the model is deployed to the cockpit domain controller A for use by applications on the domain controller.
[0056] Module software architecture such as Figure 4 As shown. On domain controller A, HandCollect is an application layer application responsible for collecting gesture information, and after the gesture information is calculated by the convolutional neural network, the predicted text results are transmitted to domain controller B through MQTT communication. In domain controller B, the results are fed back to the central control screen and instrument panel through the ShowToDriver module.
[0057] Figure 4 In domain controller A, HandCollect is an application layer application that recognizes gestures and inputs them into wordManager. WordManager is a system application that links the HandtoWordModel model preset in the machine, that is, Figure 3 The recognition model calculated by GoogleNet is obtained. The gesture recognition information obtained by the application layer is outputted into the corresponding text result through the large model, and the text result is transmitted to the domain controller B through MQTT. The voice model is preset in the domain controller B. The text result will be displayed to the application layer application ShowToDriver module of the domain controller B and presented to the driver. After the driver reads the text information, he can input the voice into the SoundToWordModel voice model through voice. The voice model converts the voice recognition into text, which is then transmitted to the domain controller A through MQTT and presented to the passengers through the screen of A, realizing barrier-free communication.
[0058] The system architecture of the cockpit domain is as follows Figure 5 ,The two domain controllers use Ethernet and MQTT communication methods to complete data exchange using long ,connections and multiple connections.
[0059] Due to the use of a two-core four-screen system architecture, this system has the characteristics of high stability and high response speed, which reduces the computational burden brought by the algorithm and can achieve long-term stable operation.
[0060] Figure 6 It demonstrated a dual-system architecture based on virtualization, with both Linux and Android systems on the domain controller. This architecture was applied for the first time to applications for the hearing impaired, achieving lightweight deployment.
[0061] The specific implementation scheme of this embodiment can refer to the relevant description in the above embodiment, which will not be repeated here.
[0062] It can be understood that the same or similar parts of the above embodiments can be referenced to each other, and the contents not described in detail in some embodiments can refer to the same or similar contents in other embodiments.
[0063] It should be noted that, in the description of the present invention, the terms "first", "second", etc. are only used for descriptive purposes and cannot be understood as indicating or implying relative importance. In addition, in the description of the present invention, unless otherwise specified, the meaning of "plurality" refers to at least two.
[0064] In the description of this specification, the description with reference to the terms "one embodiment", "some embodiments", "examples", "specific examples", or "some examples" means that the specific features, structures, materials or characteristics described in conjunction with the embodiment or example are included in at least one embodiment or example of the present invention. In this specification, the schematic representation of the above terms does not necessarily refer to the same embodiment or example. Moreover, the specific features, structures, materials or characteristics described may be combined in any one or more embodiments or examples in a suitable manner.
[0065] The system of the present invention is used to realize the intelligent cockpit communication function for hearing-impaired passengers. The two-core four-screen cockpit domain architecture provides sufficient computing resources for face recognition, separates the computing power of the rear screen and the central control screen, and the computing power of the instrument, and separates the computing power of voice recognition and image recognition, which reduces the computing power burden of the domain controller and is more versatile. The central control and instrument are implemented based on virtualization on MTK2715, with one core and two systems, and more lightweight deployment. The present invention pioneered the link from gesture to text using the GoogleNET neural network, which has a wide range of applications.
[0066] In this specification, the present invention has been described with reference to specific embodiments thereof. However, it is apparent that various modifications and variations may be made without departing from the spirit and scope of the present invention. Therefore, the specification and drawings should be regarded as illustrative rather than restrictive.
Claims
1. A system for realizing intelligent cockpit communication function for hearing-impaired passengers, characterized in that: The system includes a first domain controller and a second domain controller. The first domain controller is placed in the back row and connected to the two passenger screens behind the front seats. The second domain controller is deployed below the central control screen and connected to the central control screen and the instrument screen. The first domain controller and the second domain controller communicate with each other via MQTT. The first domain controller is used to deploy a convolutional neural network and can implement a gesture recognition algorithm. The second domain controller implements the instrument and Android dual systems based on virtualization and can implement a voice recognition algorithm.
2. The system for realizing intelligent cockpit communication function for hearing-impaired passengers according to claim 1, characterized in that: The second domain controller includes a virtualization module, and also includes an Android system and an instrument system, and the Android system and the instrument system are respectively connected to the virtualization module.
3. The system for realizing intelligent cockpit communication function for hearing-impaired passengers according to claim 1, characterized in that: The first domain controller includes a HandCollect module, a first WordManager module, a HandToWordModel module, and an MQTTServer module. The HandCollect module is an application layer application for collecting gesture information. The first WordManager module is connected to the HandCollect module, the HandToWordModel module is connected to the first WordManager module, the input end of the MQTTServer module is connected to the first WordManager module, and the output end of the MQTTServer module is connected to the passenger screen and also connected to the second domain controller.
4. The system for realizing intelligent cockpit communication function for hearing-impaired passengers according to claim 3, characterized in that: The second domain controller includes an MQTTClient module, a second WordManager module, a SoundToWordModel module, and a ShowToDriver module. The input end of the MQTTClient module is connected to the MQTTServer module of the first domain controller, the second WordManager module is connected to the MQTTClient module, the SoundToWordModel module is connected to the second WordManager module, the input end of the ShowToDriver module is connected to the MQTTClient module, and the output end of the ShowToDriver module is connected to the central control screen and the instrument screen.
5. The system for realizing intelligent cockpit communication function for hearing-impaired passengers according to claim 1, characterized in that: The first domain controller implements the gesture recognition algorithm through the GoogleNET convolutional neural network, and the GoogleNET convolutional neural network is composed of 5 GoogleNET single modules.
6. The system for realizing intelligent cockpit communication function for hearing-impaired passengers according to claim 5, characterized in that: The GoogleNET single module includes a concatenator, multiple convolution kernels, a pooling layer and a preprocessing layer, wherein the multiple convolution kernels are connected to the concatenator, the pooling layer is connected to the convolution kernel, and the preprocessing layer is connected to the convolution kernel or the pooling layer.
7. The system for realizing intelligent cockpit communication function for hearing-impaired passengers according to claim 6, characterized in that: The multiple convolution kernels include 1×1 convolution kernels, 3×3 convolution kernels and 5×5 convolution kernels, the pooling layer includes a 3×3 pooling layer, one 3×3 convolution kernel, one 5×5 convolution kernel and two 1×1 convolution kernels are connected to the series concatenator, the input end of one 1×1 convolution kernel is connected to the preprocessing layer, the input end of another 1×1 convolution kernel is connected to the 3×3 pooling layer and the input end of the 3×3 pooling layer is connected to the preprocessing layer, the input ends of the 3×3 convolution kernel and the 5×5 convolution kernel are respectively connected to a 1×1 convolution kernel, and the input end of the 1×1 convolution kernel is connected to the preprocessing layer.
8. The system for realizing intelligent cockpit communication function for hearing-impaired passengers according to claim 5, characterized in that: The GoogleNET convolutional neural network is built using the Tensorflow framework.
9. The system for realizing intelligent cockpit communication function for hearing-impaired passengers according to claim 5, characterized in that: The activation function used in each layer of the GoogleNET convolutional neural network is the ReLU activation function.
10. The system for realizing intelligent cockpit communication function for hearing-impaired passengers according to claim 5, characterized in that: The GoogleNET convolutional neural network converts the input passenger fixed gesture into text output. Training the GoogleNET convolutional neural network includes the following steps: (1) Collect a large amount of training data from hearing-impaired people by collecting questionnaires; (2) The collected training data and the corresponding Chinese information form the training set of the GoogleNet convolutional neural network; (3) Perform reinforcement learning training on the GoogleNet convolutional neural network.