Multi-mode interaction control system and method for amphibious vehicle

By combining sensors such as gesture sensors, waterproof microphones and touch screens with a multimodal interactive control system onboard controllers, the control complexity and waterproofness issues of amphibious vehicles in different environments are solved, achieving intelligent seamless switching and efficient control.

CN120802720APending Publication Date: 2025-10-17WUHU SHIPYARD CO LTD +1
View PDF 0 Cites 0 Cited by

Patent Information

Application Number
CN202510879951.9
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-06-27
Publication Date
2025-10-17

AI Technical Summary

Technical Problem

The control methods of traditional amphibious vehicles in land and water scenarios have problems such as complex operation, insufficient waterproofness, and low intelligence, which leads to driver distraction, slow response and inconvenient control.

Method used

Gesture sensors, waterproof microphones, touch screens and environmental perception sensors are combined with the on-board controller to achieve multimodal interactive control. Control instructions are generated through recognition and decision-making, and the optimal interaction method is selected through weighted decision fusion. It is integrated into the steering wheel to improve waterproofness and anti-interference capabilities.

Benefits of technology

It achieves seamless switching and efficient control between land and water scenes, improves driving safety and convenience, and reduces misoperation and slow response.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120802720A_ABST
    Figure CN120802720A_ABST
Patent Text Reader

Abstract

The invention discloses a multi-mode interaction control system and method of an amphibious vehicle, and belongs to the field of amphibious vehicles. The system comprises a gesture sensor, a waterproof microphone, a touch screen, an environment sensing sensor, a vehicle-mounted controller and an execution mechanism, wherein the gesture sensor is used for collecting a user gesture signal; the waterproof microphone is used for collecting a user voice signal; the touch screen is used for collecting a user touch signal; the environment sensing sensor is used for acquiring the current environment state of the vehicle; the vehicle-mounted controller is used for recognizing the gesture signal, the voice signal and the touch signal of the user according to the current environment state of the vehicle and deciding to generate a control instruction; the execution mechanism is used for receiving and executing the control instruction. The invention provides a multi-mode interaction control system with waterproof performance, anti-interference capability and intelligent decision making, seamless switching and efficient control of land and water scenes are realized, and driving safety and convenience are improved.
Need to check novelty before this filing date? Find Prior Art

Description

TECHNICAL FIELD

[0001] The present application belongs to the field of amphibious vehicles, and particularly relates to a multimodal interactive control system and method for amphibious vehicles. BACKGROUND

[0002] With the increasing application demand of amphibious vehicles in extreme environments, the traditional control mode exposes the following problems:

[0003] 1. Land scene: the driver needs to operate the steering wheel and physical buttons with both hands, which is easy to be mispressed or tired in rough off-road driving, and complex functions (such as four-wheel drive mode switching and wading depth monitoring) need to be operated through multiple menus, which disperses the driving attention.

[0004] 2. Water scene: the steering wheel and buttons are easy to be damaged by water in a wet environment, and the driver needs to control the direction with one hand and operate the propeller with the other hand when driving on water, which is complicated and slow to respond in an emergency. The waterproof performance of physical buttons is insufficient: the traditional sealing rubber strip is easy to age or install unevenly, causing water penetration and damage to the circuit

[0005] 3. Cross-environment switching: the water-land mode switching needs to be manually operated by a mechanical button, and lacks intelligent linkage. The existing technology is mostly single mode (such as pure voice or pure gesture), lacks dynamic fusion strategy, and cannot automatically adapt the interaction mode according to the environment. The voice interaction has weak anti-interference ability: engine noise and water flow sound can easily cause misrecognition of voice commands; the gesture recognition has poor adaptability: underwater light refraction, driver's gloves blocking or limited action amplitude can cause the traditional visual recognition algorithm to fail.

[0006] Therefore, the present application provides a multimodal interactive control system and method for amphibious vehicles. SUMMARY

[0007] The present application aims to overcome the shortcomings of the prior art and provides a multimodal interactive control system and method for amphibious vehicles to achieve the following purposes: to provide a multimodal interactive control system with waterproof performance, anti-interference ability and intelligent decision-making, to realize seamless switching and efficient control between land and water scenes, and to improve driving safety and convenience.

[0008] To achieve the above purposes, the technical solution adopted by the present application is as follows: a multimodal interactive control system for amphibious vehicles, the system comprising a gesture sensor, a waterproof microphone, a touch screen, an environment perception sensor, a vehicle-mounted controller and an execution mechanism, wherein:

[0009] The gesture sensor is used to collect user gesture signals and send them to the vehicle-mounted controller;

[0010] The waterproof microphone is used to collect user voice signals and send them to the vehicle-mounted controller;

[0011] The touch screen is used for collecting user touch signals and sending to the vehicle-mounted controller;

[0012] The environment perception sensor is used for obtaining the current environment state of the vehicle, including land, water, and transition state;

[0013] The vehicle-mounted controller is used for identifying the user gesture signal, voice signal, and touch signal and generating control instructions according to the current environment state of the vehicle;

[0014] The execution mechanism is used for receiving and executing the control instructions.

[0015] Preferably, the gesture sensor, waterproof microphone, touch screen, environment perception sensor, and execution mechanism are connected with the vehicle-mounted controller respectively.

[0016] Preferably, the gesture sensor, waterproof microphone, and touch screen are integrated on the same interactive interface, wherein the interactive interface comprises a shell composed of a front frame and a rear cover; the touch screen is fixed on the rear cover, and a flexible sealing strip is arranged between the touch screen and the front frame; the waterproof microphone and gesture sensor are installed on the front frame; and the surface of the touch screen is covered with a waterproof film.

[0017] Preferably, the surface of the gesture sensor is wrapped with a transparent waterproof cover, and the transparent waterproof cover is covered with an anti-fog coating.

[0018] Preferably, the touch screen is a capacitive pressure touch screen.

[0019] Preferably, the interactive interface is arranged on the steering wheel.

[0020] Preferably, the environment perception sensor comprises a water level sensor and an acceleration sensor, and the water level sensor and the acceleration sensor are connected with the vehicle-mounted controller respectively, wherein the water level sensor is used for collecting the water depth of the vehicle to determine the water state of the vehicle, and the acceleration sensor is used for determining the attitude of the vehicle.

[0021] The application further provides a multi-modal interactive control method for an amphibious vehicle, which uses the multi-modal interactive control system for an amphibious vehicle according to the above method, and the method comprises the following steps:

[0022] Step S1, obtaining user gesture signals, voice signals, and touch signals;

[0023] Step S2, obtaining the current environment state of the vehicle, including land, water, and transition state;

[0024] Step S3, after pre-processing and feature extraction of the user gesture signals and voice signals, identifying gestures, voices, and touch and obtaining corresponding confidence.

[0025] Step S4, according to the current environment state of the vehicle, setting the confidence weight of gesture, voice, touch control recognition result;

[0026] Step S5, combining the confidence of gesture, voice, touch control recognition result and the corresponding confidence weight, determining the user interaction mode and recognition result through weighted decision fusion and generating a control instruction;

[0027] Step S6, the actuator executes the corresponding operation according to the control instruction.

[0028] Preferably, in the step S3, the voice signal is input into a bidirectional LSTM network for voice recognition and obtains the corresponding confidence after feature extraction by the MFCC algorithm; the gesture signal is input into a CNN network for gesture recognition and obtains the corresponding confidence after feature extraction by the HOG algorithm.

[0029] Preferably, in the step S5, the system pre-establishes an "interaction input-control instruction" mapping table, that is, according to the mapping table, the input user interaction mode and recognition result can obtain the to-be-executed control instruction.

[0030] The technical effects of the present application are:

[0031] (1) The system of the present application has simple structure, multiple sensors, supports multi-modal control input, realizes multi-modal data acquisition, and provides reliable data support for subsequent control decision.

[0032] (2) The system of the present application is designed to be compatible with waterproof performance, fully adapting to water work scenarios.

[0033] (3) For multi-modal data, the present application realizes dynamic multi-modal fusion decision according to real-time environmental changes, that is, the system automatically switches the optimal interaction mode according to the environment, improving driving safety and convenience. BRIEF DESCRIPTION OF DRAWINGS

[0034] Figure 1 A multi-modal interaction control system control principle block diagram of an amphibious vehicle is provided for the embodiments of the present application.

[0035] Figure 2 An interactive interface partial sectional view is provided for the embodiments of the present application.

[0036] Figure 2 Middle: waterproof cover 1, gesture sensor 2, waterproof microphone 3, front frame 4, flexible sealing strip 5, touch screen 6, rear cover 7. DETAILED DESCRIPTION

[0037] The specific embodiments of the present application are further described in detail below with the accompanying drawings and the description of the embodiments, which are intended to help the skilled in the art to have a more complete, accurate and in-depth understanding of the inventive concept and technical solutions of the present application, and to help the implementation. It should be noted that the terms "first", "second" and the like used in the present application are only for the convenience of describing the technical solutions and are used as a distinction between components, and the corresponding component configurations may be the same or different, and the present application is not limited to this. In order to make the technical solutions of the present application more clear, the present application is explained and described by the following embodiments.

[0038] The present embodiment provides a multimodal interactive control system for an amphibious vehicle, as shown in the accompanying drawings, the system comprises a gesture sensor, a waterproof microphone, a touch screen, an environmental perception sensor, a vehicle-mounted controller, an actuator, wherein: Figure 1

[0039] The gesture sensor is used to collect user gesture signals and send them to the vehicle-mounted controller;

[0040] The waterproof microphone is used to collect user voice signals and send them to the vehicle-mounted controller;

[0041] The touch screen is used to collect user touch signals and send them to the vehicle-mounted controller;

[0042] The environmental perception sensor is used to obtain the current environmental state of the vehicle, including land, water, and transition state;

[0043] The vehicle-mounted controller is used to identify and decide the control instructions according to the current environmental state of the vehicle, the user gesture signals, voice signals, and touch signals;

[0044] The actuator is used to receive and execute the control instructions. The gesture sensor, waterproof microphone, touch screen, environmental perception sensor, and actuator are connected to the vehicle-mounted controller.

[0045] Specifically, the vehicle-mounted controller of the present embodiment uses an MCU (microcontroller) with model number NXP S32G, and other controllers can also be selected according to actual conditions during implementation. Data is transmitted between the devices of the system through Ethernet with a delay of <50ms.

[0046] In order to facilitate system layout and improve system integration, the gesture sensor 2, waterproof microphone 3, and touch screen 6 are integrated on the same interactive interface in the present embodiment. Among them, as shown in the accompanying drawings, Figure 2 ​As shown, the interactive interface includes a shell composed of a front frame 4 and a rear cover 7; the touch screen 6 is fixed on the rear cover 7, and a flexible sealing strip 5 is arranged between the touch screen 6 and the front frame 4 to prevent water seepage. In a preferred embodiment of the present application, the contact end of the flexible sealing strip 5 with the touch screen 6 is tapered. Since the contact area is small, the friction torque is small, and the tapered sealing is better than the flat sealing, further improving the waterproof effect. At the same time, the touch screen 6 of the present embodiment adopts a capacitive pressure touch screen, which performs touch control recognition through changes in capacitance and pressure, improving the touch control accuracy when touching with wet hands and reducing the false touch rate. The present embodiment further covers the surface of the touch screen 6 with a 0.3mm thick PET waterproof film, thereby preventing water from entering the touch screen 6 and protecting the equipment safety.

[0047] A waterproof microphone 3 and a gesture sensor 2 are mounted on the front frame 4 of the interactive interface shell, wherein the arrangement of the gesture sensor 2 needs to ensure that the recognition range covers the habitual area of the driver's gestures, for example, the 0-20cm operation area of the driver's right hand; the waterproof microphone 3 array (sensitivity: -38dB±2dB) can be distributed in a ring shape on the front frame 4, and the specific installation position can be flexibly adjusted according to actual needs. The gesture sensor 2 can be an infrared gesture sensor (model: OP1011), which supports 6 degrees of freedom motion capture (pitch / roll / deflection angle error <5°). To prevent the gesture sensor 2 from being damaged by water when wading, the present embodiment further wraps a transparent waterproof cover 1 around the gesture sensor 2 to achieve waterproofing, which is usually made of acrylic material. Further, the transparent waterproof cover 1 is further covered with an anti-fog coating to avoid fog affecting the recognition accuracy of the gesture sensor 2.

[0048] At the same time, the present embodiment sets the interactive interface on the steering wheel, so that the driver's gesture signals, voice signals, and touch signals can be more easily collected, further improving the system integration.

[0049] The environment perception sensor of the present embodiment includes a water level sensor and an acceleration sensor, which are respectively connected to the vehicle-mounted controller. The water level sensor is used to collect the wading depth of the vehicle and send it to the vehicle-mounted controller, so that the vehicle-mounted controller can determine the wading state of the vehicle according to the wading depth. For example, when the wading depth is greater than a preset first depth threshold, it is considered to be in a wading state; when the preset second depth threshold ≤ wading depth ≤ preset first depth threshold, it is considered to be in a transition state; and when the wading depth is less than the preset second depth threshold, it is considered to be in a land state. The specific threshold can be flexibly set according to actual conditions. The acceleration sensor is used to determine the attitude of the vehicle to assist in wading state judgment, that is, to collect the acceleration changes of the object in three directions and send them to the vehicle-mounted controller, so that the vehicle-mounted controller can calculate the attitude and state of the vehicle. For example, when the vehicle enters or exits the water, the attitude of the vehicle will affect the water depth measurement accuracy, and the acceleration sensor can provide attitude correction parameters to ensure accurate water depth calculation.

[0050] According to the aforementioned multimodal interactive control system for an amphibious vehicle, the vehicle controller of this embodiment can analyze and process the collected user gesture signals, touch signals, and touch signals in combination with the vehicle's environmental status, and generate corresponding control instructions for execution by the actuator. Accordingly, this embodiment proposes a multimodal interactive control method for an amphibious vehicle, the method comprising:

[0051] Step S1: Acquire user gesture signals, voice signals, and touch signals;

[0052] Step S2: obtaining the current environmental state of the vehicle, including land, wading, and transition states;

[0053] Step S3: After preprocessing and feature extraction of the user's gesture signal and voice signal, gesture, voice, and touch recognition are performed and corresponding confidence levels are obtained;

[0054] Step S4: setting confidence weights for gesture, voice, and touch recognition results based on the vehicle's current environmental status;

[0055] Step S5: Combining the confidence levels of the gesture, voice, and touch recognition results and their corresponding confidence weights, the user interaction mode and recognition results are determined through weighted decision fusion and a control instruction is generated.

[0056] Step S6: The execution mechanism performs corresponding operations according to the control instruction.

[0057] Specifically, the method of this embodiment supports control input of multimodal interactive data and selects the interaction mode based on environmental intelligence. In step S3 of this embodiment, the touch signal often directly points to the function instruction. For example, when the user presses the window opening button, the instruction directly points to the window opening function. The user's operation intention is clearly visible and does not require preprocessing, feature extraction, recognition, etc. For voice signals and gesture signals, preprocessing is first required, including filtering, denoising, etc. to ensure data reliability. Then, the preprocessed data is extracted for voice and gesture features, and combined with a neural network, the user's operation intention contained in the voice and gesture is identified. Finally, the recognition result output by the Softmax function in the neural network will include the corresponding confidence level (the Softmax function converts the neural network output into a probability distribution, and the probability of the highest category is the confidence level). If the signal input of a certain modality is missing in actual work, the confidence level of the corresponding signal recognition defaults to 0.

[0058] In the embodiment, the voice signal is extracted by the MFCC algorithm (Mel Frequency Cepstral Coefficient algorithm, a voice signal feature extraction algorithm based on the hearing characteristics of human ears), and then input into a bidirectional LSTM network (Long Short-Term Memory network) for voice recognition and obtaining the corresponding confidence.

[0059] In the embodiment, the gesture signal is extracted by the HOG algorithm (Histogram of Oriented Gradients algorithm, a computer vision technology for image feature extraction), and then input into a CNN network (Convolutional Neural Network) for gesture recognition and obtaining the corresponding confidence.

[0060] In addition, in the presence of a touch signal, the confidence can be preset to a fixed value according to actual needs. In the case of high touch mis-touch rate, a lower confidence is set, and in the case of low touch mis-touch rate, a higher confidence is set. For example, the touch mis-touch rate of the embodiment is less than 5%, and the embodiment can be set to 100%.

[0061] Before specific application, the bidirectional LSTM network and the CNN network are trained and verified by a large amount of training data. In order to enhance the anti-interference ability and robustness, the embodiment uses an amphibious vehicle exclusive voice library as the training data of the voice recognition model, which contains 200,000 voice samples in the noise scene of high engine load (5000 rpm) and water spray (75 dB), so as to overcome the influence of engine noise (100 Hz-5 kHz) and water flow noise (2 kHz-8 kHz). The OpenPose-based skeleton key point detection algorithm is used to process the data set as the training data of the gesture recognition model, so as to overcome the gesture deformation problem caused by underwater light refraction.

[0062] In step S4 of the embodiment, the confidence weights of gesture, voice and touch recognition results are set according to the current environmental state of the vehicle. For example, in the land scene: the voice weight is 60%, the gesture weight is 30%, and the touch weight is 10%, so as to prefer voice interaction and free hands to deal with bumps. In the water and wet hand scene: the gesture weight is 70%, the voice weight is 20%, and the touch weight is 10%, so as to avoid wet hand mis-touch and strengthen gesture quick operation. In the transition scene: the gesture weight is 100%, the voice weight is 0%, and the touch weight is 0%, at this time the system forces to input a preset gesture (such as "continuous waving hands"), and shields voice and touch input to prevent mis-instruction. In specific implementation, the above weights can be flexibly set according to actual conditions. In a preferred embodiment of the present application, the existing weight distribution algorithm based on D-S evidence theory can also be used to automatically adjust the confidence weights of voice / gesture according to environmental parameters.

[0063] The embodiment in step S5 performs weighted decision fusion according to the confidence and the corresponding weight. For example, in a land scenario, the confidence of gesture, voice, and touch recognition is 95%, 60%, and 100% respectively, and the corresponding weight is 60% for voice, 30% for gesture, and 10% for touch. Then, we have:

[0064] Voice: 95% * 60% = 57%;

[0065] Gesture: 60% * 30% = 18%;

[0066] Touch: 100% * 10% = 10%;

[0067] Obviously, the score of the voice recognition result is the highest, and correspondingly, the system finally selects the voice recognition result as the current user operation intention.

[0068] After determining the user's interactive input and recognition result, the system can obtain the control instruction to be executed according to the pre-established "interactive input-control instruction" mapping table after inputting the user interactive mode and the recognition result. Part of the mapping table of the embodiment is shown in Table 1.

[0069]

[0070] Finally, the control instruction is sent to the execution mechanism for execution. The present application solves the problems of waterproofing, anti-fogging, and impact resistance at the hardware level, optimizes the multi-modal interactive recognition algorithm in the underwater and high-noise scenarios at the software level, forms a systematic solution, and finally provides a multi-modal interactive control system with waterproof performance, anti-interference ability, and intelligent decision-making, realizes seamless switching and efficient control between land and water scenarios, and improves driving safety and convenience.

[0071] The present application has been described above with reference to the drawings. Obviously, the specific implementation of the present application is not limited to the above-described manner. Any non-essential improvement or direct application of the above-described concept and technical solution of the present application to other occasions is within the scope of protection of the present application.

Claims

1. A multimodal interactive control system for an amphibious vehicle, characterized by: The system includes a gesture sensor, a waterproof microphone, a touch screen, an environmental perception sensor, an onboard controller, and an actuator, wherein: The gesture sensor is used to collect user gesture signals and send them to the vehicle controller; The waterproof microphone is used to collect user voice signals and send them to the vehicle controller; The touch screen is used to collect user touch signals and send them to the vehicle controller; The environmental perception sensor is used to obtain the current environmental status of the vehicle, including land, wading, and transition states; The vehicle controller is used to identify the user's gesture signals, voice signals, and touch signals according to the current environmental state of the vehicle and make decisions to generate control instructions; The execution mechanism is used to receive and execute the control instruction.

2. The multimodal interactive control system for an amphibious vehicle according to claim 1, characterized in that: The gesture sensor, waterproof microphone, touch screen, environment perception sensor, and actuator are respectively connected to the vehicle-mounted controller.

3. The multimodal interactive control system for an amphibious vehicle according to claim 1, characterized in that: The gesture sensor, waterproof microphone, and touch screen are integrated on the same interactive interface, wherein the interactive interface includes a shell consisting of a front frame and a back cover; the touch screen is fixed to the back cover, and a flexible sealing strip is provided between the touch screen and the front frame; the waterproof microphone and gesture sensor are installed on the front frame; and the surface of the touch screen is covered with a waterproof film.

4. The multimodal interactive control system for an amphibious vehicle according to claim 3, characterized in that: The surface of the gesture sensor is wrapped with a transparent waterproof cover, and the transparent waterproof cover is covered with an anti-fog coating.

5. The multimodal interactive control system for an amphibious vehicle according to claim 3, characterized in that: The touch screen adopts a capacitive pressure touch screen.

6. A multimodal interactive control system for an amphibious vehicle according to any one of claims 3 to 5, characterized in that: The interactive interface is arranged on the steering wheel.

7. The multimodal interactive control system for an amphibious vehicle according to claim 1, characterized in that: The environmental perception sensor includes a water level sensor and an acceleration sensor, which are respectively connected to the vehicle controller. The water level sensor is used to collect the vehicle's wading depth to determine the vehicle's wading status; the acceleration sensor is used to determine the vehicle's posture.

8. A multimodal interactive control method for an amphibious vehicle, using a multimodal interactive control system for an amphibious vehicle according to any one of claims 1 to 7, characterized in that: The method comprises: Step S1: Acquire user gesture signals, voice signals, and touch signals; Step S2: obtaining the current environmental state of the vehicle, including land, wading, and transition states; Step S3: After preprocessing and feature extraction of the user's gesture signal and voice signal, gesture, voice, and touch recognition are performed and corresponding confidence levels are obtained; Step S4: setting confidence weights for gesture, voice, and touch recognition results based on the vehicle's current environmental status; Step S5: Combining the confidence levels of the gesture, voice, and touch recognition results and their corresponding confidence weights, the user interaction mode and recognition results are determined through weighted decision fusion and a control instruction is generated. Step S6: The execution mechanism performs corresponding operations according to the control instruction.

9. The multi-modal interactive control method for an amphibious vehicle according to claim 8, characterized in that: In step S3, after the voice signal is feature extracted by the MFCC algorithm, it is input into the bidirectional LSTM network for voice recognition and the corresponding confidence level is obtained; after the gesture signal is feature extracted by the HOG algorithm, it is input into the CNN network for gesture recognition and the corresponding confidence level is obtained.

10. The multi-modal interactive control method for an amphibious vehicle according to claim 8, characterized in that: In step S5, the system pre-establishes an "interaction input-control instruction" mapping table, that is, according to the mapping table, the control instruction to be executed can be obtained after inputting the user interaction mode and the recognition result.