In-vehicle control method and system based on multi-modal interaction
By using a multimodal interaction method and collecting data from multiple sensors to construct a three-element collaborative control model, the limitations of existing in-vehicle control schemes in terms of scenarios and the lack of safety mechanisms are solved, thereby improving the safety and accuracy of in-vehicle control and reducing the false trigger rate.
Patent Information
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- DONGFENG MOTOR GRP
- Filing Date
- 2026-02-06
- Publication Date
- 2026-05-12
AI Technical Summary
Existing in-vehicle control solutions suffer from limitations in scenarios, lack of safety mechanisms, and insufficient multimodal collaboration, leading to problems such as distraction, low recognition accuracy, and high false trigger rate during driving.
By using a multimodal interaction method, multi-source data is collected using near-infrared cameras, millimeter-wave radar, and in-vehicle environment sensors to construct a three-dimensional collaborative control model of face-gesture-vehicle status. Combined with a safety arbitration layer to arbitrate according to preset rules, the control command execution layer ultimately drives the operation of in-vehicle devices.
It improves the safety, accuracy, and scenario adaptability of in-vehicle control, reduces the false trigger rate, and is suitable for safety control in high-frequency interaction scenarios during driving.
Smart Images

Figure CN122009052A_ABST
Abstract
Description
Technical Field
[0001] This invention relates to the field of human-machine interaction technology for intelligent vehicles, and in particular to an in-vehicle control method and system based on multimodal interaction. Background Technology
[0002] With the rapid development of automotive intelligence, in-vehicle interaction control methods are gradually evolving towards multimodal approaches. Currently, the mainstream in-vehicle control methods mainly include manual operation based on vehicle infotainment system soft buttons and intelligent voice control based on voice command recognition.
[0003] The manual operation solution based on in-vehicle infotainment system soft buttons integrates various vehicle control buttons into the system, allowing users to perform operations such as turning on music, adjusting windows, and controlling the air conditioning through the infotainment interface. However, this solution has significant drawbacks: users need to use both their hands and eyes simultaneously, which can distract the driver and pose a safety hazard while the vehicle is in motion; some control buttons have complex operation paths, making them difficult to locate quickly while driving; and manual touchscreen operation is difficult to accurately select target buttons in situations such as bumpy driving conditions.
[0004] Intelligent voice control solutions based on voice command recognition allow car owners to activate intelligent voice assistants by speaking and issue various vehicle control commands to control devices. However, this solution also has many problems: voice recognition is highly dependent on the in-car environment; when the environment is noisy, voice commands are difficult to recognize accurately; when there are other people in the car, voice operation can cause embarrassment and may disturb their rest; if the car owner has a strong dialect, it will lead to a decrease in voice recognition accuracy, or even misrecognition.
[0005] Therefore, there is an urgent need for an in-vehicle control method and system based on multimodal interaction to solve the problems of existing technologies. Summary of the Invention
[0006] The present invention aims to solve at least one of the technical problems existing in the prior art, and proposes an in-vehicle control method and system based on multimodal interaction.
[0007] In a first aspect, embodiments of the present invention provide an in-vehicle control method based on multimodal interaction, comprising:
[0008] Multimodal data of the vehicle is acquired through multiple sensors, including facial image data of occupants, gesture spatial trajectory data, in-vehicle environmental data, and vehicle operating status data.
[0009] Based on the facial image data, facial liveness detection is performed to obtain facial detection results; based on gesture spatial trajectory data, gesture matching is performed to obtain gesture matching results.
[0010] A safety arbitration rule base is constructed, which contains the correspondence between control commands and constraints on face detection results, gesture matching results, in-vehicle environment data and vehicle operating status data. Based on face recognition results, gesture recognition confidence, in-vehicle environment data and vehicle operating status data, safety arbitration is performed on control commands.
[0011] Control commands approved by safety arbitration are sent to the corresponding in-vehicle devices through the control command execution layer, and the drive devices complete the control command operations.
[0012] Furthermore, multimodal data of the vehicle is acquired through various sensors. Specific methods include: acquiring facial image data of occupants through a near-infrared camera, acquiring gesture spatial trajectory data through millimeter-wave radar, acquiring in-vehicle environmental data through in-vehicle environmental sensors, and acquiring vehicle operating status data through the CAN bus.
[0013] Furthermore, the vehicle operating status data includes vehicle speed and gear position, and the in-vehicle environment data includes the interior of the vehicle. Concentration and light intensity.
[0014] Furthermore, based on the facial image data, facial liveness detection is performed to obtain facial detection results. The specific method includes: processing the facial image data using a radio frequency liveness detection method to distinguish between real people and fake targets, and improving the facial recognition accuracy under head tilt state through a local feature alignment network.
[0015] Furthermore, based on the gesture spatial trajectory data, gestures are matched to obtain gesture matching results. Specific methods include: constructing a 3D gesture trajectory model based on the fusion data of millimeter-wave radar and RGB camera to achieve gesture spatial positioning.
[0016] Furthermore, the specific rules of the security arbitration rule base include:
[0017] When the control command is to lower all windows, the following conditions must be met: facial recognition passes, gesture matching is successful, and the vehicle speed is less than the preset speed; otherwise, the command will be refused.
[0018] When the control command is to adjust the air conditioning temperature, the gesture matching must be successful, and there are no speed or gear restrictions.
[0019] When the control command is to activate emergency ventilation, facial recognition must be passed and the interior of the vehicle must be clear. If the concentration exceeds the preset threshold, the sunroof will open slightly and the air conditioner will circulate external air without requiring a gesture.
[0020] When the control command is to temporarily release the child lock, the following conditions must be met: the administrator's face detection is passed, the specific gesture is successfully matched, and the gear is in P position; otherwise, the rear seat control restriction will not be released.
[0021] Furthermore, it also includes incremental user adaptation steps, specifically including: expanding the gesture library by generating adversarial examples through GAN, realizing online learning of new gestures based on a lightweight CNN classifier; and automatically associating the user's facial ID with their commonly used gesture set to achieve personalized control command matching.
[0022] Furthermore, the gestures include non-contact hovering gestures, which include adjusting volume by rotating in the air, lowering the window by pressing down with the palm, and adjusting the temperature by clenching a fist and rotating.
[0023] Secondly, this invention also discloses an in-vehicle control system based on multimodal interaction, comprising: a data acquisition module, a multimodal fusion module, a safety arbitration module, and a control command execution module, wherein:
[0024] The data acquisition module is used to acquire multimodal data of the vehicle through various sensors. The multimodal data includes facial image data of the occupants, gesture spatial trajectory data, in-vehicle environmental data, and vehicle operating status data.
[0025] The multimodal fusion module is used to detect facial liveness based on the facial image data to obtain facial detection results; and to match gestures based on gesture spatial trajectory data to obtain gesture matching results.
[0026] The safety arbitration module is used to build a safety arbitration rule base. The rule base contains the correspondence between control commands and constraints on face detection results, gesture matching results, in-vehicle environment data and vehicle operating status data. Based on face recognition results, gesture recognition confidence, in-vehicle environment data and vehicle operating status data, the control commands are subjected to safety arbitration.
[0027] The control command execution module is used to send control commands that have passed safety arbitration to the corresponding in-vehicle devices through the control command execution layer, and drive the devices to complete the control command operations.
[0028] Thirdly, the present invention also discloses an electronic device, comprising:
[0029] One or more processors;
[0030] Memory, used to store one or more programs;
[0031] When the one or more programs are executed by the one or more processors, the one or more processors implement the control method.
[0032] This invention discloses an in-vehicle control method and system based on multimodal interaction, aiming to solve the technical problems of existing in-vehicle control solutions, such as scenario limitations, lack of safety mechanisms, and insufficient multimodal collaboration. The method constructs a three-dimensional collaborative control model of "face-gesture-vehicle status," collecting multi-source data through near-infrared cameras, millimeter-wave radar, and in-vehicle environmental sensors. Data processing and feature extraction are performed by a multimodal fusion engine, and a safety arbitration layer arbitrates control commands according to preset rules. Finally, the control command execution layer drives devices such as windows, air conditioning, and entertainment systems to complete corresponding operations. This invention introduces the vehicle CAN bus signal as a control enable condition and integrates radio frequency liveness detection and 3D gesture spatial positioning technology, effectively improving the safety, accuracy, and scenario adaptability of in-vehicle control, reducing false triggering rates, and is suitable for safety control in high-frequency interaction scenarios during driving. Attached Figure Description
[0033] Figure 1 A flowchart illustrating an in-vehicle control method based on multimodal interaction provided in an embodiment of the present invention;
[0034] Figure 2 A structural block diagram of an in-vehicle control system based on multimodal interaction provided in an embodiment of the present invention;
[0035] Figure 3 This is a structural block diagram of an electronic device provided in an embodiment of the present invention. Detailed Implementation
[0036] To enable those skilled in the art to better understand the technical solutions of the present invention, exemplary embodiments of the present invention are described below in conjunction with the accompanying drawings, including various details of the embodiments of the present invention to aid understanding. These should be considered merely exemplary. Therefore, those skilled in the art should recognize that various changes and modifications can be made to the embodiments described herein without departing from the scope and spirit of the present invention. Similarly, for clarity and brevity, descriptions of well-known functions and structures are omitted in the following description.
[0037] Where there is no conflict, the various embodiments of the present invention and the features thereof may be combined with each other.
[0038] As used herein, the term “and / or” includes any and all combinations of one or more related enumerated entries.
[0039] The terminology used herein is for the purpose of describing particular embodiments only and is not intended to limit the invention. As used herein, the singular forms “a” and “the” are intended to include the plural forms as well, unless the context clearly indicates otherwise. It will also be understood that when the terms “comprising” and / or “made of” are used in this specification, the presence of the stated feature, integral, step, operation, element, and / or component is specified, but the presence or addition of one or more other features, integrals, steps, operations, elements, components, and / or groups thereof is not excluded. Terms such as “connected” or “linked” are not limited to physical or mechanical connections but can include electrical connections, whether direct or indirect.
[0040] Unless otherwise specified, all terms used herein (including technical and scientific terms) have the same meaning as commonly understood by one of ordinary skill in the art. It will also be understood that terms such as those defined in commonly used dictionaries should be interpreted as having the meaning consistent with their meaning in the context of the relevant art and the invention, and will not be interpreted as having an idealized or overly formal meaning unless expressly so defined herein.
[0041] In the technical solution of this invention, the collection, storage, use, processing, transmission, provision, and disclosure of user personal information all comply with relevant laws and regulations and do not violate public order and good morals. The use of user data in this technical solution follows relevant national laws and regulations (e.g., the "Information Security Technology - Personal Information Security Specification"). For example: appropriate measures are taken for personal information access control; restrictions are imposed on the display of personal information; the purpose of using personal information does not exceed the scope of direct or reasonable association; and explicit identity targeting is eliminated when using personal information to avoid precisely locating a specific individual.
[0042] In related technologies, existing solutions and methods still have problems such as scenario limitations, lack of security mechanisms, and insufficient multimodal collaboration: some solutions rely on external cameras to achieve door opening control, which cannot solve the recognition problem in complex environments such as changes in in-vehicle lighting and limb obstruction, and do not cover high-frequency interaction needs during driving such as seat adjustment and air conditioning control; some gesture window control patents only determine the target window based on forearm posture, without associating with safety parameters such as vehicle speed, which poses a high risk of accidental window opening at high speeds; and some facial recognition solutions lack a liveness detection step, making them vulnerable to photo and video attacks; existing solutions are mostly single-modal independent applications, with a high false trigger rate in complex scenarios, and some liveness detection patents cannot effectively distinguish between user intent and control commands, further affecting control accuracy and user experience.
[0043] To address at least one of the technical problems existing in the aforementioned related technologies, the present invention provides an in-vehicle control method and system based on multimodal interaction.
[0044] This implementation discloses an in-vehicle control method based on multimodal interaction, such as... Figure 1 ,include:
[0045] S100. Acquire multimodal data of the vehicle through multiple sensors, including facial image data of occupants, gesture spatial trajectory data, in-vehicle environment data, and vehicle operating status data;
[0046] In this embodiment, multimodal vehicle data is acquired through multiple sensors. Specifically, the method includes: acquiring facial image data of occupants using a near-infrared camera; acquiring gesture spatial trajectory data using millimeter-wave radar; acquiring in-vehicle environmental data using in-vehicle environmental sensors; and acquiring vehicle operating status data via the CAN bus. The vehicle operating status data includes vehicle speed and gear position, and the in-vehicle environmental data includes data from the interior of the vehicle. Concentration and light intensity.
[0047] Specifically, near-infrared cameras collect facial image data of occupants inside the vehicle for subsequent facial recognition and liveness detection; millimeter-wave radar collects spatial trajectory data of gestures, which, combined with RGB cameras, enables precise capture of gestures; and in-vehicle environmental sensors collect data on the interior of the vehicle. Environmental parameters such as concentration and light intensity; vehicle operating status parameters such as vehicle speed and gear position are obtained through the CAN bus interface to provide data support for safety arbitration.
[0048] S200. Based on the facial image data, perform facial liveness detection to obtain facial detection results; based on gesture spatial trajectory data, perform gesture matching to obtain gesture matching results;
[0049] In this embodiment, facial liveness detection is performed based on the facial image data to obtain facial detection results. The specific method includes: processing the facial image data using a radio frequency liveness detection method to distinguish between real people and fake targets, and improving the facial recognition accuracy under head tilt state through a local feature alignment network.
[0050] Specifically, radio frequency liveness detection technology is used to process the collected facial image data. By using the reflection characteristics of wireless radio frequency signals, real people are distinguished from fake targets such as fake masks and photos, thus avoiding attacks. At the same time, Local Feature Alignment Network (LFA-Net) is used to improve the facial recognition accuracy when the head is turned more than 30°, so as to adapt to the recognition needs of people in different postures inside the vehicle.
[0051] In this embodiment, gestures are matched based on gesture spatial trajectory data to obtain gesture matching results. Specifically, the method includes constructing a 3D gesture trajectory model based on fused data from millimeter-wave radar and an RGB camera to achieve gesture spatial positioning. The gestures include non-contact hovering gestures, such as adjusting volume by rotating in the air, lowering a window by pressing down with the palm, and adjusting temperature by clenching a fist and rotating.
[0052] Specifically, based on the fusion data of millimeter-wave radar and RGB camera, a 3D gesture trajectory model is constructed to achieve spatial positioning of gestures with a positioning error of less than 2cm; it supports non-contact hover gesture recognition, such as adjusting volume by rotating in the air, lowering the window by pressing down with the palm, and adjusting the temperature by clenching the fist and rotating, to meet the needs of different control scenarios.
[0053] S300. Construct a safety arbitration rule base, which includes the correspondence between control commands and constraints on face detection results, gesture matching results, in-vehicle environment data and vehicle operating status data. Based on face recognition results, gesture recognition confidence, in-vehicle environment data and vehicle operating status data, perform safety arbitration on control commands.
[0054] In this implementation, the rule base clearly defines the constraints and response actions for various control commands, including facial recognition results, gesture matching results, in-vehicle environment data, and vehicle operating status data. It receives the output facial recognition results, gesture recognition results, and confidence levels, and, combined with the collected in-vehicle environment parameters and vehicle operating status parameters, performs safety arbitration on the control commands according to preset rules in the rule base to determine whether the execution of the control command is permitted.
[0055] The specific rules of the security arbitration rule base include:
[0056] When the control command is to lower all windows, the following conditions must be met: facial recognition passes, gesture matching is successful, and the vehicle speed is less than the preset speed; otherwise, the command will be refused.
[0057] When the control command is to adjust the air conditioning temperature, the gesture matching must be successful, and there are no speed or gear restrictions.
[0058] When the control command is to activate emergency ventilation, facial recognition must be passed and the interior of the vehicle must be clear. If the concentration exceeds the preset threshold, the sunroof will open slightly and the air conditioner will circulate external air without requiring a gesture.
[0059] When the control command is to temporarily release the child lock, the following conditions must be met: the administrator's face detection is passed, the specific gesture is successfully matched, and the gear is in P position; otherwise, the rear seat control restriction will not be released.
[0060] In some embodiments, the specific rules of the security arbitration rule base include:
[0061] Full window lowering command: The safety arbitration layer will only allow the lowering command to be executed if the facial liveness detection is passed, the gesture matching is successful and the vehicle speed is <30km / h. Otherwise, the command will be refused to avoid the safety risk of accidentally opening the window while driving at high speed.
[0062] Air conditioning temperature adjustment command: As long as the gesture recognition confidence level is ≥0.7, there are no speed or gear restrictions, ensuring that users can conveniently adjust the air conditioning temperature under different driving conditions;
[0063] Emergency ventilation activation command: When facial recognition passes and the vehicle interior... When the concentration is greater than 1000ppm, the sunroof will automatically open slightly and the air conditioning will automatically circulate external air without the need for hand gestures, ensuring the air quality inside the vehicle.
[0064] Temporary release command for child lock: The administrator's facial recognition must be successful, a specific gesture must be matched, and the gear must be in P position. Otherwise, the rear seat control restrictions will not be released to ensure the safety of children riding in the car.
[0065] S400. Control commands approved through safety arbitration are sent to the corresponding in-vehicle devices via the control command execution layer, driving the devices to complete the control command operations. Specifically, control commands approved through safety arbitration are sent from the control command execution layer to the corresponding in-vehicle devices, such as window control systems, air conditioning control systems, and entertainment systems, driving the devices to complete preset operations and achieving precise control of the in-vehicle devices.
[0066] In some embodiments, an in-vehicle control method based on multimodal interaction further includes an incremental user adaptation step, specifically including: generating adversarial examples to expand the gesture library through GAN, realizing online learning of new gestures based on a lightweight CNN classifier; automatically associating the user's facial ID with its commonly used gesture set to achieve personalized control command matching.
[0067] To enhance the personalized user experience, this embodiment also includes an incremental user adaptation step: after a user completes the recording of a new gesture, the system generates a 3D key point time sequence, generates adversarial examples to expand the gesture library through GAN, and then updates the weights of the lightweight CNN classifier to achieve online learning of the new gesture; at the same time, the system automatically associates the user's facial ID with its commonly used gesture set, such as binding the "V gesture" with "play music", to achieve personalized control command matching and improve the user's ease of operation.
[0068] This invention discloses an in-vehicle control method based on multimodal interaction, aiming to address the technical problems of existing in-vehicle control solutions, such as scenario limitations, lack of safety mechanisms, and insufficient multimodal collaboration. The method constructs a three-dimensional collaborative control model of "face-gesture-vehicle status," collecting multi-source data through near-infrared cameras, millimeter-wave radar, and in-vehicle environmental sensors. This data is processed and features extracted by a multimodal fusion engine. A safety arbitration layer arbitrates control commands according to preset rules, and finally, the control command execution layer drives devices such as windows, air conditioning, and the entertainment system to complete the corresponding operations. This invention introduces the vehicle CAN bus signal as a control enable condition and integrates radio frequency liveness detection and 3D gesture spatial positioning technology, effectively improving the safety, accuracy, and scenario adaptability of in-vehicle control, reducing the false trigger rate, and is suitable for safety control in high-frequency interaction scenarios during driving.
[0069] Based on the same inventive concept, embodiments of the present invention also provide an in-vehicle control system based on multimodal interaction, such as... Figure 2 It includes: a data acquisition module, a multimodal fusion module, a security arbitration module, and a control command execution module, wherein:
[0070] The data acquisition module is used to acquire multimodal data of the vehicle through various sensors. The multimodal data includes facial image data of the occupants, gesture spatial trajectory data, in-vehicle environmental data, and vehicle operating status data.
[0071] The multimodal fusion module is used to detect facial liveness based on the facial image data to obtain facial detection results; and to match gestures based on gesture spatial trajectory data to obtain gesture matching results.
[0072] The safety arbitration module is used to build a safety arbitration rule base. The rule base contains the correspondence between control commands and constraints on face detection results, gesture matching results, in-vehicle environment data and vehicle operating status data. Based on face recognition results, gesture recognition confidence, in-vehicle environment data and vehicle operating status data, the control commands are subjected to safety arbitration.
[0073] The control command execution module is used to send control commands that have passed safety arbitration to the corresponding in-vehicle devices through the control command execution layer, and drive the devices to complete the control command operations.
[0074] The specific working methods of the data acquisition module, multimodal fusion module, security arbitration module, and control command execution module have been described in detail in the above methods, and will not be repeated here in this embodiment.
[0075] Based on the same inventive concept, embodiments of the present invention also provide an electronic device. Figure 3 This is a structural block diagram of an electronic device provided in an embodiment of the present invention. Figure 3 As shown, an embodiment of the present invention provides an electronic device including: one or more processors 101, a memory 102, and one or more I / O interfaces 103. The memory 102 stores one or more programs, which, when executed by the one or more processors, cause the one or more processors to implement any of the control methods described in the above embodiments; the one or more I / O interfaces 103 are connected between the processor and the memory, configured to enable information interaction between the processor and the memory.
[0076] The processor 101 is a device with data processing capabilities, including but not limited to a central processing unit (CPU); the memory 102 is a device with data storage capabilities, including but not limited to random access memory (RAM, more specifically SDRAM, DDR, etc.), read-only memory (ROM), electrically erasable programmable read-only memory (EEPROM), and flash memory (FLASH); the I / O interface (read / write interface) 103 is connected between the processor 101 and the memory 102, and can realize information interaction between the processor 101 and the memory 102, including but not limited to a data bus (Bus).
[0077] In some embodiments, the processor 101, memory 102, and I / O interface 103 are interconnected via bus 104, and thus connected to other components of the computing device.
[0078] In some embodiments, the one or more processors 101 include a field-programmable gate array.
[0079] This invention also provides a computer-readable medium. The computer-readable medium stores a computer program, which, when executed by a processor, implements the steps of any of the control methods described in the above embodiments. The computer-readable storage medium may be volatile or non-volatile.
[0080] This invention also provides a computer program product, including computer-readable code, or a non-volatile computer-readable storage medium carrying computer-readable code. When the computer-readable code is run in the processor of an electronic device, the processor in the electronic device executes the above-described control method.
[0081] Those skilled in the art will understand that all or some of the steps, systems, and apparatuses disclosed above, and their functional modules / units, can be implemented as software, firmware, hardware, or suitable combinations thereof. In hardware implementations, the division between functional modules / units mentioned above does not necessarily correspond to the division of physical components; for example, a physical component may have multiple functions, or a function or step may be performed collaboratively by several physical components. Some or all physical components may be implemented as software executed by a processor, such as a central processing unit, digital signal processor, or microprocessor, or as hardware, or as an integrated circuit, such as an application-specific integrated circuit (ASIC). Such software can be distributed on a computer-readable storage medium, which may include computer storage media (or non-transitory media) and communication media (or transient media).
[0082] As is known to those skilled in the art, the term computer storage medium includes volatile and non-volatile, removable and non-removable media implemented in any method or technology for storing information, such as computer-readable program instructions, data structures, program modules, or other data. Computer storage media includes, but is not limited to, random access memory (RAM), read-only memory (ROM), erasable programmable read-only memory (EPROM), static random access memory (SRAM), flash memory or other memory technologies, portable compact disc read-only memory (CD-ROM), digital versatile disc (DVD) or other optical disc storage, magnetic cartridges, magnetic tape, disk storage or other magnetic storage devices, or any other medium that can be used to store desired information and is accessible to a computer. Furthermore, it is known to those skilled in the art that communication media typically contain computer-readable program instructions, data structures, program modules, or other data in modulated data signals such as carrier waves or other transmission mechanisms, and may include any information delivery medium.
[0083] The computer-readable program instructions described herein can be downloaded from computer-readable storage media to various computing / processing devices, or downloaded via a network, such as the Internet, local area network, wide area network, and / or wireless network, to an external computer or external storage device. The network may include copper transmission cables, fiber optic transmission, wireless transmission, routers, firewalls, switches, gateway computers, and / or edge servers. A network adapter card or network interface in each computing / processing device receives the computer-readable program instructions from the network and forwards them to the computer-readable storage media in the respective computing / processing device.
[0084] The computer program instructions used to perform the operations of this invention may be assembly instructions, instruction set architecture (ISA) instructions, machine instructions, machine-dependent instructions, microcode, firmware instructions, state setting data, or source code or object code written in any combination of one or more programming languages, including object-oriented programming languages such as Smalltalk, C++, etc., and conventional procedural programming languages such as the "C" language or similar programming languages. The computer-readable program instructions may be executed entirely on the user's computer, partially on the user's computer, as a standalone software package, partially on the user's computer and partially on a remote computer, or entirely on a remote computer or server. In cases involving a remote computer, the remote computer may be connected to the user's computer via any type of network—including a local area network (LAN) or a wide area network (WAN)—or may be connected to an external computer (e.g., via the Internet using an Internet service provider). In some embodiments, electronic circuitry, such as programmable logic circuitry, field-programmable gate arrays (FPGAs), or programmable logic arrays (PLAs), is personalized by utilizing state information from the computer-readable program instructions. This electronic circuitry can execute the computer-readable program instructions to implement various aspects of the invention.
[0085] The computer program product described herein can be implemented specifically through hardware, software, or a combination thereof. In one alternative embodiment, the computer program product is specifically embodied in a computer storage medium; in another alternative embodiment, the computer program product is specifically embodied in a software product, such as a software development kit (SDK), etc.
[0086] Various aspects of the present invention are described herein with reference to flowchart illustrations and / or block diagrams of methods, apparatus (systems), and computer program products according to embodiments of the invention. It should be understood that each block of the flowchart illustrations and / or block diagrams, and combinations of blocks in the flowchart illustrations and / or block diagrams, can be implemented by computer-readable program instructions.
[0087] These computer-readable program instructions can be provided to a processor of a general-purpose computer, a special-purpose computer, or other programmable data processing apparatus to produce a machine such that, when executed by the processor of the computer or other programmable data processing apparatus, they create means for implementing the functions / actions specified in one or more blocks of the flowchart and / or block diagram. These computer-readable program instructions can also be stored in a computer-readable storage medium that causes a computer, programmable data processing apparatus, and / or other device to operate in a particular manner; thus, the computer-readable medium storing the instructions comprises an article of manufacture that includes instructions for implementing aspects of the functions / actions specified in one or more blocks of the flowchart and / or block diagram.
[0088] Computer-readable program instructions may also be loaded onto a computer, other programmable data processing apparatus, or other device to cause a series of operational steps to be performed on the computer, other programmable data processing apparatus, or other device to produce a computer-implemented process, thereby causing the instructions executed on the computer, other programmable data processing apparatus, or other device to perform the functions / actions specified in one or more boxes of a flowchart and / or block diagram.
[0089] The flowcharts and block diagrams in the accompanying drawings illustrate the architecture, functionality, and operation of possible implementations of systems, methods, and computer program products according to various embodiments of the present invention. In this regard, each block in a flowchart or block diagram may represent a module, segment, or portion of an instruction, which contains one or more executable instructions for implementing a specified logical function. In some alternative implementations, the functions marked in the blocks may occur in a different order than those shown in the drawings. For example, two consecutive blocks may actually be executed substantially in parallel, and they may sometimes be executed in reverse order, depending on the functions involved. It should also be noted that each block in the block diagrams and / or flowcharts, and combinations of blocks in the block diagrams and / or flowcharts, may be implemented using a dedicated hardware-based system that performs the specified function or action, or using a combination of dedicated hardware and computer instructions.
[0090] Example embodiments have been disclosed herein, and while specific terminology has been used, it is for illustrative purposes only and should be construed as such, and is not intended to be limiting. In some instances, it will be apparent to those skilled in the art that features, characteristics, and / or elements described in conjunction with particular embodiments may be used alone, or in combination with features, characteristics, and / or elements described in conjunction with other embodiments, unless otherwise expressly indicated. Therefore, those skilled in the art will understand that various changes in form and detail may be made without departing from the scope of the invention as set forth in the appended claims.
Claims
1. A vehicle in-vehicle control method based on multimodal interaction, characterized in that, include: Multimodal data of the vehicle is acquired through multiple sensors, including facial image data of occupants, gesture spatial trajectory data, in-vehicle environmental data, and vehicle operating status data. Based on the facial image data, facial liveness detection is performed to obtain facial detection results; Based on gesture spatial trajectory data, gestures are matched to obtain gesture matching results; A safety arbitration rule base is constructed, which contains the correspondence between control commands and constraints on face detection results, gesture matching results, in-vehicle environment data and vehicle operating status data. Based on face recognition results, gesture recognition confidence, in-vehicle environment data and vehicle operating status data, safety arbitration is performed on control commands. Control commands approved by safety arbitration are sent to the corresponding in-vehicle devices through the control command execution layer, and the drive devices complete the control command operations.
2. The control method according to claim 1, characterized in that, Multimodal data of the vehicle is acquired through a variety of sensors, including: acquiring facial image data of occupants through a near-infrared camera, acquiring gesture spatial trajectory data through millimeter-wave radar, acquiring in-vehicle environmental data through in-vehicle environmental sensors, and acquiring vehicle operating status data through the CAN bus.
3. The control method according to claim 1, characterized in that, The vehicle operating status data includes vehicle speed and gear position, and the in-vehicle environment data includes the interior of the vehicle. Concentration and light intensity.
4. The control method according to claim 1, characterized in that, Based on the facial image data, facial liveness detection is performed to obtain facial detection results. The specific method includes: processing the facial image data using radio frequency liveness detection method to distinguish between real people and fake targets, and improving the facial recognition accuracy under head tilt state through local feature alignment network.
5. The control method according to claim 1, characterized in that, Based on gesture spatial trajectory data, gestures are matched to obtain gesture matching results. Specific methods include: constructing a 3D gesture trajectory model based on the fusion data of millimeter-wave radar and RGB camera to achieve gesture spatial positioning.
6. The control method according to claim 1, characterized in that, The specific rules of the security arbitration rule base include: When the control command is to lower all windows, the following conditions must be met: facial recognition passes, gesture matching is successful, and the vehicle speed is less than the preset speed; otherwise, the command will be refused. When the control command is to adjust the air conditioning temperature, the gesture matching must be successful, and there are no speed or gear restrictions. When the control command is to activate emergency ventilation, facial recognition must be passed and the interior of the vehicle must be clear. If the concentration exceeds the preset threshold, the sunroof will open slightly and the air conditioner will circulate external air without requiring a gesture. When the control command is to temporarily release the child lock, the following conditions must be met: the administrator's face detection is passed, the specific gesture is successfully matched, and the gear is in P position; otherwise, the rear seat control restriction will not be released.
7. The control method according to claim 1, characterized in that, It also includes incremental user adaptation steps, specifically: expanding the gesture library by generating adversarial examples through GAN, and realizing online learning of new gestures based on a lightweight CNN classifier; automatically associating the user's facial ID with their commonly used gesture set to achieve personalized control command matching.
8. The control method according to claim 1, characterized in that, The gestures include non-contact hovering gestures, which include adjusting volume by rotating in the air, lowering the window by pressing down with the palm, and adjusting the temperature by clenching a fist and rotating.
9. An in-vehicle control system based on multimodal interaction, characterized in that, include: The module comprises a data acquisition module, a multimodal fusion module, a security arbitration module, and a control command execution module, among which: The data acquisition module is used to acquire multimodal data of the vehicle through various sensors. The multimodal data includes facial image data of the occupants, gesture spatial trajectory data, in-vehicle environmental data, and vehicle operating status data. The multimodal fusion module is used to detect facial liveness based on the facial image data to obtain facial detection results; and to match gestures based on gesture spatial trajectory data to obtain gesture matching results. The safety arbitration module is used to build a safety arbitration rule base. The rule base contains the correspondence between control commands and constraints on face detection results, gesture matching results, in-vehicle environment data and vehicle operating status data. Based on face recognition results, gesture recognition confidence, in-vehicle environment data and vehicle operating status data, the control commands are subjected to safety arbitration. The control command execution module is used to send control commands that have passed safety arbitration to the corresponding in-vehicle devices through the control command execution layer, and drive the devices to complete the control command operations.
10. An electronic device, characterized in that, include: One or more processors; Memory, used to store one or more programs; When the one or more programs are executed by the one or more processors, the one or more processors implement the control method as described in any one of claims 1 to 8.