Multi-modal interaction method, system and device based on intelligent cabin and vehicle
By combining multimodal interaction methods and signal priority rules with biometric verification, the smart cockpit achieves flexible integration and personalized management of multiple interaction methods, solving the problems of single interaction methods and insufficient security in existing technologies, and improving user experience and driving safety.
Patent Information
- Application Number
- CN202510673822.4
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-05-23
- Publication Date
- 2025-10-17
AI Technical Summary
Existing smart cockpit interaction methods mainly rely on voice and touch, which are difficult to adapt to the diverse interaction needs of users. Furthermore, they have limited signal fusion and processing capabilities in complex environments, and lack sufficient security and personalized verification, which affects user experience and driving safety.
It adopts a multimodal interaction method, integrating interaction methods such as voice, gesture, eye movement and facial expression. Combined with user permission management and signal priority rules, it dynamically adjusts signal weights through cross-validation binding of biometric data, generates interaction intent commands, and performs permission verification to control the operation of the execution mechanism.
It improves the flexibility, safety, and user experience of intelligent cockpit interaction, ensures the accuracy and safety of operation, and enhances the system's adaptability and personalized service capabilities.
Smart Images

Figure CN120803249A_ABST
Abstract
Description
TECHNICAL FIELD
[0001] The present application relates to the technical field of intelligent cockpit, and in particular to a multi-modal interaction method, system, device and vehicle based on intelligent cockpit. BACKGROUND
[0002] With the development of automobiles towards more intelligent direction, intelligent cockpit gradually becomes an indispensable part of modern vehicles. However, the current interaction means of intelligent cockpit is still relatively limited, mainly relying on traditional ways such as voice and touch, which is difficult to meet the growing demand of users for diversified interaction means. For example, in a specific situation, the user may prefer to use gestures or eye movements to interact, but the existing intelligent cockpit solutions usually lack support for these functions. In addition, the current interaction mode may not be natural and smooth in some cases, and the efficiency needs to be improved. Especially when driving, the driver may need to constantly adjust their line of sight or operate the touch screen, which not only distracts their attention, but also may adversely affect driving safety. SUMMARY
[0003] The purpose of the present application is to provide a multi-modal interaction method, system, device and vehicle based on intelligent cockpit, to solve one or more technical problems existing in the prior art, and at least provide a beneficial choice or create conditions.
[0004] The solution to the technical problem of the present application is: on the one hand, the present application provides a multi-modal interaction method based on intelligent cockpit, comprising the following steps: displaying a multi-modal interaction main interface of the intelligent cockpit; wherein the multi-modal interaction main interface comprises a multi-modal interaction instruction setting control, a user permission management control and a real-time feedback display area; in response to a trigger instruction of the user permission management control, displaying a user permission management sub-interface, establishing a user permission database, collecting and encrypting the user's biological feature data, and cross-verifying and binding the biological feature data with the configured permission level; in response to a trigger instruction of the multi-modal interaction instruction setting control, displaying a multi-modal interaction instruction setting sub-interface; through the multi-modal interaction instruction setting sub-interface, configuring a signal priority rule for the interaction mode of the intelligent cockpit; the signal priority rule is used to adjust the fusion weight of multi-modal input signals; the multi-modal input signals include voice signals, gesture signals, eye movement signals and facial expression signals; obtaining the multi-modal input signals of the current user, calling a multi-modal fusion processing algorithm, and generating an interaction intent instruction based on the signal priority rule; matching the multi-modal input signals with the biological feature data of the user permission database to obtain the permission level of the current user; The interaction intention instruction is matched with the permission level of the current user, and if the verification is passed, the instruction execution state is displayed in the real-time feedback display area, and the execution mechanism of the intelligent cockpit is controlled to complete the corresponding operation.
[0005] Further, the user permission management sub-interface includes a permission level configuration control, a biometric verification control, a permission range configuration control, and a permission setting saving control. In response to the trigger instruction of the user permission management control, a user permission management sub-interface is displayed, and a user permission database is established, including: In response to the trigger instruction of the permission level configuration control, a permission level configuration window is displayed, and the user permission level is set; wherein the permission level includes high-level permission, medium-level permission, and low-level permission; the high-level permission needs to be bound through cross verification of face and voiceprint; In response to the trigger instruction of the biometric verification control, a biometric feature input window is displayed, and the biometric feature data of the user is collected and encrypted; wherein the biometric feature data includes face image, voiceprint waveform, and live detection parameters; In response to the trigger instruction of the permission range configuration control, a permission range setting window is displayed, and a permission range whitelist of the operable function is configured for the specified user; the permission range whitelist includes the operation type allowed to be executed and the operation permission validity period; In response to the trigger instruction of the permission setting saving control, the permission level configuration, biometric feature data, and permission range whitelist are saved to the user permission database.
[0006] Further, according to the multi-modal interaction method based on the intelligent cockpit of claim 1, in response to the trigger instruction of the multi-modal interaction instruction setting control, a multi-modal interaction instruction setting sub-interface is displayed, including: In response to the trigger instruction of the multi-modal interaction instruction setting control, a multi-modal interaction instruction template selection window corresponding to the trigger instruction of the multi-modal interaction instruction setting control is displayed; wherein the multi-modal interaction instruction template selection window includes a custom creation control; In response to the trigger instruction of the custom creation control, a multi-modal interaction instruction setting sub-interface is displayed.
[0007] Further, the multi-modal interaction instruction setting sub-interface includes a gesture action library selection control, an eye movement tracking calibration control, and a multi-modal interaction instruction saving control. The multi-modal interaction instruction setting sub-interface is used to configure signal priority rules for the interaction mode of the intelligent cockpit, including: In response to a triggering instruction of the gesture action library selection control, a gesture action library configuration window is displayed to bind preset gesture action libraries to different scenes; wherein the gesture action library includes a basic gesture action set and an extended gesture action set, and the extended gesture action set is only enabled in a parking scene; In response to a triggering instruction of the eye movement tracking calibration control, an eye movement tracking calibration window is displayed to execute an eye movement trajectory calibration program and generate eye movement calibration parameters; the eye movement calibration parameters are used to correct the focal point coordinate deviation output by the eye movement tracking sensor; In response to a triggering instruction of the multi-modal interaction instruction saving control, the gesture action library binding relationship and the eye movement calibration parameters are saved and input into a multi-modal fusion processing algorithm to generate a scene-adaptive signal priority rule.
[0008] Further, the signal priority rule includes a signal quality evaluation rule, a scene classification rule, a dynamic weight allocation rule, and a conflict processing rule; The signal quality evaluation rule is used to quantitatively evaluate the quality of the multi-modal input signal to obtain a quality level of the multi-modal input signal; The scene classification rule is used to divide scene types according to the current state of the intelligent cockpit; The dynamic weight allocation rule is used to dynamically adjust the weights of signals of each modality according to the quality level of the current multi-modal input signal and the scene type; The conflict processing rule is used to solve the conflict of the multi-modal input signal according to a preset signal priority order and a user confirmation mechanism.
[0009] Further, the multi-modal input signal of the current user is obtained, a multi-modal fusion processing algorithm is called, and an interaction intent instruction is generated based on the signal priority rule, including: The multi-modal input signal is preprocessed and feature extracted to obtain a multi-modal signal feature vector; Based on the dynamic weight allocation rule defined in the signal priority rule, the multi-modal fusion processing algorithm is called to perform fusion processing on the multi-modal signal feature vector to generate a joint feature vector; The joint feature vector is input into an intent analysis model to generate an intent analysis result, and the intent analysis result is matched with a structured operation instruction in a preset instruction library to obtain an interaction intent instruction.
[0010] Further, the interaction intent instruction is matched with the permission level of the current user for verification, and if the verification is passed, an instruction execution state is displayed in the real-time feedback display area, and an execution mechanism of the intelligent cockpit is controlled to complete a corresponding operation, including: According to the permission level of the current user, the executable of the interactive intention instruction is queried from the permission range whitelist in the user permission database; If the interactive intention instruction is within the permission range whitelist, the verification is passed, the instruction execution state is displayed in the real-time feedback display area, and the execution mechanism of the intelligent cockpit is controlled to complete the corresponding operation; If the interactive intention instruction is not within the permission range whitelist, the verification is not passed, a permission conflict popup window is triggered, and the permission conflict reason is highlighted in the real-time feedback display area.
[0011] In another aspect, the present application provides a multi-modal interaction system based on an intelligent cockpit, which is used to implement the multi-modal interaction method based on an intelligent cockpit as described above.
[0012] In another aspect, the present application provides a multi-modal interaction device based on an intelligent cockpit, which includes an intelligent cockpit hardware device, a central processing unit, a display device, and an execution mechanism. The display device is used to display a multi-modal interaction main interface of the intelligent cockpit; wherein the multi-modal interaction main interface includes a multi-modal interaction instruction setting control, a user permission management control, and a real-time feedback display area; in response to a trigger instruction of the user permission management control, a user permission management sub-interface is displayed, a user permission database is established to collect and encrypt the biological feature data of the user, and the biological feature data is cross-verified and bound with the configured permission level; in response to a trigger instruction of the multi-modal interaction instruction setting control, a multi-modal interaction instruction setting sub-interface is displayed; through the multi-modal interaction instruction setting sub-interface, a signal priority rule is configured for the interaction mode of the intelligent cockpit; the signal priority rule is used to adjust the fusion weight of the multi-modal input signal; the multi-modal input signal includes a voice signal, a gesture signal, an eye movement signal, and a facial expression signal; The intelligent cockpit hardware device is used to obtain the multi-modal input signal of the current user; the intelligent cockpit hardware device includes a voice recognition device, a gesture recognition device, an eye movement tracking device, and a facial expression recognition device; The central processing unit is used to call a multi-modal fusion processing algorithm according to the multi-modal input signal, generate an interactive intention instruction based on the signal priority rule, match the multi-modal input signal with the biological feature data of the user permission database to obtain the permission level of the current user, and match and verify the interactive intention instruction with the permission level of the current user; if the verification is passed, the instruction execution state is displayed in the real-time feedback display area, and the execution mechanism is controlled to complete the corresponding operation.
[0013] In another aspect, the present application provides a vehicle integrated with the multi-modal interaction system based on an intelligent cockpit as described above.
[0014] The beneficial effects of the present application are: the intelligent cockpit based multi-modal interaction method provided by the present application realizes flexible configuration and real-time feedback of the intelligent cockpit interaction mode by displaying a main interface including a multi-modal interaction instruction setting control, a user permission management control and a real-time feedback display area. The method binds the cross verification of biological feature data and permission levels to ensure the operation permission security and personalized setting of different users. At the same time, through the fusion processing of multi-modal input signals such as voice, gesture, eye movement and facial expression, and according to the signal priority rule, the weight of each signal is dynamically adjusted to generate an interaction instruction in line with the user's intention, which not only improves the naturalness and efficiency of the interaction, but also enhances the user experience. Finally, the interaction intention instruction is matched and verified with the user permission level to control the execution mechanism to complete the corresponding operation, further ensuring the safety and accuracy of the operation. Therefore, the present application effectively improves the flexibility, safety and user experience of the intelligent cockpit interaction. The present application also provides a corresponding device, system and vehicle, and the beneficial effects of the device, system and vehicle are similar to those of the method, which will not be described here.
[0015] Other features and advantages of the present application will be set forth in the following description, and in part will become apparent from the description, or can be learned by practice of the present application. The objects and other advantages of the present application will be realized and achieved by the structure particularly pointed out in the description, claims and drawings. BRIEF DESCRIPTION OF DRAWINGS
[0016] The accompanying drawings are included to provide a further understanding of the technical solutions of the present application, and constitute a part of the specification, and are used together with the embodiments of the present application to explain the technical solutions of the present application, and do not constitute a limitation on the technical solutions of the present application.
[0017] Figure 1 is a flowchart of the intelligent cockpit based multi-modal interaction method provided by the present application; Figure 2 is a schematic diagram of the multi-modal interaction main interface provided by the present application; Figure 3 is a schematic diagram of the user permission management sub-interface provided by the present application; Figure 4 is a schematic diagram of the multi-modal interaction instruction setting sub-interface provided by the present application; Figure 5 is a schematic diagram of the multi-modal interaction instruction template selection window provided by the present application; Figure 6 is a structural diagram of the intelligent cockpit based multi-modal interaction system provided by the present application; Figure 7 is a structural diagram of the intelligent cockpit based multi-modal interaction device provided by the present application. DETAILED DESCRIPTION
[0018] For the purposes of the present application, the technical solutions and advantages will be more clearly apparent from the following detailed description in conjunction with the accompanying drawings and examples. It should be understood that the specific examples described herein are intended to explain the present application and are not intended to limit the present application.
[0019] The present application will be further described in conjunction with the accompanying drawings and specific examples. The described examples should not be considered as limiting the present application, and all other examples obtained by those of ordinary skill in the art without making creative efforts fall within the scope of the present application.
[0020] In the following description, "some embodiments" are referred to, which describe a subset of all possible embodiments, but it can be understood that "some embodiments" can be the same subset or different subsets of all possible embodiments, and can be combined with each other without conflict.
[0021] Unless otherwise defined, all technical and scientific terms used herein have the same meaning as commonly understood by one of ordinary skill in the art to which the present application belongs. The terms used herein are only for the purpose of describing the embodiments of the present application and are not intended to limit the present application.
[0022] With the improvement of the intelligent level of automobiles, the intelligent cockpit, as an important bridge connecting the driver and the vehicle, gradually changes its interaction mode from traditional physical buttons to more advanced technologies such as touch screens and voice recognition. However, most current intelligent cockpit systems still mainly rely on single or limited interaction modes for information input and output, such as setting navigation through a touch screen or adjusting the temperature in the vehicle using a voice assistant. Although this approach improves user experience to some extent, it is limited by the singularity of the interaction mode, making it difficult to meet the diverse needs of users in different scenarios, and also lacks consideration of personalized preferences of users, resulting in insufficient consistency and flexibility of operation experience.
[0023] Another major challenge of existing intelligent cockpit systems lies in the limitations of signal fusion and processing capabilities. When facing input signals from multiple modalities (such as voice, gesture, eye movement, etc.), traditional systems often cannot effectively evaluate and assign appropriate weights to each signal, leading to misjudgment or conflict in multi-task processing or complex environments. In addition, although some systems integrate biometric recognition technology to enhance security, the biometric verification mechanisms of these systems are usually simple and do not fully consider the importance of data encryption storage and cross-validation, thus there are certain security risks. This not only affects the trust of users, but also to some extent limits the potential of intelligent cockpit to develop towards higher levels of autonomous driving.
[0024] To solve the above problems, the application provides a multi-modal interaction method, system, device and vehicle based on an intelligent cockpit. The method integrates multiple interaction modes such as voice, gesture, eye movement and facial expression, and combines a personalized user permission management mechanism, thereby greatly enhancing the flexibility and security of the system. The method also uses an advanced multi-modal fusion processing algorithm to dynamically adjust the weight of each modal signal according to signal quality evaluation rules and scene classification rules, thereby ensuring the accuracy of the interaction intent instruction. Meanwhile, the user's biological feature data is strictly encrypted and stored, and a cross-verification binding mechanism of face and voiceprint is used to determine the permission level, thereby further improving the security performance of the system and providing a safer, more convenient and personalized interaction environment for the user.
[0025] First, the multi-modal interaction method based on an intelligent cockpit provided by the embodiments of the application will be described in detail below with reference to the accompanying drawings.
[0026] With reference to Figures 1 to 4 The implementation process of the multi-modal interaction method based on an intelligent cockpit provided by the embodiments of the application includes but is not limited to the following steps.
[0027] S110, display a multi-modal interaction main interface 100 of the intelligent cockpit.
[0028] With reference to Figure 2 The multi-modal interaction main interface 100 includes a multi-modal interaction instruction setting control 102, a user permission management control 103 and a real-time feedback display area 101.
[0029] In step S110, displaying the multi-modal interaction main interface 100 of the intelligent cockpit is the first step of user interaction with the system, which provides an intuitive and easy-to-operate platform for the user. The interface includes the multi-modal interaction instruction setting control 102, the user permission management control 103 and the real-time feedback display area 101, so that the user can conveniently adjust the interaction mode, manage personal permissions and immediately check the system response status. Through this design, the user's experience is improved, and the transparency and timeliness of information transmission are ensured, so that the user can always master the running status of the system.
[0030] S120, in response to a trigger instruction of the user permission management control 103, display a user permission management sub-interface 200, establish a user permission database, collect and encrypt the biological feature data of the user, and cross-verify and bind the biological feature data with the configured permission level.
[0031] In step S120, when the user triggers the user permission management control 103, the system will display the user permission management sub-interface 200, allowing the user to personalize the permission configuration. In this step, the system will create a user permission database that specifically stores the user's biometric data (such as facial recognition or fingerprint, etc.), and encrypt these sensitive data to protect user privacy. In addition, the user's biometric data is cross-verified and bound to the permission level configured by the user, which is crucial for ensuring that only authorized users can perform specific operations, thereby enhancing the security and reliability of the system.
[0032] In step S130, in response to the triggering instruction of the multi-modal interaction instruction setting control 102, the multi-modal interaction instruction setting sub-interface 300 is displayed.
[0033] In step S130, once the user triggers the multi-modal interaction instruction setting control 102, the system will display a multi-modal interaction instruction setting sub-interface 300, providing the user with an opportunity to customize the interaction mode. This interface allows users to adjust the priority rules of various input methods (such as voice, gestures, etc.) according to their own preferences and needs, thereby optimizing the convenience and efficiency of human-computer interaction. This function greatly enhances user experience, allowing different users to obtain the most suitable interaction mode in different scenarios.
[0034] In step S140, the multi-modal interaction instruction setting sub-interface 300 is used to configure signal priority rules for the interaction mode of the intelligent cockpit.
[0035] The signal priority rules are used to adjust the fusion weight of multi-modal input signals. The multi-modal input signals include voice signals, gesture signals, eye movement signals, and facial expression signals.
[0036] In step S140, the user can configure signal priority rules for different interaction modes of the intelligent cockpit through the multi-modal interaction instruction setting sub-interface 300. These rules determine how to dynamically adjust the fusion weight of multiple types of input signals (such as voice, gestures, eye movements, and facial expressions) when they are received, in order to more accurately understand the user's intent. In this way, the system can more flexibly handle complex user input situations, improve the accuracy of interaction instruction generation, and thus enhance the overall user experience.
[0037] In step S150, the multi-modal input signals of the current user are obtained, and a multi-modal fusion processing algorithm is called to generate an interaction intent instruction based on the signal priority rules.
[0038] In step S150, the system first acquires the multi-modal input signals from the user, and then analyzes and integrates these signals through a multi-modal fusion processing algorithm using pre-set signal priority rules. This not only enhances the system's understanding of complex environments, but also ensures that the user's true intentions can be accurately captured even in the presence of interference. Finally, based on the above analysis results, specific interaction intent instructions are generated to provide clear guidance for subsequent operations.
[0039] In step S160, the multi-modal input signals are matched with the biometric data in the user permission database to obtain the permission level of the current user. The interaction intent instruction is matched and verified with the permission level of the current user. If the verification is passed, the instruction execution state is displayed in the real-time feedback display area 101, and the execution mechanism of the intelligent cockpit is controlled to complete the corresponding operation.
[0040] In step S160, the multi-modal input signals of the user are first compared with the biometric data in the user permission database to confirm the identity and corresponding permission level of the user. Then, the system checks whether the generated interaction intent instruction complies with the user's permission range. If the verification is successful, the system will display the instruction execution state in the real-time feedback display area 101 and instruct the relevant execution mechanism of the intelligent cockpit to complete the specified operation. This not only ensures the safety of the operation, but also ensures that all actions are carried out within the legal and compliant framework.
[0041] In some embodiments of the present application, with reference to Figure 3 The user permission management sub-interface 200 includes a permission level configuration control 201, a biometric verification control 202, a permission range configuration control 203, and a permission setting saving control 204. In step S120, in response to the trigger instruction of the user permission management control 103, the user permission management sub-interface 200 is displayed, and the implementation process of establishing the user permission database includes but is not limited to the following steps.
[0042] In step S210, in response to the trigger instruction of the permission level configuration control 201, a permission level configuration window is displayed to set the user permission level.
[0043] The permission level includes high-level permission, medium-level permission, and low-level permission. High-level permission requires cross-verification binding through face and voiceprint.
[0044] In step S210, when the user triggers the permission level configuration control 201, the system displays a window dedicated to setting the user's permission level. This window allows the user to select different permission levels (such as high, medium, and low) according to actual needs. In particular, for high-level permissions that require access to sensitive information or perform critical operations, further confirmation of user identity is required through cross-validation binding of facial and voiceprint. This hierarchical permission management approach not only improves the security of the system, but also ensures that users with different permission levels can only access functions and services that match their permissions.
[0045] In step S220, in response to the triggering instruction of the biometric verification control 202, a biometric entry window is displayed, and the user's biometric data is collected and encrypted.
[0046] The biometric data includes facial images, voiceprint waveforms, and liveness detection parameters.
[0047] In step S220, when the user triggers the biometric verification control 202, the system guides the user to enter the biometric entry window, where various biometric data including facial images, voiceprint waveforms, and liveness detection parameters are collected. These data are then encrypted for storage to protect the user's privacy and security. In this way, the system can establish a reliable user identity verification mechanism to ensure that only verified users can obtain corresponding permissions, thereby providing solid security for subsequent operations.
[0048] In step S230, in response to the triggering instruction of the permission range configuration control 203, a permission range setting window is displayed to configure the permission range whitelist of the specified user for the operable functions.
[0049] The permission range whitelist includes the types of operations allowed to perform and the operation permission validity period.
[0050] In step S230, the user can access the permission range setting window by triggering the permission range configuration control 203, which allows the user to customize the detailed permission range whitelist for a specific account. This whitelist lists all operation types that the user can perform and their corresponding permission validity periods. For example, some users may be limited to viewing vehicle status and cannot make changes. This approach makes permission management more detailed and flexible, allowing for adjustments in permission allocation based on specific application scenarios and user needs, enhancing the adaptability and security of the system.
[0051] In step S240, in response to the triggering instruction of the permission setting save control 204, the permission level configuration, biometric data, and permission range whitelist are saved to the user permission database.
[0052] In step S240, once the user has completed all the necessary permission settings, they can save this information into the user permission database by triggering the permission setting save control 204. This includes the previously configured permission level, the encrypted biometric data, and the detailed permission scope whitelist. The saved information not only provides the basis for subsequent identity verification and permission checks, but also ensures the persistence and consistency of the user's settings. This process ensures that the user's personalized permission settings are accurately maintained even over a long period of use, improving the user experience while also enhancing the overall security of the system.
[0053] In some embodiments of the present application, the advanced permission is the highest level of permission setting, mainly granted to users who need to access sensitive information or perform critical operations, such as the owner or primary user of the vehicle. To ensure the security of these high-risk activities, users who obtain advanced permissions must undergo a strict authentication process - cross-verification of facial and voiceprint. This process requires the user to provide real-time facial images and voiceprint waveforms, which are analyzed and compared by the system's built-in algorithms. Only when both are verified and the correlation is confirmed to be correct can the user obtain advanced permissions. This dual verification mechanism greatly reduces the risk of identity fraud, providing a high level of security. Advanced permissions are suitable for scenarios that require a high level of trust, such as core setting adjustments of the vehicle (e.g., engine parameter adjustments), access to sensitive data (e.g., driving records, personal privacy settings), etc.
[0054] In some embodiments of the present application, the intermediate permission is between the advanced permission and the low-level permission, suitable for users who do not need to access the most sensitive information but still need to perform some important operations. The intermediate permission is authorized by the user with advanced permission (usually the vehicle owner) through the user permission management sub-interface 200 of the intelligent cockpit. The advanced permission user first needs to confirm his identity through strict identity verification, then selects the authorized person and defines the specific operation types and validity period that the authorized person can perform. The intermediate permission user may also need to provide certain biometric data for subsequent identity verification. The intermediate permission allows users to perform functions such as navigation settings, entertainment system control, etc., while ensuring the flexibility and security of the system.
[0055] In some embodiments of the present application, the low-level permission is the most basic permission level, usually assigned to temporary users or visitors, allowing them to perform the most basic operations, such as viewing the current time, weather forecast, etc. non-sensitive information. Low-level permissions may not require any form of verification or only very basic verification (e.g. agreement of terms of use) to minimize the user's usage barrier. This permission design ensures that even unverified users can enjoy a certain level of service while avoiding any potential security risks. Low-level permissions are mainly used for basic functions that do not affect vehicle operation or user privacy, such as querying vehicle status, adjusting seat position, etc., to provide convenience for users while maintaining the security and stability of the system.
[0056] In some embodiments of the present application, with reference to Figure 5 , in step S130, in response to the trigger instruction of the multi-modal interaction instruction setting control 102, the implementation process of displaying the multi-modal interaction instruction setting sub-interface 300 includes but is not limited to the following steps.
[0057] S310, in response to the trigger instruction of the multi-modal interaction instruction setting control 102, a multi-modal interaction instruction template selection window 400 corresponding to the trigger instruction of the multi-modal interaction instruction setting control 102 is displayed.
[0058] Among them, the multi-modal interaction instruction template selection window 400 includes a custom creation control 401.
[0059] In step S310, when the user triggers the multi-modal interaction instruction setting control 102, the system will display a multi-modal interaction instruction template selection window 400. This window provides a variety of preset interaction instruction templates 402 for users to choose from, each template is optimized for a specific use case or operation type. The use template control 403 is set below the interaction instruction template 402, click the use template control 403, use the interaction instruction template 402 to set the multi-modal interaction instruction. In addition, the window also includes a custom creation control 401, allowing users to create new interaction instruction templates according to individual needs. In this way, users can flexibly choose the appropriate interaction mode according to the actual application scenario, ensuring the convenience and adaptability of the system.
[0060] S320, in response to the trigger instruction of the custom creation control 401, the multi-modal interaction instruction setting sub-interface 300 is displayed.
[0061] In step S320, if the user selects the "custom creation control 401", the system will display a multi-modal interaction instruction setting sub-interface 300. This sub-interface provides a highly customized environment for users, allowing them to configure various interaction parameters according to their own preferences and specific needs. For example, users can adjust voice sensitivity, select different gesture action libraries, perform eye tracking calibration, etc. This flexibility not only allows users to create interaction methods that conform to their habits, but also further enhances the system's personalized service capabilities, improving overall user experience satisfaction. Through these detailed setting options, users can precisely define how to interact through voice, gestures, eye movements, and other input methods, ensuring that each instruction accurately conveys the user's intent.
[0062] In some embodiments of the present application, with reference to Figure 4 , the multi-modal interaction instruction setting sub-interface 300 includes a gesture action library selection control 301, an eye tracking calibration control 302, and a multi-modal interaction instruction saving control 303. In step S140, the process of configuring signal priority rules for the interaction mode of the intelligent cockpit through the multi-modal interaction instruction setting sub-interface 300 includes but is not limited to the following steps.
[0063] S410, in response to the trigger instruction of the gesture action library selection control 301, a gesture action library configuration window is displayed to bind preset gesture action libraries for different scenarios.
[0064] Wherein, the gesture action library includes a basic gesture action set and an extended gesture action set, and the extended gesture action set is only enabled in the parking scenario.
[0065] In step S410, when the user triggers the gesture action library selection control 301, the system will display a gesture action library configuration window, where the user can select or bind corresponding gesture action libraries according to different application scenarios. The gesture action library includes a basic gesture action set and an extended gesture action set, and the extended gesture action set is only enabled in the parking and other safety states to avoid safety hazards caused by misoperation during driving. In this way, users can flexibly adjust the available gesture command set according to the current environment, improving the convenience and safety of interaction.
[0066] S420, in response to the trigger instruction of the eye tracking calibration control 302, an eye tracking calibration window is displayed, an eye movement trajectory calibration program is executed, and eye calibration parameters are generated.
[0067] Wherein, the eye calibration parameters are used to correct the focal point coordinate deviation output by the eye tracking sensor.
[0068] In step S420, once the user touches the eye tracking calibration control 302, the system guides the user to enter the eye tracking calibration window. In this window, the user needs to complete a series of eye movement trajectory calibration tasks according to the instructions, and the system generates eye tracking calibration parameters accordingly. These parameters are mainly used to correct the focal point coordinate deviation of the eye tracking sensor output, ensuring the accuracy of the eye tracking data. Correct eye tracking calibration not only improves the user experience, but also effectively improves the reliability and accuracy of the eye tracking-based interaction mode, so that the user can easily control various functions in the intelligent cockpit through eye contact.
[0069] In step S430, in response to the triggering instruction of the multi-modal interaction instruction saving control 303, the gesture action library binding relationship and the eye tracking calibration parameters are saved and input into the multi-modal fusion processing algorithm to generate scene-adaptive signal priority rules.
[0070] In step S430, the user saves all the parameters previously set, including the gesture action library binding relationship and the eye tracking calibration parameters, by triggering the multi-modal interaction instruction saving control 303. These information will be subsequently input into the multi-modal fusion processing algorithm to generate a set of scene-adaptive signal priority rules. This rule can dynamically adjust the weight of each modal input signal according to the specific needs in different scenarios, ensuring that the system can accurately understand and respond to the user's intention whether in driving or parking state. This not only improves the interaction efficiency, but also enhances the flexibility and adaptability of the system, meeting the diversified and personalized user needs.
[0071] In some embodiments of the present application, the signal priority rules include signal quality evaluation rules, scene classification rules, dynamic weight allocation rules, and conflict processing rules.
[0072] In some embodiments of the present application, the signal quality evaluation rules are used to quantitatively evaluate the quality of multi-modal input signals to obtain the quality level of multi-modal input signals. This process is achieved by analyzing the characteristics of each input signal, such as signal-to-noise ratio and clarity for voice signals, accuracy and continuity of actions for gesture signals, and focus positioning accuracy for eye movement signals. By comprehensively evaluating these parameters, the system can assign a quality level to each input signal, which helps to dynamically adjust the importance or weight of the signal in the subsequent steps according to the actual quality of the signal, ensuring that the system can respond based on the most reliable data.
[0073] In some embodiments of the present application, the scene classification rule is used to divide scene types according to the current state of the intelligent cockpit. The intelligent cockpit can be in a variety of different states, such as driving, parking, navigation mode, etc., and the user's needs and interaction methods are different in each state, and each scene type is bound to a corresponding safety limit rule for the principle of safe driving. Through the scene classification rule, the system can accurately determine the specific scene it is currently in and adjust the corresponding interaction strategy accordingly. This scene-based adaptive mechanism enables the intelligent cockpit to provide the most suitable operation interface and service in different situations, not only improving the user experience, but also enhancing the practicality and safety of the system.
[0074] In some embodiments of the present application, the dynamic weight allocation rule is used to dynamically adjust the weight of each modality signal according to the quality level of the current multi-modal input signal and the scene type. This means that the system does not treat all input signals statically, but adjusts the importance of each signal flexibly according to real-time conditions. For example, in a noisy environment, the quality of the voice signal may be low, at which time the system may reduce the weight of the voice signal and increase the weight of the gesture or eye movement signal. Similarly, in a specific scene (such as when parking), certain types of input signals (such as an extended gesture set) may become more important. This approach ensures that regardless of environmental changes, the system can prioritize the most reliable and most suitable input signal for the current scene, thereby improving the accuracy and reliability of the interaction instruction.
[0075] In some embodiments of the present application, the conflict processing rule is used to solve the conflict of multi-modal input signals according to the preset signal priority order and user confirmation mechanism. The preset signal priority order defines which types of signals should be given priority when a conflict occurs, but to further ensure the correctness of the decision, the system also introduces a user confirmation mechanism. For example, if a voice command and a gesture operation are issued at the same time and the intentions are inconsistent, the system first selects one signal as the main basis according to the preset priority order, and then prompts the user to confirm whether to perform the operation or select another input signal. This approach not only effectively avoids errors caused by misrecognition, but also enhances the user's sense of control and trust, ensuring that each interaction is the result the user truly wants.
[0076] In some embodiments of the present application, the conflict handling rules explicitly define the priority order of each modality signal as voice signal, gesture signal, eye movement signal, and facial expression signal. Voice signal is given the highest priority because it is the most direct and explicit way of expressing user intent, especially suitable for conveying complex instructions quickly during driving. Gesture signal comes second, suitable for providing intuitive interaction in situations where voice commands are not appropriate or convenient, such as adjusting volume or switching music. Eye movement signal ranks third, mainly used to capture the user's focus of attention and line of sight, enabling operations without physical contact, optimizing the interaction process in specific scenarios. Facial expression signal has the lowest priority, mainly used to perceive the user's emotional state, enhancing user experience and personalized services, rather than directly controlling vehicle functions. Through this priority-based conflict handling mechanism, the system can make reasonable and consistent responses when multiple modality input signals conflict, ensuring the reliability of operations and the overall satisfaction of user experience. This mechanism not only improves the accuracy and convenience of interaction, but also effectively avoids unnecessary operation interference caused by misidentification.
[0077] In some embodiments of the present application, when the instruction content corresponding to multiple signals does not conflict, the system executes the composite instruction.
[0078] In some embodiments of the present application, when multiple modality signals exist simultaneously and may conflict, the system will process them according to the above priority order. For example, if the user issues a voice command and a gesture action simultaneously, the system will prioritize the voice command; if the voice signal is unclear or invalid, it will turn to processing the gesture signal. Similarly, if gesture and eye movement signals conflict simultaneously, the system will give priority to the gesture signal. This priority-based conflict handling mechanism ensures that even in the case of complex interweaving of multi-modal input signals, the system can make reasonable and consistent responses, improving the reliability of interaction and the overall satisfaction of user experience.
[0079] In some embodiments of the present application, during vehicle driving, a senior user (vehicle owner) issues a "close window" voice command, while a junior user (passenger) draws an "open window" gesture. The system first performs permission verification: the junior user's permission range whitelist does not contain the window control function, and his gesture instruction is invalid due to insufficient permissions; the senior user's voice instruction passes the permission verification. Then, according to the signal priority rule (voice priority higher than gesture priority), the system prioritizes the execution of the legal senior user's voice instruction, and prompts "window closed (voice instruction effective)" in the real-time feedback display area 101. Finally, only the senior user's instruction is executed, and the junior user is intercepted by the system due to insufficient permissions.
[0080] In some embodiments of the present application, while the vehicle is in motion, a senior user (owner) issues a voice command to "open the door," while an intermediate user (passenger) issues a voice command to "close the window." The system categorizes the current scenario type as "vehicle in motion." The senior user's "open the door" command is forcibly intercepted because it triggers the safety restriction rules (dangerous operations are prohibited) for the "vehicle in motion" scenario type, even though its permission level is legal. The system prompts in the real-time feedback display area 101: "Operation prohibited: Doors locked while driving." The intermediate user's "close the window" command passes permission verification (the permission range whitelist includes window control) and there is no security conflict, so it is considered a legal command. The window closing operation is executed normally, and the real-time feedback displays "Window closed (voice command effective)." In short, the system prioritizes the execution of safety rules to intercept dangerous operations, while independently processing non-conflicting commands within legal permissions to ensure a balance between safety and functionality.
[0081] In some embodiments of the present application, while the vehicle is parked, a low-level user (visitor) uses a voice command to "start the engine," while a high-level user uses eye movement to gaze at the "lock vehicle" button. During permission verification, the system discovers that the low-level user's permissions are limited to querying basic information and do not allow the engine start operation, so the voice command is rejected. However, the high-level user's eye movement command passes permission verification. Although the voice signal takes priority over the eye movement signal, because the low-level user's permissions are invalid, the system executes the high-level user's legitimate eye movement command instead, completing the vehicle locking operation. The real-time feedback display area 101 displays a prompt to the low-level user: "Insufficient permissions: Unable to execute engine start," and the vehicle status is updated to locked.
[0082] In some embodiments of the present application, in step S150, the multimodal input signal of the current user is obtained, the multimodal fusion processing algorithm is called, and the implementation process of generating the interaction intention instruction based on the signal priority rule includes but is not limited to the following steps.
[0083] S510 , preprocessing and feature extraction are performed on the multimodal input signal to obtain a multimodal signal feature vector.
[0084] In step S510, the system first preprocesses and extracts features from the user's various input signals (such as speech, gestures, and eye movements). This process includes removing noise, standardizing the data format, and extracting key features from the raw signals, such as intonation in speech, the trajectory of gestures, or the direction and speed of eye movements. Through these operations, the system converts different types of input signals into a unified feature vector representation—a multimodal signal feature vector. This step is crucial because it provides a structured and easy-to-process data foundation for subsequent signal fusion and intent analysis, ensuring that different types of input signals can be effectively integrated.
[0085] S520, based on the dynamic weight distribution rule defined in the signal priority rule, calling a multi-modal fusion processing algorithm to perform fusion processing on the multi-modal signal feature vector to generate a joint feature vector.
[0086] In step S520, the multi-modal fusion processing algorithm is used to fuse the multi-modal signal feature vector according to the previously defined dynamic weight distribution rule. In this step, the system dynamically adjusts the weight of each input signal according to the current environmental conditions and the specific situation of the user (such as signal quality evaluation results and scene classification). This means that in some cases the voice signal may be considered more important than the gesture signal, and in other scenarios the opposite may be true. In this way, the system can consider all available information sources and generate a joint feature vector representing the overall user intent. This context-based adaptive mechanism greatly improves the accuracy and reliability of the interactive instruction generation.
[0087] S530, inputting the joint feature vector into an intent analysis model to generate an intent analysis result, and matching it with structured operation instructions in a preset instruction library to obtain an interactive intent instruction.
[0088] In step S530, the system inputs the generated joint feature vector into the intent analysis model to generate a specific intent analysis result. This model uses machine learning or deep learning techniques to understand complex user intent and convert it into executable operation instructions. Then, the system matches the parsed intent with the structured operation instructions in the preset instruction library to find the operation command that best matches the current intent. This completes the conversion from user input to specific operation instructions, forming a clear and explicit interactive intent instruction. This step not only ensures that the user's intent can be accurately identified and executed, but also makes the entire interaction process more intuitive and efficient, improving the overall satisfaction of the user experience.
[0089] In some embodiments of the present application, the multi-modal fusion processing algorithm includes a Transformer model based on cross-modal attention mechanism, a multi-modal graph neural network, and a reinforcement learning driven dynamic fusion model.
[0090] The Transformer model based on cross-modal attention mechanism is mainly used to capture the complex relationships between different modal signals. This model can weight the information within each modality through self-attention mechanisms, while assessing the mutual influence between different modalities through cross-modal attention mechanisms. This design not only improves the feature extraction capability of each modality signal, but also effectively integrates multiple input signals to generate more comprehensive and accurate joint feature representations. For example, in an intelligent cockpit, voice and gesture signals may convey the user's intent simultaneously, and the cross-modal attention mechanism can automatically identify and integrate these information to ensure that the system's response is more accurate and consistent.
[0091] Multimodal Graph Neural Network (M-GNN) is used to model and analyze the structured relationships between different modal signals. By representing various input signals as graph nodes and utilizing the powerful representation capabilities of graph neural networks, M-GNN can capture local and global dependencies between signals. This method is particularly suitable for handling complex interaction scenarios, such as situations where multiple input sources need to be processed simultaneously during driving. M-GNN constructs a graph structure of multi-modal signals, enabling the system to better understand the intrinsic relationships between signals and generate more interpretable and robust interaction instructions. For example, it can help the system better understand how voice and gesture signals work together to complete specific tasks.
[0092] The reinforcement learning-driven dynamic fusion model focuses on continuously optimizing the fusion strategy based on real-time feedback. This model uses reinforcement learning algorithms to adjust and optimize the fusion weights of multi-modal signals through interaction with the environment, achieving the best interaction effect. In the application scenario of an intelligent cockpit, the system can gradually improve its understanding of user intent and execution efficiency through continuous learning and adjustment. For example, when the system detects a decrease in the quality of a certain modality signal, it can dynamically adjust the weights of other modality signals to ensure the reliability of the overall input signal. This adaptive capability enables the system to maintain high performance in complex and variable real-world environments, providing more personalized and flexible services.
[0093] In some embodiments of the present application, the implementation process of matching the multi-modal input signal with the biometric data of the user permission database in step S160 to obtain the permission level of the current user includes but is not limited to the following steps.
[0094] In some embodiments of the present application, the implementation process of matching the interaction intent instruction with the permission level of the current user in step S160 for verification, if the verification is passed, displaying the instruction execution status in the real-time feedback display area 101, and controlling the execution mechanism of the intelligent cockpit to complete the corresponding operation includes but is not limited to the following steps.
[0095] S610, query the executability of the interactive intent instruction from the permission scope whitelist in the user permission database according to the permission level of the current user.
[0096] In step S610, the system first queries whether the user has the right to execute the interactive intent instruction generated by the multi-modal input signal within the permission scope whitelist in the user permission database according to the permission level of the current user. The permission scope whitelist lists all operation types that the user can perform and their corresponding permission validity period in detail. Through this query process, the system can determine whether the operation requested by the user is within its permission scope. This step ensures that only users with the corresponding permission level can access or operate specific functions, thereby effectively protecting the security and data privacy of the system.
[0097] S620, if the interactive intent instruction is within the permission scope whitelist, the verification is passed, the instruction execution state is displayed in the real-time feedback display area 101, and the execution mechanism of the intelligent cockpit is controlled to complete the corresponding operation.
[0098] When the interactive intent instruction is confirmed to be within the user's permission scope whitelist (i.e., the verification is passed) in step S610, the system will enter step S620. At this time, the system will display the execution state of the instruction in the real-time feedback display area 101, allowing the user to immediately understand the processing progress and results of their request. At the same time, the system will send control commands to the relevant execution mechanism of the intelligent cockpit to complete the operation requested by the user. For example, if the user requests to adjust the temperature in the car, the system will instruct the air conditioning system to make the corresponding adjustment. This way not only improves the transparency of user experience, but also ensures that all operations are efficiently executed under the premise of legal authorization.
[0099] S630, if the interactive intent instruction is not within the permission scope whitelist, the verification is not passed, a permission conflict popup is triggered, and the reason for the permission conflict is highlighted in the real-time feedback display area 101.
[0100] If it is found in step S610 that the user's interactive intent instruction exceeds the allowed operations within the permission scope whitelist (i.e., the verification is not passed), the system will execute step S630. At this time, the system will trigger a permission conflict popup to clearly show the user the reason for the permission conflict, and highlight the relevant instructions in the real-time feedback display area 101. This design helps users quickly understand why their request cannot be executed, and provides clear information to guide them on how to solve the problem (such as contacting a user with higher permissions for authorization). This not only enhances the security of the system, but also improves the user's understanding and acceptance of the system's rules, reducing confusion and dissatisfaction caused by permission issues.
[0101] Secondly, with reference to Figure 6The embodiment of the application provides a multi-modal interaction system based on an intelligent cockpit, which comprises a multi-modal interaction main interface module 710, a user permission management module 720, a multi-modal interaction instruction setting module 730, an interaction intention instruction generation module 740 and an interaction intention instruction execution module 750.
[0102] The multi-modal interaction main interface module 710 is used for displaying a multi-modal interaction main interface of the intelligent cockpit. The multi-modal interaction main interface comprises a multi-modal interaction instruction setting control, a user permission management control and a real-time feedback display area.
[0103] The multi-modal interaction main interface module 710 is responsible for displaying the multi-modal interaction main interface of the intelligent cockpit, which is the main entrance for user interaction with the system. The interface integrates the multi-modal interaction instruction setting control, the user permission management control and the real-time feedback display area, so that the user can conveniently adjust the interaction mode, manage personal permissions and instantly check the response status of the system. By providing an intuitive and easy-to-operate interface, the module not only improves the user experience, but also ensures the transparency and timeliness of information transmission, so that the user can always master the running status of the system.
[0104] The user permission management module 720 is used for displaying a user permission management sub-interface in response to a trigger instruction of the user permission management control, establishing a user permission database, collecting and encrypting biological feature data of the user and cross-verifying and binding the biological feature data with a configured permission level.
[0105] The user permission management module 720 is used for responding to the permission setting needs of the user. When the user triggers the user permission management control, it displays the user permission management sub-interface. This module allows the user to establish and manage the user permission database, including collecting and encrypting the biological feature data (such as facial recognition or fingerprint) of the user, and cross-verifying and binding these data with the configured permission level. In this way, the module ensures that only the verified user can obtain the corresponding permission, thereby enhancing the security and privacy protection capability of the system.
[0106] The multi-modal interaction instruction setting module 730 is used for displaying a multi-modal interaction instruction setting sub-interface in response to a trigger instruction of the multi-modal interaction instruction setting control; through the multi-modal interaction instruction setting sub-interface, signal priority rules are configured for the interaction mode of the intelligent cockpit; the signal priority rules are used for adjusting the fusion weight of multi-modal input signals; the multi-modal input signals comprise voice signals, gesture signals, eye movement signals and facial expression signals.
[0107] The multi-modal interaction instruction setting module 730 is responsible for responding to the user's interaction instruction setting needs. When the user triggers the multi-modal interaction instruction setting control, it will display the multi-modal interaction instruction setting sub-interface, which defines how to dynamically adjust the fusion weights of multiple types of input signals (such as voice, gestures, eye movements, and facial expressions) when they are received. Through this customized setting, the module improves the flexibility and accuracy of the interaction, ensuring optimal user experience in different scenarios.
[0108] The interaction intent instruction generation module 740 is used to obtain the multi-modal input signals of the current user, call the multi-modal fusion processing algorithm, and generate the interaction intent instruction based on the signal priority rules.
[0109] The core task of the interaction intent instruction generation module 740 is to process the multi-modal input signals of the user and generate the interaction intent instruction based on the pre-set signal priority rules. The module first pre-processes and extracts features from the input signals, then calls the multi-modal fusion processing algorithm to fuse these feature vectors, and finally generates the instruction representing the user's intent. This process ensures that even in complex environments, the system can accurately understand and execute the user's actual intent, greatly improving the efficiency and accuracy of the interaction.
[0110] The interaction intent instruction execution module 750 is used to match the multi-modal input signals with the biometric data of the user permission database to obtain the permission level of the current user, match and verify the interaction intent instruction with the permission level of the current user, and if the verification is passed, display the instruction execution status in the real-time feedback display area and control the execution mechanism of the intelligent cockpit to complete the corresponding operation.
[0111] The interaction intent instruction execution module 750 is responsible for verifying the user's permission and executing the interaction intent instruction. First, it matches the user's multi-modal input signals with the biometric data in the user permission database to determine the user's permission level. Then, the module matches and verifies the interaction intent instruction with the permission level of the current user. If the verification is passed, the system will display the instruction execution status in the real-time feedback display area and control the relevant execution mechanism of the intelligent cockpit to complete the corresponding operation; if the verification is not passed, it will trigger the corresponding permission conflict prompt. This step ensures that all operations are efficiently executed under the premise of legal authorization, while providing a clear feedback mechanism to enhance user experience.
[0112] Furthermore, with reference to Figure 7 The embodiments of the present application provide a multi-modal interaction device based on an intelligent cockpit, which comprises a display device 810, an intelligent cockpit hardware device 820, a central processing unit 830, and an execution mechanism 840.
[0113] The display device 810 is used to display the multi-modal interaction main interface of the intelligent cockpit. The multi-modal interaction main interface includes multi-modal interaction instruction setting controls, user permission management controls, and real-time feedback display areas. In response to a trigger instruction for the user permission management controls, a user permission management sub-interface is displayed, a user permission database is established, and user biometric data is collected and encrypted, and the biometric data is cross-verified and bound with the configured permission level. In response to a trigger instruction for the multi-modal interaction instruction setting controls, a multi-modal interaction instruction setting sub-interface is displayed. Through the multi-modal interaction instruction setting sub-interface, signal priority rules are configured for the interaction mode of the intelligent cockpit. The signal priority rules are used to adjust the fusion weight of the multi-modal input signals. The multi-modal input signals include voice signals, gesture signals, eye movement signals, and facial expression signals.
[0114] The display device 810 is used to display the multi-modal interaction main interface of the intelligent cockpit, which integrates multi-modal interaction instruction setting controls, user permission management controls, and real-time feedback display areas. These controls allow users to easily adjust the interaction mode, manage personal permissions, and instantly view the response status of the system. In addition, when the user triggers the user permission management controls, the display device 810 presents the user permission management sub-interface, supports the establishment of a user permission database and the encrypted storage of user biometric data, and cross-verification and binding of these data with the configured permission level. Similarly, after triggering the multi-modal interaction instruction setting controls, the display device 810 presents the multi-modal interaction instruction setting sub-interface, allowing users to configure signal priority rules for different interaction modes, thereby optimizing the fusion weight of input signals.
[0115] The intelligent cockpit hardware device 820 is used to obtain multi-modal input signals of the current user. The intelligent cockpit hardware device includes a voice recognition device 821, a gesture recognition device 822, an eye movement tracking device 823, and a facial expression recognition device 824.
[0116] The intelligent cockpit hardware device 820 is responsible for obtaining multi-modal input signals of the current user, including voice signals, gesture signals, eye movement signals, and facial expression signals. The device is composed of various specialized devices, such as voice recognition devices, gesture recognition devices, eye movement tracking devices, and facial expression recognition devices. Through these devices, the system can comprehensively capture various input information of the user, ensuring that no matter which interaction method the user uses, it can be accurately recognized and processed. This multi-dimensional data collection not only improves the flexibility of interaction, but also enhances the realism and immersion of user experience.
[0117] In some embodiments of the present application, the voice recognition device employs a high-sensitivity microphone array that can accurately capture the user's voice commands. This design not only improves the quality of voice signal collection but also effectively filters background noise in noisy environments, ensuring the clarity and accuracy of voice commands. By using a microphone array, the system can achieve sound source positioning, further improving the accuracy of voice recognition. This is particularly important for intelligent cockpit, as it needs to accurately understand the voice commands of the driver or passenger in various complex driving environments, providing a safer and more convenient operation experience.
[0118] In some embodiments of the present application, the gesture recognition device utilizes a deep learning-based gesture recognition algorithm and an infrared camera to capture and recognize user gestures in real time. The device determines the user's specific intent by analyzing features such as the shape and motion trajectory of the gesture. The use of infrared cameras allows accurate gesture recognition even in low-light conditions, improving the stability and reliability of the system. In addition, the deep learning algorithm continuously optimizes the accuracy of gesture recognition, enabling efficient performance in complex and variable real-world application scenarios. Gesture recognition devices provide an intuitive and non-physical contact interaction method for users, enhancing the user experience of intelligent cockpit.
[0119] In some embodiments of the present application, the eye tracking device employs high-precision eye tracking sensors and advanced algorithms to accurately monitor the user's eye movement trajectory. These sensors are usually installed near the instrument panel or display screen to minimize the impact on the user's line of sight. By accurately tracking eye movements, the system can infer the user's gaze point and attention direction, enabling more natural human-computer interaction. For example, in an intelligent cockpit, eye tracking technology can be used to control the infotainment system or navigation interface, allowing users to complete operations simply by looking at a specific area with their eyes, greatly improving the convenience and safety of interaction.
[0120] In some embodiments of the present application, the facial expression recognition device uses computer vision-based facial expression recognition algorithms and high-definition cameras to capture and recognize user facial expressions in real time. High-definition cameras provide high-quality image input, while computer vision algorithms analyze changes in facial feature points such as eyebrow lifting and mouth bending to determine the user's emotional state (e.g., happy, sad, surprised, etc.). This function not only enhances the emotional interaction of human-computer interaction but also automatically adjusts the cockpit environment settings (such as music playlists, light brightness, etc.) according to the user's emotional state, creating a more comfortable and harmonious driving experience. The application of facial expression recognition devices enables intelligent cockpit to not only understand users' language and actions but also perceive their emotional needs, further improving the level of personalized service.
[0121] The central processing unit 830 is used to call a multi-modal fusion processing algorithm according to the multi-modal input signal, generate an interactive intent instruction based on a signal priority rule; match the multi-modal input signal with the biometric data of the user permission database to obtain the permission level of the current user; match and verify the interactive intent instruction with the permission level of the current user, and if the verification is passed, display the instruction execution state in the real-time feedback display area and control the execution mechanism 840 to complete the corresponding operation.
[0122] The central processing unit 830 is the core of the entire system, responsible for processing multi-modal input signals received from the intelligent cockpit hardware device 820. First, it calls a multi-modal fusion processing algorithm to generate an interactive intent instruction based on a pre-set signal priority rule. Next, the central processing unit 830 matches the multi-modal input signal with the biometric data in the user permission database to determine the permission level of the current user. Then, it matches and verifies the generated interactive intent instruction with the user's permission level. If the verification is passed, the central processing unit 830 displays the instruction execution state in the real-time feedback display area and sends control commands to the execution mechanism 840 to complete the corresponding operation. This process ensures that all operations are efficiently executed under the premise of legal authorization, while also providing a clear feedback mechanism to enhance user experience.
[0123] The execution mechanism 840 executes specific operational tasks according to the instructions of the central processing unit 830. Once the central processing unit 830 confirms that the user's interactive intent instruction matches the current user's permission level and the verification is passed, the execution mechanism 840 will take action according to the instruction content. For example, if the user requests to adjust the temperature in the car or start the navigation system, the execution mechanism 840 will directly control the relevant vehicle-mounted equipment (such as the air conditioning system or navigation instrument) to complete the corresponding operation. This step ensures that the user's intent can be quickly and accurately realized, greatly improving the convenience and satisfaction of user experience. In this way, the execution mechanism 840 ensures the safety and efficiency of operations while providing a seamless human-machine interaction experience for users.
[0124] In addition, the embodiment of the present application provides a vehicle integrated with the aforementioned multi-modal interaction system based on an intelligent cockpit.
[0125] It should be noted that in various specific embodiments of the present application, when relevant processing needs to be performed on data related to the identity or characteristics of the user, such as user information, user behavior data, user history data, and user location information, the user's permission or consent will be obtained first, and the collection, use, and processing of such data will comply with relevant laws, regulations, and standards in the relevant countries and regions. In addition, when the embodiments of the present application need to obtain sensitive personal information of the user, the separate permission or separate consent of the user will be obtained through a pop-up window or a jump to a confirmation page, and after obtaining the separate permission or separate consent of the user, the necessary user-related data for the normal operation of the embodiments of the present application will be obtained.
[0126] In summary, the multi-modal interaction method, system, device and vehicle based on the intelligent cockpit provided by the embodiments of the present application have the following technical effects.
[0127] The multi-modal interaction method, system, device and vehicle based on the intelligent cockpit provided by the embodiments of the present application effectively enhance the user experience by integrating multiple input methods such as voice, gesture, eye movement and facial expression. Users can choose the most suitable interaction method according to the specific situation, whether it is to quickly convey complex instructions during driving or to perform detailed operations in a parked state, and can provide intuitive and convenient interaction experience. High-precision signal processing and fusion technology ensures efficient processing and accurate fusion of multi-modal input signals, improving the response speed and accuracy of the system.
[0128] In addition, the embodiments of the present application have a powerful permission management and security protection mechanism. By establishing a user permission database and cross-verifying and binding with biometric data, the permission levels of different users can be effectively managed to ensure that only authorized users can perform specific operations. The pre-set signal priority order and priority-based conflict processing rules enable the system to make reasonable and consistent decisions when multi-modal input signals conflict, avoiding operation interference caused by misidentification. The design of the real-time feedback display area allows users to immediately understand the response progress and results of the system, enhancing the user's sense of control and trust.
[0129] Finally, the embodiments of the present application also provide flexible personalized configuration functions, and users can adjust the priority rules of various input methods according to actual needs, and dynamically adjust parameters according to personal preferences or application scenarios. The application of facial expression recognition devices not only enhances the emotional interactivity of human-computer interaction, but also automatically adjusts the cockpit environment settings according to the emotional state of the user, further improving the level of personalized services.
[0130] In some alternative embodiments, the function / operations described in the block diagrams can not occur in the order described in the operational illustrations. For example, two blocks shown in succession can in fact be executed substantially concurrently or the blocks can sometimes be executed in the reverse order, depending upon the functionality / operations involved. Also, although the embodiments presented in the flow diagrams are shown as a sequence of operations, it is to be understood that the logical flow is merely illustrative of alternative embodiments. The disclosed methods are not limited to the order of operations presented herein. Alternative embodiments are contemplated in which the order of operations is changed, and in which sub-operations are performed in different orders or in parallel.
[0131] Further, although the present application has been described in the context of functional modules, it is to be understood that one or more of the functions and / or features can be integrated in a single physical device and / or software module, or one or more functions and / or features can be implemented in separate physical devices or software modules. It is also to be understood that detailed discussion of the actual implementation of each module is unnecessary to an understanding of the present application. Rather, the actual implementation is within the routine of an engineer's knowledge given the property, functionality and internal relationships of the various functional modules disclosed in the devices shown herein. Accordingly, the present application is not limited to purely hardware implementations, but also encompasses software implementations and / or firmware implementations. It is also to be understood that the disclosed specific concepts are merely illustrative and not intended to limit the scope of the present application, which is defined by the appended claims and their equivalents.
[0132] If the functions are implemented in the form of software function units and sold or used as independent products, they can be stored in a computer readable storage medium. Based on such understanding, the technical solutions of the present application or the parts of the technical solutions that essentially contribute to the prior art can be embodied in the form of a software product. The computer software product is stored in a storage medium and includes a plurality of programs for causing a computer device (which can be a personal computer, a server, or a network device, etc.) to execute all or part of the steps of the embodiments of the present application. The aforementioned storage medium includes: a U disk, a mobile hard disk, a read-only memory (ROM, Read-Only Memory), a random access memory (RAM, Random Access Memory), a magnetic disk or an optical disk, and various media that can store program codes.
[0133] The logic and / or steps represented in the flow diagrams or otherwise described herein, for example, can be considered as a sequence of executable program instructions for implementing logic functions, and can be embodied in any computer-readable medium for use by or in connection with an instruction execution system, apparatus, or device, such as a computer-based system, processor- containing system, or other system that can fetch the instructions from the instruction execution system, apparatus, or device and execute the instructions. For the purposes of this specification, a "computer-readable medium" can be any apparatus that can contain, store, communicate, propagate, or transport the program for use by or in connection with the instruction execution system, apparatus, or device. The computer-readable medium can be, for example but not limited to, an electronic, magnetic, optical, electromagnetic, infrared, or semiconductor system, apparatus, or device. More specific examples (a non-exhaustive list) of the computer-readable medium include the following: an electrical connection (electronic) having one or more wires, a portable computer diskette (magnetic), a random access memory (RAM), a read-only memory (ROM), an erasable programmable read-only memory (EPROM or Flash memory), an optical fiber, and a portable compact disc read-only memory (CDROM). In addition, the computer-readable medium can even be paper or another suitable medium upon which the program is printed, as the program can be electronically captured, for example via an optical scanner, then compiled, interpreted, or otherwise processed in a suitable manner, if necessary, and stored in a computer memory.
[0134] More specific examples (a non-exhaustive list) of the computer-readable medium include the following: an electrical connection (electronic) having one or more wires, a portable computer diskette (magnetic), a random access memory (RAM), a read-only memory (ROM), an erasable programmable read-only memory (EPROM or Flash memory), an optical fiber, and a portable compact disc read-only memory (CDROM). In addition, the computer-readable medium can even be paper or another suitable medium upon which the program is printed, as the program can be electronically captured, for example via an optical scanner, then compiled, interpreted, or otherwise processed in a suitable manner, if necessary, and stored in a computer memory.
[0135] It should be understood that aspects of the application can be implemented in hardware, software, firmware, or combinations thereof. In the embodiments described above, various steps or methods can be implemented, for example, using software or firmware stored in memory and executed by a suitable instruction execution system. For example, if implemented in hardware, and in another embodiment, any of the following can be used: discrete logic circuitry, application specific integrated circuits (ASICs), programmable gate arrays (PGAs), field programmable gate arrays (FPGAs), and / or the like.
[0136] In the above-description of various embodiments of the present application, it is to be understood that the terminology used herein is for the purpose of describing particular embodiments only and is not intended to be limiting. It will be understood that the use of "one embodiment," "another embodiment," or "some embodiments" in the above description indicates that the feature(s) being described in the embodiment is included in at least one embodiment of the present application. Thus, the use of "one embodiment," "another embodiment," or "some embodiments" in the above description does not necessarily indicate that the same embodiment(s) is included in all embodiments of the present application. In addition, it will be understood that the use of "one embodiment," "another embodiment," or "some embodiments" in the above description indicates that the feature(s) being described in the embodiment is included in at least one embodiment of the present application. Thus, the use of "one embodiment," "another embodiment," or "some embodiments" in the above description does not necessarily indicate that the same embodiment(s) is included in all embodiments of the present application.
[0137] While the embodiments of the application have been shown and described, it is to be understood that the embodiments described are merely exemplary and are not to be taken as limiting the scope of the application. The scope of the application is defined by the claims and their equivalents.
[0138] The above is a specific description of the preferred embodiment of the present application, but the present application is not limited to the embodiment, and those skilled in the art can make various equivalent modifications or replacements without departing from the spirit of the present application, and these equivalent modifications or replacements are all included in the scope defined by the claims of the present application.
Claims
1. A multimodal interaction method based on a smart cockpit, characterized in that: The steps include: Displaying the multimodal interaction main interface of the smart cockpit; wherein the multimodal interaction main interface includes a multimodal interaction instruction setting control, a user authority management control and a real-time feedback display area; In response to a trigger instruction of the user authority management control, display a user authority management sub-interface, establish a user authority database to collect and encrypt and store the user's biometric data, and cross-verify and bind the biometric data with the configured authority level; In response to a trigger instruction for the multimodal interaction instruction setting control, a multimodal interaction instruction setting sub-interface is displayed; through the multimodal interaction instruction setting sub-interface, a signal priority rule is configured for the interaction mode of the smart cockpit; the signal priority rule is used to adjust the fusion weight of the multimodal input signal; the multimodal input signal includes a voice signal, a gesture signal, an eye movement signal, and a facial expression signal; Obtaining the multimodal input signal of the current user, calling the multimodal fusion processing algorithm, and generating an interaction intention instruction based on the signal priority rule; Matching the multimodal input signal with the biometric data in the user authority database to obtain the authority level of the current user; The interaction intention instruction is matched and verified with the current user's authority level. If the verification is successful, the instruction execution status is displayed in the real-time feedback display area, and the actuator of the smart cockpit is controlled to complete the corresponding operation.
2. The multimodal interaction method based on the smart cockpit according to claim 1 is characterized in that: The user authority management sub-interface includes an authority level configuration control, a biometric verification control, an authority range configuration control, and an authority setting save control; The method of displaying a user rights management sub-interface and establishing a user rights database in response to a trigger instruction of the user rights management control includes: In response to a trigger instruction for the permission level configuration control, a permission level configuration window is displayed to set the user permission level; wherein the permission levels include high-level permission, medium-level permission and low-level permission; the high-level permission must be bound by cross-verification of face and voiceprint; In response to a trigger instruction for the biometric verification control, a biometric entry window is displayed, and the user's biometric data is collected and encrypted for storage; wherein the biometric data includes a face image, a voiceprint waveform, and liveness detection parameters; In response to a trigger instruction for the permission range configuration control, a permission range setting window is displayed to configure a permission range whitelist of operable functions for a specified user; the permission range whitelist includes the types of operations allowed to be performed and the validity period of the operation permissions; In response to a trigger instruction for saving the permission setting control, the permission level configuration, biometric data and permission range whitelist are saved to the user permission database.
3. The multimodal interaction method based on the smart cockpit according to claim 1, characterized in that: The displaying of the multimodal interaction instruction setting sub-interface in response to the triggering instruction of the multimodal interaction instruction setting control includes: In response to a trigger instruction of the multimodal interaction instruction setting control, displaying a multimodal interaction instruction template selection window corresponding to the trigger instruction of the multimodal interaction instruction setting control; wherein the multimodal interaction instruction template selection window includes a custom creation control; In response to a trigger instruction for the custom creation control, a multimodal interaction instruction setting sub-interface is displayed.
4. The multimodal interaction method based on the smart cockpit according to claim 1, characterized in that: The multimodal interaction instruction setting sub-interface includes a gesture action library selection control, an eye tracking calibration control, and a multimodal interaction instruction saving control; The step of setting a sub-interface through the multimodal interaction instruction and configuring a signal priority rule for the interaction mode of the smart cockpit includes: In response to a trigger instruction on the gesture action library selection control, a gesture action library configuration window is displayed to bind preset gesture action libraries for different scenarios; wherein the gesture action library includes a basic gesture action set and an extended gesture action set, and the extended gesture action set is only enabled in the parking scenario; In response to a trigger instruction for the eye tracking calibration control, an eye tracking calibration window is displayed, an eye movement trajectory calibration procedure is executed, and eye movement calibration parameters are generated; the eye movement calibration parameters are used to correct the focus coordinate deviation output by the eye tracking sensor; In response to a trigger instruction for saving a control for the multimodal interaction instruction, the gesture action library binding relationship and the eye movement calibration parameters are saved and input into a multimodal fusion processing algorithm to generate a scene-adaptive signal priority rule.
5. The multimodal interaction method based on the smart cockpit according to claim 4 is characterized in that: The signal priority rules include signal quality assessment rules, scene classification rules, dynamic weight allocation rules and conflict handling rules; The signal quality assessment rule is used to quantitatively assess the quality of the multimodal input signal to obtain a quality level of the multimodal input signal; The scene classification rule is used to classify scene types according to the current state of the smart cockpit; The dynamic weight allocation rule is used to dynamically adjust the weight of each modal signal according to the quality level and scene type of the current multimodal input signal; The conflict handling rule is used to resolve the conflict of the multimodal input signals according to a preset signal priority order.
6. The multimodal interaction method based on the smart cockpit according to claim 1, characterized in that: The acquiring of the multimodal input signal of the current user, calling the multimodal fusion processing algorithm, and generating the interaction intention instruction based on the signal priority rule includes: Preprocessing and feature extraction are performed on the multimodal input signal to obtain a multimodal signal feature vector; Based on the dynamic weight allocation rule defined in the signal priority rule, calling the multimodal fusion processing algorithm, fusing the multimodal signal feature vectors, and generating a joint feature vector; The joint feature vector is input into the intention parsing model to generate an intention parsing result, which is then matched with the structured operation instructions in the preset instruction library to obtain the interaction intention instruction.
7. The multimodal interaction method based on the smart cockpit according to claim 1, characterized in that: The matching and verification of the interaction intention command with the current user's authority level is performed. If the verification is successful, the command execution status is displayed in the real-time feedback display area, and the actuator of the smart cockpit is controlled to complete the corresponding operation, including: According to the authority level of the current user, query the executability of the interaction intention instruction from the authority range whitelist in the user authority database; If the interaction intention command is within the permission range whitelist, the verification is passed, the command execution status is displayed in the real-time feedback display area, and the actuator of the smart cockpit is controlled to complete the corresponding operation; If the interaction intention instruction is not in the permission range whitelist, the verification fails, triggering a permission conflict pop-up window, and highlighting the reason for the permission conflict in the real-time feedback display area.
8. A multimodal interactive system based on a smart cockpit, characterized in that: The multimodal interaction system is used to implement the multimodal interaction method based on the smart cockpit as described in any one of claims 1 to 7.
9. A multimodal interactive device based on a smart cockpit, characterized in that: Includes smart cockpit hardware devices, central processing unit, display equipment and actuators; The display device is used to display a multimodal interaction main interface of the smart cockpit; wherein the multimodal interaction main interface includes a multimodal interaction instruction setting control, a user authority management control, and a real-time feedback display area; in response to a trigger instruction of the user authority management control, a user authority management sub-interface is displayed, a user authority database is established to collect and encrypt and store the user's biometric data, and the biometric data is cross-validated and bound to the configured authority level; in response to a trigger instruction of the multimodal interaction instruction setting control, a multimodal interaction instruction setting sub-interface is displayed; through the multimodal interaction instruction setting sub-interface, a signal priority rule is configured for the interaction mode of the smart cockpit; the signal priority rule is used to adjust the fusion weight of the multimodal input signal; the multimodal input signal includes a voice signal, a gesture signal, an eye movement signal, and a facial expression signal; The smart cockpit hardware device is used to obtain the multimodal input signal of the current user; the smart cockpit hardware device includes a voice recognition device, a gesture recognition device, an eye tracking device and a facial expression recognition device; The central processing unit is used to call the multimodal fusion processing algorithm according to the multimodal input signal, and generate an interaction intention instruction based on the signal priority rule; match the multimodal input signal with the biometric data of the user authority database to obtain the current user's authority level; match and verify the interaction intention instruction with the current user's authority level. If the verification is successful, the instruction execution status is displayed in the real-time feedback display area, and the execution mechanism is controlled to complete the corresponding operation.
10. A vehicle, characterized in that: The multimodal interaction system based on the smart cockpit as claimed in claim 8 is integrated.
Citation Information
Cited By
Interaction control method of intelligent glasses
CN121277365A
Instruction monitoring method and device, equipment, storage medium and program product
CN121387679A