Automobile intelligent cabin child safety monitoring system based on multi-mode interaction

By using deep learning algorithms with multimodal sensor arrays and edge processing layers, the limitations of existing automotive intelligent cockpit systems in recognizing children's behavior using a single modality and the privacy risks have been solved, achieving accurate recognition and security in multiple scenarios.

CN121893897APending Publication Date: 2026-04-21CHERY AUTOMOBILE CO LTD
View PDF 0 Cites 0 Cited by

Patent Information

Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
CHERY AUTOMOBILE CO LTD
Filing Date
2026-01-06
Publication Date
2026-04-21

AI Technical Summary

Technical Problem

Existing automotive intelligent cockpit systems have limitations in recognizing children's behavior due to their single-modality nature, especially in low-light or obstructed environments. They cannot accurately distinguish between the physiological characteristics of children and adults, and lack linkage with vehicle control, posing privacy risks.

Method used

Multi-modal sensor arrays (such as millimeter-wave radar, array microphones, TOF cameras, and seat pressure sensors) are used to collect multi-dimensional data. Combined with deep learning algorithms and privacy protection units in the edge processing layer, child feature recognition and safety policy output are achieved, driving the vehicle to perform active intervention.

Benefits of technology

It enables accurate identification and safety assurance of children's behavior in multiple scenarios, ensures data privacy, and can quickly execute vehicle intervention actions in emergency situations, thereby improving the system's security and privacy protection.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN121893897A_ABST
    Figure CN121893897A_ABST
Patent Text Reader

Abstract

The invention belongs to the technical field of new energy automobile intelligent cabins, and particularly relates to an automobile intelligent cabin child safety monitoring system based on multi-mode interaction. The invention discloses an automobile intelligent cabin child safety monitoring system based on multi-modal interaction, and the system comprises a data collection layer which collects and processes the multi-modal information of an automobile intelligent cabin user through a multi-modal sensor array; the edge processing layer is connected with the data acquisition layer, identifies child features in the multi-modal information, carries out local encryption on the child features, carries out fusion calculation on the multi-modal information and outputs a security policy instruction; the execution control layer is connected with the edge processing layer and drives the vehicle to execute an active intervention behavior according to the security policy instruction; the technical defects that an existing system cannot accurately recognize specific behaviors of children due to the limitation of a single mode, and a pure vision scheme fails in a weak light or shielding scene are overcome.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] This invention belongs to the field of intelligent cockpit technology for new energy vehicles, specifically a child safety monitoring system for intelligent cockpits based on multimodal interaction. Background Technology

[0002] With the popularization of new energy vehicles, through multimodal interaction and scenario-based services, intelligent cockpits have shifted from "passive response" to "proactive service," providing users with a safer, more efficient, and personalized driving experience. This has become a higher requirement for users' future automotive intelligent cockpit experience.

[0003] The limitations of existing systems, being limited by their single modality, prevent them from accurately identifying children's specific behaviors. Purely visual solutions fail in low-light or occluded environments. Furthermore, the systems exhibit poor scene adaptability, failing to effectively distinguish between children's and adults' physiological characteristics (such as differences in heart rate / voiceprint), and lack linkage with vehicle control during emergency interventions (e.g., automatic window locking). Additionally, the existing systems pose privacy risks in the collection and processing of the aforementioned data; children's facial data is only processed locally, failing to meet privacy protection requirements.

[0004] Therefore, there is an urgent need to design a child safety monitoring system for automotive intelligent cockpits based on multimodal interaction to address the technical deficiencies in existing solutions. Summary of the Invention

[0005] The purpose of this invention is to propose a child safety monitoring system that integrates multimodal data such as visual, acoustic, and physiological signals, compared with existing technical solutions, and is particularly suitable for monitoring and intervening in the behavior of children in vehicles.

[0006] To achieve the above objectives, the present invention provides the following technical solution: A child safety monitoring system for an automotive intelligent cockpit based on multimodal interaction, comprising: The data acquisition layer collects and processes multimodal information from users of the intelligent cockpit of the car through a multimodal sensor array; The edge processing layer connects to the data acquisition layer, identifies child features in multimodal information, performs local encryption on child features, and simultaneously performs fusion calculations on multimodal information to output security policy instructions; The execution control layer connects to the edge processing layer and drives the vehicle to perform proactive intervention actions according to safety policy instructions.

[0007] The above technical solution produces the following technical effects: This application utilizes a multimodal sensor array to enable the system to collect user information from within the vehicle's smart cockpit from all angles and perspectives, including but not limited to visual images, acoustic signals, and physiological signals, thereby ensuring the comprehensiveness and accuracy of the data. Secondly, the edge processing layer of this application processes the collected multimodal information to accurately identify child characteristics. Simultaneously, encryption is performed locally during the identification process, protecting user privacy and ensuring data security. Furthermore, this application can also perform fusion calculations on the multimodal information to output targeted security policy instructions.

[0008] Finally, the execution control layer, based on the safety policy instructions, quickly drives the vehicle to perform corresponding active intervention behaviors, such as automatic window locking and emergency braking, thereby effectively ensuring the safety of children inside the vehicle.

[0009] As a further improvement to the multimodal interaction-based child safety monitoring system for automotive intelligent cockpits of this application, the multimodal sensor array includes: Millimeter-wave radar collects the range of body movements of users in a smart car cockpit; An array of microphones is used to identify the voiceprints of users in a smart car cockpit. TOF camera for 3D skeletal tracking of users in smart car cockpits; Seat pressure sensors detect the sitting posture of users in a car's smart cockpit.

[0010] As a further improvement to the child safety monitoring system for a car intelligent cockpit based on multimodal interaction proposed in this application, the edge processing layer includes a child feature recognition engine. The child feature recognition engine is based on a deep learning algorithm to fuse and analyze multimodal information, identify and output child features.

[0011] As a further improvement to the multimodal interaction-based child safety monitoring system for automotive intelligent cockpits of this application, the TOF camera has a horizontal viewing angle of 120° and a vertical viewing angle of 90°.

[0012] As a further improvement to the child safety monitoring system for a car intelligent cockpit based on multimodal interaction proposed in this application, the edge processing layer is embedded with a safety policy library. The safety policy library is preset with risk levels and penalty conditions. When the child's characteristics meet the risk level and penalty conditions, a safety policy instruction is output.

[0013] As a further improvement to the multimodal interaction-based child safety monitoring system for automotive intelligent cockpits proposed in this application, a privacy desensitization module is provided in the edge processing layer; The privacy desensitization module desensitizes sensitive information in security policy instructions and outputs the desensitized security policy instructions.

[0014] As a further improvement to the child safety monitoring system for a car intelligent cockpit based on multimodal interaction proposed in this application, the edge processing layer includes a privacy protection unit. The privacy protection unit performs local encryption on sensitive data in the child's characteristics, and the local encryption method is AES-256 encryption.

[0015] As a further improvement to the multimodal interaction-based intelligent cockpit child safety monitoring system of this application, the execution control layer includes a vehicle domain controller and an execution device; The body domain controller is connected to the edge processing layer via the CAN-FD bus, and outputs safety policy commands to the body domain controller. The vehicle domain controller drives the actuators to perform proactive intervention actions based on safety policy instructions.

[0016] As a further improvement to the multimodal interaction-based child safety monitoring system for automotive intelligent cockpits of this application, the actuators include a window controller, an aromatherapy system, and a voice prompt module; The window controller is used to control the operation of the vehicle's window lift motors. Fragrance system, used to control the opening and closing of the solenoid valve of the fragrance release module in the vehicle; The voice prompt module outputs specified voice prompts according to security policy instructions.

[0017] As a further improvement to the multimodal interaction-based child safety monitoring system for automotive intelligent cockpits of this application, the actuator also includes an entertainment center console, which, after receiving safety policy instructions, calls up media resources stored locally in the vehicle. Attached Figure Description

[0018] Figure 1 This is a system architecture diagram of the present invention; Figure 2 This is a flowchart of the system security intervention process of the present invention; in: 1-Data Acquisition Layer; 2-Edge processing layer; 3-Execution control layer. Detailed Implementation

[0019] The technical solutions of the embodiments of the present invention will be clearly and completely described below with reference to the accompanying drawings. Obviously, the described embodiments are only some embodiments of the present invention, and not all embodiments. Based on the embodiments of the present invention, all other embodiments obtained by those skilled in the art without creative effort are within the scope of protection of the present invention.

[0020] To facilitate an accurate understanding of the solutions provided in the following embodiments of the present invention, the terms involved in the present invention are explained as follows before describing the technical solutions provided by the present invention: Multimodal sensor array: A hardware unit that collects physiological and behavioral data of children in the vehicle through the collaborative collection of multiple sensors, including millimeter-wave radar, array microphone, TOF camera, seat pressure sensor, etc., to achieve multi-dimensional data fusion.

[0021] Child Feature Recognition Engine: A core processing unit that accurately identifies children's physiological features (voiceprint, skeleton, heart rate) and behavioral intentions based on multimodal data fusion algorithms and liveness detection technology.

[0022] Safety strategy library: A decision-making system that pre-sets proactive protection strategies based on risk levels, including automated intervention schemes such as vehicle control linkage (e.g., window locking) and environmental regulation (e.g., fragrance release).

[0023] Privacy Protection Unit: A module that enables secure processing of children's sensitive data (facial and voice information) through local data encryption, edge computing, and differential privacy technology.

[0024] The following detailed description is exemplary and intended to provide further detailed explanation of the invention. Unless otherwise specified, all technical terms used in this invention have the same meaning as commonly understood by one of ordinary skill in the art. The terminology used in this invention is for describing particular embodiments only and is not intended to limit the scope of exemplary embodiments according to the invention.

[0025] Example 1 like Figure 1 As shown, this application understands that traditional solutions in the existing technology often use a single modality (such as a pure vision camera), which has the problem of failure in low light / occlusion scenes and lacks high-precision limb movement detection (such as millimeter-wave radar is not widespread, and the limb movement detection accuracy is mostly > ±20mm); in addition, the voiceprint recognition in the existing technology is not optimized for the high-frequency crying characteristics of children (adult voiceprint models cannot distinguish the crying frequency of children >300Hz).

[0026] To address the technical deficiencies in the aforementioned solutions, this application designs a multimodal interaction-based intelligent cockpit child safety monitoring system for automobiles, comprising: Data acquisition layer 1 collects and processes multimodal information of users in the intelligent cockpit of the car through a multimodal sensor array; Edge processing layer 2, connected to data acquisition layer 1, identifies child features in multimodal information, performs local encryption on child features, and simultaneously performs fusion calculation on multimodal information to output security policy instructions; The execution control layer 3, connected to the edge processing layer 2, drives the vehicle to perform proactive intervention actions according to safety policy instructions.

[0027] In its implementation, the system of this application first employs a multimodal sensor array during the data acquisition phase. These sensors include, but are not limited to, a high-precision visual camera, a microphone array with optimized voiceprint recognition based on the characteristics of a child's high-frequency crying, and a high-precision millimeter-wave radar. Specifically, the multimodal sensor array of this application includes: Millimeter-wave radar collects the range of limb movements of users in a smart cockpit of a car. The specific installation coordinates in the vehicle can be defined as the millimeter-wave radar installation coordinates (X=1250±5mm, Y=±350mm, Z=850mm).

[0028] An array of microphones is used to identify the voiceprints of users in a smart car cockpit. The CAN bus is used to collect information on the door lock status or vehicle speed. The TOF camera performs 3D skeletal tracking for users in the smart cockpit of a car; the TOF camera has a horizontal viewing angle of 120° and a vertical viewing angle of 90°.

[0029] Seat pressure sensors detect the sitting posture of users in a car's smart cockpit.

[0030] The system includes a visual camera that captures images from inside the cockpit, providing clear images even in low-light conditions and preventing detection failures due to insufficient light. Millimeter-wave radar focuses on the precise detection of children's movements, with an accuracy within ±5mm. It captures subtle changes in children's movements in real time, providing accurate data for subsequent safety assessments. The microphone array is optimized for high-frequency crying sounds in children, accurately identifying cries above 300Hz. This contrasts sharply with traditional adult voiceprint models, effectively distinguishing children's cries from adult voices.

[0031] Next, the collected multimodal information is transmitted to edge processing layer 2. Edge processing layer 2 analyzes and processes the multimodal information. First, it uses feature recognition technology to accurately identify the child's characteristics from image, sound, and radar data. Once the child's characteristics are identified, the system immediately encrypts this feature information locally to ensure the security and privacy of the child's information. Simultaneously, edge processing layer 2 performs fusion calculations on the multimodal information, comprehensively analyzing the child's body movements, vocal states, and other information to determine whether the child is currently in a safe state. If the determination results in a safety hazard, the system will quickly output corresponding safety policy instructions.

[0032] Finally, after receiving the safety policy instructions output by the edge processing layer 2, the execution control layer 3 immediately drives the vehicle to perform active intervention actions. For example, when the system detects that a child is engaging in dangerous actions in the cabin, such as attempting to open the door or touching a dangerous area, the execution control layer 3 will quickly lock the door and adjust the seat position to secure the child in a safe area. If the system detects that a child is crying incessantly due to emotional agitation, it will play soothing music or a story the child likes through the car's audio system to calm the child down, while simultaneously sending a notification to the parents' mobile phones to keep them informed about the child's situation in the car. Through this multi-layered collaborative operation, the multimodal interactive automotive intelligent cabin child safety monitoring system of this application can effectively ensure the safety of children in the car cabin.

[0033] Furthermore, the edge processing layer 2 includes a child feature recognition engine, which is based on a deep learning algorithm to fuse and analyze multimodal information, identify and output child features.

[0034] Furthermore, edge processing layer 2 includes a child feature recognition engine. This engine, based on a deep learning algorithm, fuses and analyzes multimodal information to identify and output child features. In practice, the deep learning algorithm continuously learns and optimizes, improving the accuracy of child feature recognition through extensive data training. The deep learning algorithm in this application is based on a lightweight deep neural network using MobileNetV3 and incorporates a multimodal feature attention fusion mechanism, specifically including: Basic network architecture: The backbone network is MobileNetV3 (small version), which reduces the amount of computation through depthwise separable convolution and inverse residual structure, adapts to in-vehicle edge computing resources (such as J5 chip NPU), and supports INT8 quantization acceleration (computing power increased by 2.1 times and power consumption reduced to 0.9W).

[0035] To address the characteristics of multimodal data, the branch network is extended: Visual branch: 3D convolutional layer (inputting 3D skeletal coordinates from the TOF camera) + residual block (extracting limb motion features); Acoustics branch: 1D convolution (processing the acoustic spectrum of array microphones) + LSTM (capturing the temporal characteristics of crying frequencies > 300Hz); Physiological characteristics branch: Fully connected layer (integrating posture data from seat pressure sensors with limb amplitude data from millimeter-wave radar).

[0036] Specifically, the optimization strategy for the network framework described above in this application can be as follows: Sparse neural network pruning: Remove redundant neurons (pruning rate 40%), reducing the power consumption of the voiceprint recognition module to 0.9W (traditional solution 3.5W). Dynamic weight allocation: The weights of each modality are automatically adjusted through an attention mechanism (e.g., visual weight 0.7 and radar weight 0.3 in the windowing scene, acoustic weight 0.6 and pressure sensing weight 0.4 in the crying scene).

[0037] In its implementation, edge processing layer 2 includes a privacy protection unit that locally encrypts sensitive data in the child's characteristics using AES-256 encryption. Furthermore, edge processing layer 2 is equipped with a privacy desensitization module; this module desensitizes sensitive information in security policy instructions and outputs the desensitized security policy instructions.

[0038] It is worth noting that the privacy protection unit in this application is designed to achieve local encryption of data within the system, addressing the security of static data storage and preventing the leakage of original sensitive data (such as children's facial images and voiceprint audio) on local hardware. Specifically, it uses the AES-256 encryption algorithm to encrypt and store the original data (e.g., using a TPM 2.0 security chip), ensuring that even if the device is physically stolen, the data cannot be cracked. The privacy desensitization module aims to address the security of dynamic command transmission, preventing sensitive information from being associated and identified during internal system flow or cloud interaction. Intermediate data in security policy commands is desensitized (e.g., removing the identity identifier from "Child A's Window," retaining only "Window Behavior + Risk Level"), and outputting non-sensitive tagged commands. Both of these parts are necessary for the edge processing layer 2 of this application, mainly because even if the original data is encrypted, if the command generation contains associated information of "child's facial features + behavior tags" (e.g., "child's Window with User ID: 001 identified"), it is still possible to reverse engineer and crack privacy through data association.

[0039] Furthermore, the control layer 3 of this application includes a vehicle body domain controller and an execution device; The vehicle body domain controller is connected to edge processing layer 2 via the CAN-FD bus, outputting safety policy commands to the vehicle body domain controller. The vehicle body domain controller then drives the actuators to perform proactive intervention actions based on these safety policy commands. These actuators include a window controller, a fragrance system, and a voice prompt module. The window controller controls the operation of the vehicle's window lift motors; the fragrance system controls the opening and closing of the solenoid valve of the fragrance release module; and the voice prompt module outputs specified voice prompts based on the safety policy commands. Additionally, the actuators also include an infotainment system, which, upon receiving safety policy commands, accesses media resources stored locally in the vehicle.

[0040] In the specific implementation process, safety strategy commands include voice interaction, vehicle control such as windows and seats, playing soothing music, releasing fragrance, hazard lights, and notifying the front center console screen. According to the safety strategy commands, the vehicle domain controller controls the relevant equipment in the smart cockpit to execute the safety protection plan, such as playing soothing music, notifying the front center console screen, releasing fragrance, and adjusting window functions.

[0041] Example 2 Furthermore, the edge processing layer 2 in this application has a security policy library embedded in it. The security policy library has preset risk levels and penalty conditions. When the child's characteristics meet the risk level and penalty conditions, a security policy instruction is output.

[0042] like Figure 2 The diagram shows the system security intervention flowchart of this application. As can be seen from the diagram, the process starts with continuous multimodal monitoring. After the process starts from the "Start" node, it enters the continuous multimodal monitoring stage. The system collects vehicle and in-vehicle environment data in real time through multi-dimensional sensors (such as millimeter-wave radar, TOF, voiceprint, vision, pressure sensors, etc.) to provide a basis for subsequent risk assessment.

[0043] Furthermore, the core operation of this application's system involves risk assessment and tiered processing. By continuously monitoring data flowing into the risk assessment node, the system classifies risks into three levels—Level I (limb probing), Level II (continuous crying), and Level III (unlocking while driving)—based on preset risk levels and penalty conditions in the security policy library, triggering differentiated processing procedures for each level. (1) Level I risk: Limb exploration Triggering condition: The system determines that there is a risk of "limbs reaching out of the window" (such as a passenger extending their limbs out of the car window).

[0044] Detection and verification: Initiate "millimeter-wave radar + TOF cross-verification" to confirm the limb windowing behavior through cross-validation of data from the two sensors.

[0045] Threshold judgment: If the confidence level after verification is >0.99 (i.e., highly confirming limb protrusion), then "trigger window limit" is executed (e.g., limit the window opening to prevent the limb from protruding further); if the confidence level is ≤0.99, it is judged as non-risk and "continuous multimodal monitoring" is returned.

[0046] (2) Level II risk: persistent crying Triggering condition: The system determines that there is a risk of "continuous crying" (such as a child or passenger crying for a long time).

[0047] Detection and Analysis: Activate "Voiceprint + Visual Emotion Analysis" to jointly analyze emotional state through voiceprint features (such as the frequency and volume of crying sounds) and visual images (such as facial expressions and body movements).

[0048] Threshold judgment: If the crying behavior lasts for more than 30 seconds, then "start soothing strategy" (such as playing soothing music, adjusting the air conditioner temperature, etc.) will be executed; if the duration is less than or equal to 30 seconds, it will be judged as non-risk and return to "continuous multimodal monitoring".

[0049] (3) Level III risk: unlocking while driving Triggering condition: The system determines that there is a risk of "unlocking while driving" (such as a passenger attempting to unlock the door while the vehicle is in motion).

[0050] Detection and verification: Start the joint monitoring of "pressure sensor + CAN vehicle speed" to detect the door unlocking operation through the pressure sensor and read the vehicle speed data transmitted in real time through the CAN bus.

[0051] Threshold judgment: If an unlocking operation is detected and the vehicle speed is >5km / h (i.e., unlocking while driving), then “hazard lights warning + central control pop-up” will be executed (hazard lights warn surrounding vehicles and a safety prompt will pop up on the central control screen); if the vehicle speed is ≤5km / h (such as unlocking at low speed or when parked), it is judged as non-risk and returns to “continuous multimodal monitoring”.

[0052] Furthermore, the system in this application also includes a closed-loop feedback mechanism for real-time monitoring and strategy optimization. After the above three levels of risk handling, all processes enter the "real-time feedback monitoring" stage, where the system continuously tracks the effectiveness of risk handling. Risk Relief: If the risk is eliminated after processing (e.g., limb withdrawal, cessation of crying, termination of unlocking behavior), then “Record Learning Data” (stores the risk characteristics, processing process and results of this case), and “Update Strategy Weights” (optimizes the risk judgment model and processing strategy based on historical data), and finally returns to “Continuous Multimodal Monitoring” to enter the next cycle.

[0053] No improvement handling: If the risk does not improve within 10 seconds (e.g., the limbs are not retracted after the window is locked, or the crying does not stop after the soothing strategy), then “activate the alternative strategy” (e.g., upgrade the window lock, enhance the soothing measures, etc.), and return to “continuous multimodal monitoring” after handling.

[0054] Other aspects that are the same as those in Embodiment 1 of this application will not be described again in this embodiment.

[0055] Example 3 Regarding the system described in Embodiment 1 of this application, the following experiments were conducted. Specifically, this embodiment mainly focuses on testing children's crying, and the multimodal verification includes the following specific contents: Voiceprint analysis: Fundamental frequency 350Hz (characteristic of children crying); Visual confirmation: Tear detection is performed using a wide-angle TOF camera (enhanced periorbital reflection). Personalized adjustments: Account history check: Preference for Peppa Pig theme songs Execution: Play animated music while simultaneously controlling the release of fragrance (concentration 0.2ml / min).

[0056] In the specific experiments described above, the hardware deployment scheme of this application is shown in Table 1 below:

[0057] Table 1 Through the above detection, the system can identify the Level III risk behavior of a child unlocking a car door while the vehicle is in motion. When the millimeter-wave radar detects abnormal movements of a child in a specific area, and the flexible pressure sensor provides feedback on pressure changes in the corresponding area of ​​the rear seat, combined with images of the child's limb movements captured by the wide-angle TOF camera, the system can quickly determine the risk of a child unlocking a car door while the vehicle is in motion through multimodal data fusion analysis.

[0058] Once a risk is confirmed, the system will immediately trigger an alarm, emitting a warning sound through the car's audio system to alert the driver. At the same time, it will send a notification to the driver's smart terminal to ensure that the driver can take timely measures to protect the safety of children in the car.

[0059] Other aspects that are the same as those in Embodiment 1 of this application will not be described again in this embodiment.

[0060] It is noteworthy that those skilled in the art will understand that embodiments of the present invention can be provided as methods, systems, or computer program products. Therefore, the present invention can take the form of a completely hardware embodiment, a completely software embodiment, or an embodiment combining software and hardware aspects. Furthermore, the present invention can take the form of a computer program product embodied on one or more computer-usable storage media (including, but not limited to, disk storage, CD-ROM, optical storage, etc.) containing computer-usable program code.

[0061] This invention is described with reference to flowchart illustrations and / or block diagrams of methods, apparatus (systems), and computer program products according to embodiments of the invention. It will be understood that each block of the flowchart illustrations and / or block diagrams, as well as combinations of blocks in the flowchart illustrations and / or block diagrams, can be implemented by computer program instructions. These computer program instructions can be provided to a processor of a general-purpose computer, special-purpose computer, embedded processor, or other programmable data processing apparatus to produce a machine, such that the instructions, which execute via the processor of the computer or other programmable data processing apparatus, generate instructions for implementing the flowchart illustrations and / or block diagrams. Figure 1 One or more processes and / or boxes Figure 1 A device that provides the functions specified in one or more boxes.

[0062] These computer program instructions may also be stored in a computer-readable storage medium that can direct a computer or other programmable data processing device to function in a particular manner, such that the instructions stored in the computer-readable storage medium produce an article of manufacture including instruction means, which are implemented in a process Figure 1 One or more processes and / or boxes Figure 1 The function specified in one or more boxes.

[0063] These computer program instructions can also be loaded onto a computer or other programmable data processing equipment to cause a series of operational steps to be performed on the computer or other programmable equipment to produce a computer-implemented process, thereby providing instructions that execute on the computer or other programmable equipment for implementing the process. Figure 1 One or more processes and / or boxes Figure 1 The steps of the function specified in one or more boxes.

[0064] Finally, it should be noted that the above embodiments are only used to illustrate the technical solutions of the present invention and not to limit it. Although the present invention has been described in detail with reference to the above embodiments, those skilled in the art should understand that modifications or equivalent substitutions can still be made to the specific implementation of the present invention. Any modifications or equivalent substitutions that do not depart from the spirit and scope of the present invention should be covered within the scope of protection of the claims of the present invention.

[0065] It should be noted that, in this document, relational terms such as "first" and "second" are used only to distinguish one entity or operation from another, and do not necessarily require or imply any such actual relationship or order between these entities or operations. Furthermore, the terms "comprising," "including," or any other variations thereof are intended to cover non-exclusive inclusion, such that a process, method, article, or apparatus that comprises a list of elements includes not only those elements but also other elements not expressly listed, or elements inherent to such a process, method, article, or apparatus.

[0066] Although embodiments of the invention have been shown and described, it will be understood by those skilled in the art that various changes, modifications, substitutions and alterations can be made to these embodiments without departing from the principles and spirit of the invention, the scope of which is defined by the appended claims and their equivalents.

Claims

1. A child safety monitoring system for an intelligent car cockpit based on multimodal interaction, characterized in that, include: The data acquisition layer collects and processes multimodal information of the users of the intelligent cockpit of the car through a multimodal sensor array; An edge processing layer, connected to the data acquisition layer, identifies child features in the multimodal information, performs local encryption on the child features, and simultaneously performs fusion calculation on the multimodal information to output security policy instructions; The execution control layer, connected to the edge processing layer, drives the vehicle to perform proactive intervention actions according to the safety policy instructions.

2. The child safety monitoring system for an intelligent car cockpit based on multimodal interaction according to claim 1, characterized in that, The multimodal sensor array includes: Millimeter-wave radar collects the range of limb movements of users in the intelligent cockpit of the car; An array of microphones is used to identify the voiceprints of users in the intelligent cockpit of the car. A TOF camera is used to perform 3D skeletal tracking on the user of the car's smart cockpit. The seat pressure sensor detects the sitting posture of the user in the intelligent cockpit of the car.

3. The child safety monitoring system for an intelligent car cockpit based on multimodal interaction according to claim 2, characterized in that, The edge processing layer includes a child feature recognition engine, which is based on a deep learning algorithm to fuse and analyze the multimodal information, identify and output the child features.

4. A child safety monitoring system for an intelligent car cockpit based on multimodal interaction according to claim 2, characterized in that, The TOF camera has a horizontal viewing angle of 120° and a vertical viewing angle of 90°.

5. A child safety monitoring system for an intelligent car cockpit based on multimodal interaction as described in claim 1, characterized in that, The edge processing layer has an embedded security policy library, which is preset with risk levels and penalty conditions. When the child's characteristics meet the risk level and penalty conditions, a security policy instruction is output.

6. A child safety monitoring system for an intelligent car cockpit based on multimodal interaction as described in claim 5, characterized in that, The edge processing layer is equipped with a privacy desensitization module; The privacy desensitization module desensitizes sensitive information in the security policy instructions and outputs the desensitized security policy instructions.

7. A child safety monitoring system for an intelligent car cockpit based on multimodal interaction according to claim 1, characterized in that, The edge processing layer includes a privacy protection unit, which locally encrypts sensitive data in the child's characteristics using the AES-256 encryption method.

8. A child safety monitoring system for an intelligent car cockpit based on multimodal interaction according to claim 1, characterized in that, The execution control layer includes a vehicle domain controller and execution devices; The body domain controller is connected to the edge processing layer via a CAN-FD bus, and outputs the safety policy instructions to the body domain controller. The vehicle domain controller drives the actuator to perform proactive intervention actions according to the safety policy instructions.

9. A child safety monitoring system for an intelligent car cockpit based on multimodal interaction as described in claim 7, characterized in that, The actuator includes a window controller, a fragrance system, and a voice prompt module; The window controller is used to control the operation of the vehicle's window lift motors. Fragrance system, used to control the opening and closing of the solenoid valve of the fragrance release module in the vehicle; The voice prompt module outputs specified voice prompts according to the security policy instructions.

10. A child safety monitoring system for an intelligent car cockpit based on multimodal interaction according to claim 7, characterized in that, The execution device also includes an entertainment control center, which, upon receiving the security policy instruction, calls up media resources stored locally in the vehicle.