Multi-modal data fusion cabin safety monitoring method and system

By integrating multimodal data and performing deep semantic analysis, a safety situation profile is generated, which addresses the shortcomings of existing systems in perception, decision-making, and execution, enabling proactive safety protection within the cockpit and enhancing the driving experience.

CN121469583APending Publication Date: 2026-02-06CHANGCHUN AUTOMOTIVE TEST CENT

Patent Information

Application Number
CN202511631971.0
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-11-10
Publication Date
2026-02-06

AI Technical Summary

Technical Problem

Existing driver and passenger monitoring systems have shortcomings in perception, decision-making, and execution. They cannot achieve synchronous collection and deep fusion of multimodal data, lack contextual understanding and predictive risk assessment, and each subsystem works independently, making it difficult to achieve proactive safety protection.

Method used

The system employs a combination of multimodal sensors to acquire data on personnel, vehicles, and the environment within the cockpit. Through deep feature extraction and cross-modal feature fusion, it generates a cockpit scenario feature vector, performs deep semantic analysis and risk assessment, generates a safety situation profile, and generates intervention strategies through a hierarchical decision-making mechanism, which are then transformed into specific control commands.

Benefits of technology

It has achieved scenario understanding and security situation profiling assessment, improved the adaptability of the decision-making system, and formed a multi-level, collaborative security protection system to ensure the safety of personnel in the cockpit.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN121469583A_ABST
    Figure CN121469583A_ABST
Patent Text Reader

Abstract

The invention discloses a multi-modal data fusion cabin safety monitoring method and system, and relates to the technical field of automobile safety driving. The method comprises the following steps: obtaining personnel, vehicle and environment data in a cabin; the data of personnel in the cabin comprises a driver and passengers, the vehicle data comprises a vehicle speed, an acceleration, a steering angle, a brake pressure and a light state, and the data of environment in the cabin comprises an audio signal in the cabin; performing depth feature extraction and cross-modal feature fusion on the acquired personnel, vehicle and environment data in the cockpit to obtain a cockpit scene feature vector; performing deep semantic analysis and risk assessment on the cockpit scene feature vector to obtain a security situation portrait; a corresponding intervention strategy is generated through a hierarchical decision-making mechanism according to the security situation portrait, the strategy is converted into a specific control instruction, and then the control instruction is converted into actual cabin and vehicle behaviors; a complete safety protection system is formed, and the safety of personnel in the cabin is guaranteed.
Need to check novelty before this filing date? Find Prior Art

Description

TECHNICAL FIELD

[0001] The present application relates to the technical field of automobile safe driving, in particular to a cabin safety monitoring method and system based on multi-modal data fusion. BACKGROUND

[0002] With the increasing intelligence of automobiles, the existing driver state monitoring system (DMS) and passenger monitoring system (OMS) have been difficult to meet the complex cabin safety requirements. The traditional technical solutions mainly have the following technical defects: firstly, in the perception dimension, the existing system mostly uses single vision or simple sound detection, lacks the synchronous collection and deep fusion of multi-modal data such as vital signs and fine behaviors, resulting in one-sided perception information; secondly, in the decision-making level, most systems use fixed threshold-based judgment rules or simple mathematical models, which cannot realize context-based scene understanding and predictive risk assessment, and the decision-making process is rigid; finally, in the execution level, the existing technology mostly stays in single intervention methods such as sound and light alarms, and lacks a multi-level and collaborative intervention mechanism linked with the vehicle control system. In addition, each subsystem often works independently, forming an "information island", which is difficult to achieve active safety protection from the system level. The existing technology with the publication number CN120229260A introduces multi-parameter risk assessment, but it still cannot solve the above-mentioned fundamental technical problems due to its reliance on indirect perception methods of WiFi modules and fixed mathematical calculation formulas. SUMMARY

[0003] In view of the above-mentioned prior art, the present application provides a cabin safety monitoring method and system based on multi-modal data fusion, mainly solving the technical problems existing in the background art.

[0004] To achieve the above-mentioned purpose, the technical scheme of the embodiment of the present application is as follows: In a first aspect, the present application provides a cabin safety monitoring method based on multi-modal data fusion, which comprises the following steps: Step S1, obtaining personnel, vehicle and environment data in the cabin; the personnel data in the cabin includes the driver and the passenger, the vehicle data includes the vehicle speed, the acceleration, the steering angle, the brake pressure, the light state, and the in-cabin environment data includes the in-cabin audio signal.

[0005] As a preferred scheme of the present application, a plurality of image sensors are arranged at different spatial positions in the vehicle cabin, the image sensors at least include a near-infrared image sensor facing the driver area and an RGB color image sensor covering the whole cabin, the near-infrared image sensor collects an infrared image sequence of the driver's face, and the RGB color image sensor collects a visible light image sequence containing the driver and all passengers; A uniform linear microphone array is arranged on the top of the cockpit to collect multi-channel audio signals from inside the cockpit. A communication connection is established with the vehicle chassis domain controller through the controller area network bus interface, and multiple sets of data frames on the vehicle bus are read. The data frames include the vehicle longitudinal speed value calculated by the wheel speed sensor, the vehicle three-axis acceleration value measured by the inertial measurement unit, the steering wheel angle value fed back by the electric power steering system, the master cylinder pressure value detected by the brake pressure sensor, and various headlight switch status quantities fed back by the body controller.

[0006] As a preferred embodiment of the present invention, the near-infrared image sensor acquires an infrared image sequence of the driver's face, including facial features, head posture, and hand position; The Euler angles of the driver's and passenger's head posture and the three-dimensional coordinates of the limb joints were obtained through the skeletal key point detection algorithm. The RGB color image sensor acquires a visible light image sequence containing all passengers, including personnel distribution, body posture, and identity recognition; and obtains a passenger distribution heatmap and identity classification confidence based on deep convolutional features by performing real-time instance segmentation of the passenger area through a convolutional neural network.

[0007] As a preferred embodiment of the present invention, the step of arranging a uniform linear microphone array on the top of the cockpit to collect multi-channel audio signals in the cockpit includes: calculating the signal arrival time difference between each microphone unit through a generalized cross-correlation algorithm, and then determining the three-dimensional spatial coordinates of the sound source in the cockpit through a direction-of-arrival estimation algorithm.

[0008] As a preferred embodiment of the present invention, a radar sensor is also installed on the top of the cockpit. The radar sensor transmits frequency-modulated continuous wave signals and receives echo signals reflected from the body surfaces of the driver and passengers. Based on the echo signal, the phase difference algorithm is used to extract the micro-Doppler features caused by the periodic fluctuations of the human chest cavity. Then, the respiratory rate waveform and heart rate pulse waveform corresponding to each occupant are separated and calculated through bandpass filtering and spectral peak detection algorithms. Furthermore, cluster analysis was used to associate different vital sign signal sources with their spatial locations within the cockpit, establishing a mapping table between occupant positions and vital sign signals.

[0009] Step S2 involves performing deep feature extraction and cross-modal feature fusion on the acquired data on personnel, vehicles, and environment within the cockpit to obtain a cockpit scenario feature vector.

[0010] As a preferred embodiment of the present invention, the step of performing deep feature extraction and cross-modal feature fusion on the acquired data on occupants, vehicles, and the environment in the cabin to obtain a cabin scenario feature vector specifically includes: The feature vectors extracted from each modality are normalized; The normalized feature vectors are input into a multi-head self-attention module, and the feature vectors of each modality are respectively taken as query vectors, key vectors and value vectors to calculate cross-attention weights between different modality features; the feature vectors of each modality are weighted and summed according to the calculated attention weights for feature fusion; The fused features are then input into a feedforward neural network for nonlinear transformation and feature enhancement; The original features and the enhanced features are combined by using a residual connection layer to retain the feature information at different abstraction levels, and a cockpit scene feature vector is obtained.

[0011] Step S3, performing deep semantic analysis and risk assessment on the cockpit scene feature vector to obtain a safety situation portrait.

[0012] As a preferred scheme of the present application, the cockpit scene feature vector is mapped to an initial representation in a hidden space through a linear transformation layer, and deep semantic encoding is performed through a multi-layer Transformer decoder module, and finally the hidden representation is mapped to a predefined scene semantic category space through an output projection layer; An instant risk value is calculated based on the current situation state, and a risk element mapping table is established to match the detected behavior patterns, environmental states and preset risk elements; The predicted risk value is calculated by weighted summation to obtain a safety situation portrait.

[0013] Step S4, generating a corresponding intervention strategy according to the safety situation portrait through a hierarchical decision mechanism, and converting the strategy into specific control instructions, and then converting the control instructions into actual cockpit and vehicle behaviors.

[0014] As a preferred scheme of the present application, the feature vector of the safety situation portrait is matched with the rule conditions of the predefined typical scene mode to generate a basic strategy set, and the typical scene mode includes fatigue driving mode, distraction driving mode, emotional abnormal mode and passenger abnormal behavior mode. Validated effective strategies are obtained based on similarity calculation by searching for similar scenes in a historical case library. The expected effect of different strategy combinations is evaluated through a strategy optimization algorithm, and the strategy combination with the best comprehensive performance is selected.

[0015] As a preferred scheme of the present application, the optimal strategy combination is converted into specific device control instructions, which specifically includes: A strategy-instruction mapping knowledge base is established to map strategy elements to control parameters. For human-computer interaction strategy, a control instruction set including display content, audio parameters, haptic mode is generated; For cabin environment strategy, environment regulation instructions including temperature setting, wind speed control, fragrance selection, lighting parameters are generated; For vehicle control strategy, vehicle control instructions including dynamics parameters, control system activation state, intervention intensity are generated.

[0016] In a second aspect, the present application also provides a multi-modal data fusion cabin safety monitoring system, which comprises: A multi-modal data acquisition module is configured to acquire cabin personnel, vehicle and environment data; the cabin personnel data includes drivers and passengers, the vehicle data includes vehicle speed, acceleration, steering angle, brake pressure, light state, and the cabin environment data includes cabin audio signals; A feature extraction and fusion module is configured to perform deep feature extraction and cross-modal feature fusion on the acquired cabin personnel, vehicle and environment data to obtain a cabin scenario feature vector; A safety situation assessment module is configured to perform deep semantic analysis and risk assessment on the cabin scenario feature vector to obtain a safety situation portrait; A strategy control module is configured to generate corresponding intervention strategies through a hierarchical decision mechanism according to the safety situation portrait, and to convert the strategies into specific control instructions and convert the control instructions into actual cabin and vehicle behaviors.

[0017] The present application has the beneficial effects that: through scenario understanding and safety situation portrait assessment, a mapping relationship from data features to semantic understanding is established, which has scenario cognitive ability and predictive risk assessment ability, and improves the adaptability of the decision system. Through hierarchical decision making, human-computer interaction, cabin environment regulation and vehicle dynamic control are deeply integrated, forming a complete safety protection system to ensure the safety of personnel in the cabin. BRIEF DESCRIPTION OF DRAWINGS

[0018] Figure 1 The present application provides a multi-modal data fusion cabin safety monitoring method step flow chart; Figure 2 The present application provides a multi-modal data fusion cabin safety monitoring system structure diagram. DETAILED DESCRIPTION

[0019] The technical solutions of the present application will be further described in detail below in conjunction with the accompanying drawings and specific embodiments. Unless otherwise defined, all technical and scientific terms used herein have the same meaning as commonly understood by one of ordinary skill in the art to which the present application belongs. The terms used in the specification of the present application are only for the purpose of describing specific embodiments and are not intended to limit the present application. In the following description, the expression "some embodiments" describes a subset of all possible embodiments, but it should be understood that "some embodiments" can be the same subset or different subsets of all possible embodiments and can be combined with each other without conflict.

[0020] In the following description, a large number of specific details are given in order to provide a more thorough understanding of the present application. However, it is obvious to those skilled in the art that the present application can be implemented without one or more of these details. In other cases, some technical features known in the art are not described in order to avoid obscuring the present application.

[0021] It should be understood that the present application can be implemented in different forms and should not be interpreted as being limited to the embodiments presented herein. On the contrary, these embodiments are provided to make the disclosure complete and full and to fully convey the scope of the present application to those skilled in the art. And the purpose of the terms used herein is only to describe specific embodiments and not as a limitation of the present application. As used herein, the singular forms "a", "an" and "the" are also intended to include the plural forms unless the context clearly indicates otherwise. It should also be understood that the terms "comprise" and / or "include", when used in the specification, determine the presence of the stated features, integers, steps, operations, elements, and / or components, but do not exclude the presence or addition of one or more other features, integers, steps, operations, elements, components, and / or groups. As used herein, the term "and / or" includes any and all combinations of the associated listed items.

[0022] It should also be noted that when an element is referred to as being "fixed to" another element, it can be directly on the other element or there can be an intervening element. When an element is referred to as being "connected to" another element, it can be directly connected to the other element or there can be an intervening element. The terms "vertical", "horizontal", "inner", "outer", "left", "right", and similar expressions used herein are for illustrative purposes only and are not intended to be the only implementation.

[0023] In order to fully understand the present application, detailed structures will be presented in the following description in order to illustrate the technical solutions presented by the present application. The alternative embodiments of the present application are described in detail as follows, however, in addition to these detailed descriptions, the present application can have other implementation manners.

[0024] In a first aspect, see the accompanying drawings Figure 1 The application provides a multi-modal data fusion cabin safety monitoring method, which comprises the following steps: Step S1, obtaining cabin personnel, vehicle and environment data; the cabin personnel data includes a driver and passengers, the vehicle data includes vehicle speed, acceleration, steering angle, brake pressure, light state, and the cabin environment data includes cabin audio signals.

[0025] As a preferred scheme of the application, a plurality of image sensors are arranged at different spatial positions in the vehicle cabin, the image sensors at least include one near-infrared image sensor facing the driver area and one RGB color image sensor covering the whole cabin, the near-infrared image sensor is configured to continuously work under low light conditions, by actively emitting a near-infrared waveband light source and receiving reflected light, and then collecting an infrared image sequence of the driver's face, and the RGB color image sensor collects a visible light image sequence containing the driver and all passengers at a sampling frequency of no less than 30 frames per second.

[0026] In the embodiment, the near-infrared image sensor collects an infrared image sequence of the driver's face, which includes facial features, head posture, and hand position; wherein the facial features include eyelid opening parameter, line of sight direction vector determined by geometric center coordinates of eyeball iris and pupil, and mouth opening and closing state parameter. The RGB color image sensor collects a visible light image sequence containing all passengers, which includes personnel distribution, body posture, and identity recognition; and obtains a passenger distribution heat map through real-time instance segmentation of the passenger area by a convolutional neural network and an identity classification confidence based on deep convolutional features.

[0027] In some embodiments, during the image acquisition process, the head posture Euler angle and limb joint three-dimensional coordinates of the driver and passengers are obtained through a skeletal key point detection algorithm; and the passenger distribution heat map obtained through real-time instance segmentation of the passenger area by a convolutional neural network and the identity classification confidence based on deep convolutional features.

[0028] As a preferred scheme of the application, a uniform linear microphone array is arranged at the top of the cabin to collect multi-channel audio signals in the cabin. In this embodiment, the step of arranging a uniform linear microphone array on the top of the cockpit to collect multi-channel audio signals in the cockpit includes: calculating the signal arrival time difference between each microphone unit using a generalized cross-correlation algorithm, determining the three-dimensional spatial coordinates of the sound source in the cockpit using a direction-of-arrival estimation algorithm, performing a short-time Fourier transform on the collected audio signals to obtain time-spectrum features, and inputting the time-spectrum features into a pre-trained deep neural network model for acoustic event classification, identifying predefined abnormal audio event types such as speech signals, baby cries, glass breaking sounds, and metal impact sounds, while outputting the type label of each audio event and the corresponding temporal positioning information.

[0029] As a preferred embodiment of the present invention, a communication connection is established with the vehicle chassis domain controller through the controller area network bus interface, and multiple sets of data frames on the vehicle bus are read. The data frames include the vehicle longitudinal speed value calculated by the wheel speed sensor, the vehicle three-axis acceleration value measured by the inertial measurement unit, the steering wheel angle value fed back by the electric power steering system, the master cylinder pressure value detected by the brake pressure sensor, and various headlight switch status quantities fed back by the body controller.

[0030] As a preferred embodiment of the present invention, a radar sensor is also installed on the top of the cockpit. The radar sensor transmits frequency-modulated continuous wave signals and receives echo signals reflected from the body surfaces of the driver and passengers. Based on the echo signal, the phase difference algorithm is used to extract the micro-Doppler features caused by the periodic fluctuations of the human chest cavity. Then, the respiratory rate waveform and heart rate pulse waveform corresponding to each occupant are separated and calculated through bandpass filtering and spectral peak detection algorithms. Furthermore, cluster analysis was used to associate different vital sign signal sources with their spatial locations within the cockpit, establishing a mapping table between occupant positions and vital sign signals.

[0031] Step S2 involves performing deep feature extraction and cross-modal feature fusion on the acquired data on personnel, vehicles, and environment within the cockpit to obtain a cockpit scenario feature vector.

[0032] As a preferred embodiment of the present invention, the step of performing deep feature extraction and cross-modal feature fusion on the acquired data on occupants, vehicles, and the environment in the cabin to obtain a cabin scenario feature vector specifically includes: The feature vectors extracted from each modality are standardized.

[0033] Specifically, for the driver's facial region, a facial detection algorithm based on a region proposal network is used to locate the facial bounding box. Then, a key point detection network is used to extract the geometric feature parameters of the eyelid contour, iris boundary, and lip contour, and calculate the real-time physiological state features of three dimensions: eyelid opening ratio, gaze deflection angle, and mouth opening degree. For the physical characteristics of the driver and passengers, a skeletal key point detection algorithm based on deep convolutional network is adopted to extract the two-dimensional coordinate sequence of the main joints of the head, neck and limbs from the full-body image. Then, the two-dimensional coordinates are converted into three-dimensional spatial coordinates with the vehicle cabin as the reference frame through the inverse perspective transformation algorithm to form a posture feature vector. The last fully connected layer of the deep convolutional network outputs a fixed-dimensional visual feature tensor, which integrates facial expression features, gaze direction features, and body posture features.

[0034] Specifically, the extraction of vehicle dynamic features involves first performing sliding window sampling on the input multi-dimensional time series data, such as vehicle speed, acceleration, steering angle, and braking pressure, to form time series data segments of equal length. Temporal convolutional networks are used to extract local features from each temporal data segment, capturing short-term driving behavior patterns such as rapid acceleration, sudden braking, and frequent steering. The feature sequences output by the temporal convolutional network are input into the long short-term memory network. Through its gating mechanism, the network learns long-term driving behavior dependencies and identifies persistent dangerous states such as fatigued driving and distracted driving. Finally, a fixed-dimensional driving behavior feature vector is extracted from the last time step of the Long Short-Term Memory network. This vector encodes the driving style and risk propensity in the current and historical periods.

[0035] Specifically, for the extraction of vital signs, firstly, the original intermediate frequency signal received by the radar is subjected to range-Doppler processing, and the joint spectrum of the range dimension and the Doppler dimension is obtained through two-dimensional fast Fourier transform. Moving targets with vital signs are located in the range-Doppler spectrum using a constant false alarm rate detection algorithm, and the Doppler spectrum of the corresponding range gate is extracted. Empirical mode decomposition was performed on the extracted Doppler spectrum to separate the low-frequency component caused by respiratory motion and the high-frequency component caused by heartbeat. Envelope detection and peak identification were performed on respiratory and heart rate signals respectively, and statistical characteristics of respiratory rate, heart rate and their variability were calculated. The above features are combined into a vital signs feature vector, which represents the physiological state and stress response level of the occupants.

[0036] Specifically, for the extraction of audio features, firstly, the audio signal of each channel is pre-emphasized, framed, and windowed, and then the acoustic features are obtained through the Mel frequency cepstral coefficient extraction algorithm. Meanwhile, a short-time Fourier transform is performed on the audio signal to obtain a time-frequency graph, which is then input into a pre-trained convolutional neural network to extract advanced audio features. Traditional acoustic features are combined with deep learning features to form an audio feature vector.

[0037] In this embodiment, the dimensions of each modality feature vector are mapped to the same feature space through a fully connected layer, and the scale of each feature vector is normalized by a layer normalization method to eliminate the dimensional differences between different modal features and to accurately align the feature sequences of different modalities on the time axis. The standardized feature vectors are used in a multi-head self-attention module. The feature vector of each modality is used as the query vector, key vector and value vector respectively. The cross attention weights between features of different modalities are calculated. The feature vectors of each modality are weighted and summed according to the calculated attention weights to perform feature fusion. The fused features are then input into a feedforward neural network for nonlinear transformation and feature enhancement. A residual connection layer is used to combine the original features with the enhanced features, retaining feature information at different levels of abstraction, to obtain the cockpit scenario feature vector.

[0038] Step S3: Perform deep semantic analysis and risk assessment on the cockpit scenario feature vector to obtain a safety situation profile.

[0039] As a preferred embodiment of the present invention, the cockpit scenario feature vector is subjected to deep semantic analysis and risk assessment to obtain a safety situation profile. Specifically, the cockpit scenario feature vector is mapped to an initial representation of the latent space through a linear transformation layer, and deep semantic encoding is performed through a multi-layer Transformer decoder module. Finally, the latent representation is mapped to a predefined scenario semantic category space through an output projection layer. In some embodiments, a structured scenario description is generated based on the scenario semantic category space, as follows: The main actors in the current scenario are identified by a softmax classifier, and the output includes the categories of actors such as driver, front passenger, and rear passenger, as well as their confidence scores. A multi-label classification method is used to identify the currently detected behavioral patterns, including predefined behavioral categories such as normal driving, frequent head turning, holding a mobile phone, closing eyes and dozing off, and abnormal body movements. And through attention weight analysis, the objects that the actor interacts with are identified, including in-vehicle equipment, other passengers, mobile electronic devices, etc. Based on facial expression features, voice tone features, and physiological indicators, the system estimates the emotional state dimensions of passengers through a regression model, including scores for three emotional dimensions: valence, arousal, and dominance. Combining vehicle dynamic features and external environmental information, the system identifies the current driving scenario, such as urban roads, highways, parking status, and other environmental context factors. Calculate the real-time risk value based on the current situation and establish a risk element mapping table to match the detected behavioral patterns and environmental conditions with the preset risk elements. In this embodiment, the risk element mapping table stores the correspondence between behavioral patterns, environmental states, and risk elements. Each entry includes a risk element identifier, a basic risk value, an environmental adjustment coefficient, and a duration coefficient. By querying the risk element mapping table using the contextual semantic description output by deep semantic parsing, a set of risk elements matching the currently detected behavioral patterns and environmental states can be obtained.

[0040] In this embodiment, calculating the immediate risk value based on the current scenario state includes extracting the current behavior pattern set from the scenario semantic parsing results, extracting the current environmental state parameters from the vehicle environmental data, and querying the risk element mapping table to obtain a matching risk element set. Then, a weighted summation algorithm is used to calculate the immediate risk value, where the contribution value of each risk element is the product of its base risk value and the corresponding adjustment coefficient.

[0041] By combining various risk factors and calculating the predicted risk value through a weighted summation method, a security situation profile is obtained.

[0042] In this embodiment, the calculation of the predicted risk value by integrating various risk factors includes considering the time evolution characteristics of the risk factors and introducing a risk trend factor to characterize the rate of risk change; considering the coupling effect between different risk factors and establishing a coupling coefficient matrix to describe the mutual influence relationship between the risk factors. The predicted risk value is calculated by integrating the real-time risk value, the risk coupling term, and the risk trend term using a multi-factor weighted fusion algorithm.

[0043] In this embodiment, the data structure of the security situation profile includes four main parts: timestamp information, contextual semantic description, risk assessment results, and risk element details. The timestamp information records the precise time point of the situation assessment; the contextual semantic description includes complete semantic annotations of the acting subject, behavioral patterns, interaction objects, and emotional states; the risk assessment results include immediate risk values, predicted risk values, risk trends, and confidence indices; and the risk element details list all activated risk elements and their contribution.

[0044] Step S4: Based on the security situation profile, generate corresponding intervention strategies through a hierarchical decision-making mechanism, transform the strategies into specific control commands, and then transform the control commands into actual cockpit and vehicle behavior.

[0045] As a preferred embodiment of the present invention, the feature vector of the safety situation profile is matched with the rule conditions of predefined typical scenario patterns to generate a basic strategy set. The typical scenario patterns include fatigue driving mode, distracted driving mode, abnormal emotion mode, and abnormal passenger behavior mode. By searching for similar scenarios in a historical case database, validated and effective strategies are obtained based on similarity calculations. The expected effects of different strategy combinations are evaluated using a strategy optimization algorithm, and the strategy combination with the best overall performance is selected.

[0046] As a preferred embodiment of the present invention, the optimal strategy combination is transformed into specific device control instructions, specifically including: Establish a policy-instruction mapping knowledge base to map policy elements to control parameters; In this embodiment, the knowledge base adopts a three-layer mapping structure: the first layer establishes a mapping relationship from strategy type to device type; the second layer defines the conversion rules from strategy parameters to device control parameters; and the third layer provides a specific numerical calculation model for the device control parameters. For example, for the "attention reminder" strategy, the strategy elements include reminder intensity, reminder method, and duration; the mapping rules include the mathematical relationship between tactile vibration intensity, display icon size, and prompt volume based on the reminder intensity; and the control parameters include the vibration motor voltage value, icon pixel size, and volume decibel value.

[0047] For human-computer interaction strategies, a set of control instructions is generated, including display content, audio parameters, and haptic modes. In this embodiment, regarding the display content instruction, the corresponding icon type is selected from the preset icon library according to the strategy type, the optimal display position coordinates are calculated based on the attention guidance principle, and the blinking frequency parameter, color change curve, and size animation sequence of the icon are defined. Regarding the audio parameter instruction, the corresponding tone template is selected from the tone template library according to the urgency of the strategy, the output volume gain is adaptively adjusted based on real-time data from the environmental noise sensor, and the frequency and duration parameters of the text prompt information are converted into speech signals through the speech synthesis engine. Regarding the haptic mode instruction, the frequency value, amplitude value, and rhythm mode parameters of the vibration motor are defined, the collaborative working sequence of multiple vibrators is specified, and the gradual curve function of haptic intensity is designed.

[0048] For cabin environment strategies, generate environmental adjustment instructions including temperature setting, wind speed control, fragrance selection, and lighting parameters; In this embodiment, regarding temperature control commands, the target temperature value is calculated based on strategy requirement parameters and occupant preferences. Wind speed level parameters are set according to the difference between the current temperature and the target temperature, as well as the urgency of the strategy. Directional airflow is achieved by controlling the angle of the damper actuator. Regarding fragrance control commands, the appropriate fragrance type is selected from the fragrance type library according to the strategy target type. The opening of the solenoid valve of the fragrance generator is precisely controlled to adjust the fragrance release concentration. The diffusion range of the fragrance is adjusted through the damper control parameters of the air conditioning system. Regarding lighting control commands, the color temperature value of the LED light source is adjusted according to the strategy type. The brightness value is adjusted based on data from the ambient light intensity sensor and the target effect requirements. The gradient curve of the light color and the brightness pulsation frequency are designed.

[0049] For vehicle control strategies, vehicle control commands are generated, including dynamic parameters, control system activation status, and intervention intensity.

[0050] In this embodiment, regarding dynamic parameter control, the power assist characteristic curve parameters and damping coefficient of the steering system are set, the response curve parameters and braking force distribution ratio of the braking system are set, and the torque mapping table of the engine management system and the shift strategy parameters of the transmission control unit are adjusted. Regarding control system activation commands, the activation command and parameter settings of the lane keeping assist system are sent, the following distance adjustment command of the adaptive cruise system is triggered, and the working state of the seat belt pretensioner and the target position of the seat posture adjustment motor are controlled. Regarding intervention intensity control, the intensity gradient of the progressive intervention is calculated based on the risk level parameters, the intervention parameters are fine-tuned based on the driver characteristic profile data, and the vehicle stability control module ensures that all intervention operations are within the preset safety boundary range.

[0051] Therefore, this invention establishes a cockpit safety active protection system, which achieves early identification and accurate prediction of safety risks through multimodal data fusion and deep scenario understanding; it achieves personalized and humanized safety protection through hierarchical decision-making and collaborative intervention mechanisms; and it ultimately constructs a complete closed loop from risk perception to active protection, thereby improving the safety protection level of the vehicle cockpit and the driving experience.

[0052] Secondly, this invention also provides a cockpit safety monitoring system based on multimodal data fusion, please refer to the appendix. Figure 2 The system includes: A multimodal data acquisition module is used to acquire data on personnel, vehicles, and the environment inside the cabin; the personnel data includes the driver and passengers, the vehicle data includes vehicle speed, acceleration, steering angle, braking pressure, and lighting status, and the cabin environment data includes cabin audio signals; The feature extraction and fusion module is used to perform deep feature extraction and cross-modal feature fusion on the acquired data of people, vehicles and environment in the cockpit to obtain cockpit scene feature vectors; The safety situation assessment module is used to perform deep semantic analysis and risk assessment on the cockpit scenario feature vector to obtain a safety situation profile. The strategy control module is used to generate corresponding intervention strategies based on the security situation profile through a hierarchical decision-making mechanism, and to transform the strategies into specific control commands, and then into actual cockpit and vehicle behavior.

[0053] Thirdly, the present invention also provides a computer-readable storage medium having a computer program stored thereon, wherein the computer program, when executed by a processor, implements the above-mentioned multimodal data fusion cockpit safety monitoring method.

[0054] In this embodiment, the computer-readable storage medium may be a non-transitory computer-readable storage medium, such as a read-only memory (ROM), random access memory (RAM), CD-ROM, magnetic tape, floppy disk, and optical data storage device.

[0055] Fourthly, embodiments of this application also provide an electronic device, including a processor and a memory, wherein the memory stores a plurality of instructions; the processor loads instructions from the memory to execute steps in any of the multimodal data fusion cockpit safety monitoring methods provided in embodiments of this application.

[0056] Fifthly, embodiments of this application also provide a computer program product, including a computer program / instructions, which, when executed by a processor, implement the steps in any of the multimodal data fusion cockpit safety monitoring methods provided in embodiments of this application.

[0057] In this embodiment, those skilled in the art will understand that all or part of the steps in the various methods of the above embodiments can be accomplished by instructions, or by instructions controlling related hardware. These instructions can be stored in a computer-readable storage medium and loaded and executed by a processor.

[0058] Therefore, embodiments of this application provide a computer-readable storage medium storing a plurality of instructions that can be loaded by a processor to execute steps in any of the multimodal data fusion cockpit safety monitoring methods provided in embodiments of this application.

[0059] It should be noted that, through the above description of the implementation methods, those skilled in the art can clearly understand that, for the sake of convenience and brevity, only the division of the above functional modules is used as an example. In actual applications, the above functions can be assigned to different functional modules as needed, that is, the internal structure of the device can be divided into different functional modules to complete all or part of the functions described above.

[0060] The above are merely specific embodiments of the present invention, but the scope of protection of the present invention is not limited thereto. Any variations or substitutions that can be easily conceived by those skilled in the art within the technical scope disclosed in the present invention should be included within the scope of protection of the present invention. The scope of protection of the present invention should be determined by the scope of the claims.

Claims

1. A cockpit safety monitoring method based on multimodal data fusion, characterized in that, The method includes the following steps: Acquire data on occupants, vehicles, and the environment within the cabin; the occupant data includes the driver and passengers, the vehicle data includes vehicle speed, acceleration, steering angle, braking pressure, and lighting status, and the cabin environment data includes cabin audio signals; The acquired data on people, vehicles, and environment inside the cockpit are subjected to deep feature extraction and cross-modal feature fusion to obtain a cockpit scene feature vector. Deep semantic analysis and risk assessment are performed on the cockpit scenario feature vectors to obtain a safety situation profile; Based on the security situation profile, a hierarchical decision-making mechanism is used to generate corresponding intervention strategies, which are then transformed into specific control commands, and finally into actual cockpit and vehicle behavior.

2. The security monitoring method for multimodal data fusion according to claim 1, characterized in that, Multiple image sensors are arranged in different spatial locations within the vehicle cabin. The image sensors include at least one near-infrared image sensor facing the driver's area and an RGB color image sensor covering the entire cabin. The near-infrared image sensor acquires an infrared image sequence of the driver's face, and the RGB color image sensor acquires a visible light image sequence containing the driver and all passengers. A uniform linear microphone array is arranged on the top of the cockpit to collect multi-channel audio signals from inside the cockpit. A communication connection is established with the vehicle chassis domain controller through the controller area network bus interface, and multiple sets of data frames on the vehicle bus are read. The data frames include the vehicle longitudinal speed value calculated by the wheel speed sensor, the vehicle three-axis acceleration value measured by the inertial measurement unit, the steering wheel angle value fed back by the electric power steering system, the master cylinder pressure value detected by the brake pressure sensor, and various headlight switch status quantities fed back by the body controller.

3. The cockpit safety monitoring method based on multimodal data fusion according to claim 2, characterized in that, The near-infrared image sensor acquires a sequence of infrared images of the driver's face, including facial features, head posture, and hand position. The Euler angles of the driver's and passenger's head posture and the three-dimensional coordinates of the limb joints were obtained through the skeletal key point detection algorithm. The RGB color image sensor acquires a visible light image sequence containing all passengers, including personnel distribution, body posture, and identity recognition; and obtains a passenger distribution heatmap and identity classification confidence based on deep convolutional features by performing real-time instance segmentation of the passenger area through a convolutional neural network.

4. The cockpit safety monitoring method based on multimodal data fusion according to claim 3, characterized in that, The method of arranging a uniform linear microphone array on the top of the cockpit to collect multi-channel audio signals in the cockpit includes: calculating the signal arrival time difference between each microphone unit through a generalized cross-correlation algorithm, and then determining the three-dimensional spatial coordinates of the sound source in the cockpit through a direction-of-arrival estimation algorithm.

5. The cockpit safety monitoring method based on multimodal data fusion according to claim 4, characterized in that, A radar sensor is also installed on the ceiling inside the cockpit. This radar sensor transmits frequency-modulated continuous wave signals and receives echo signals reflected from the bodies of the driver and passengers. Based on the echo signal, the phase difference algorithm is used to extract the micro-Doppler features caused by the periodic fluctuations of the human chest cavity. Then, the respiratory rate waveform and heart rate pulse waveform corresponding to each occupant are separated and calculated through bandpass filtering and spectral peak detection algorithms. Furthermore, cluster analysis was used to associate different vital sign signal sources with their spatial locations within the cockpit, establishing a mapping table between occupant positions and vital sign signals.

6. The cockpit safety monitoring method based on multimodal data fusion according to claim 5, characterized in that, The process of extracting deep features and fusing cross-modal features from the acquired data on occupants, vehicles, and the environment within the cockpit to obtain a cockpit scenario feature vector specifically includes: The feature vectors extracted from each modality are standardized. The standardized feature vectors are used in a multi-head self-attention module. The feature vector of each modality is used as the query vector, key vector and value vector respectively. The cross attention weights between features of different modalities are calculated. The feature vectors of each modality are weighted and summed according to the calculated attention weights to perform feature fusion. The fused features are then input into a feedforward neural network for nonlinear transformation and feature enhancement. A residual connection layer is used to combine the original features with the enhanced features, retaining feature information at different levels of abstraction, to obtain the cockpit scenario feature vector.

7. The cockpit safety monitoring method based on multimodal data fusion according to claim 6, characterized in that, The cockpit scenario feature vector is subjected to deep semantic analysis and risk assessment to obtain a safety situation profile. Specifically, the cockpit scenario feature vector is mapped to an initial representation of the latent space through a linear transformation layer, and deep semantic encoding is performed through a multi-layer Transformer decoder module. Finally, the latent representation is mapped to a predefined scenario semantic category space through an output projection layer. Calculate the real-time risk value based on the current situation and establish a risk element mapping table to match the detected behavioral patterns and environmental conditions with the preset risk elements. By combining various risk factors and calculating the predicted risk value through a weighted summation method, a security situation profile is obtained.

8. The cockpit safety monitoring method based on multimodal data fusion according to claim 7, characterized in that, The feature vector of the safety situation profile is matched with the rule conditions of predefined typical scenario patterns to generate a set of basic strategies. The typical scenario patterns include fatigue driving mode, distracted driving mode, abnormal emotion mode, and abnormal passenger behavior mode. By searching for similar scenarios in a historical case database, validated and effective strategies are obtained based on similarity calculations. The expected effects of different strategy combinations are evaluated using a strategy optimization algorithm, and the strategy combination with the best overall performance is selected.

9. A cockpit safety monitoring method based on multimodal data fusion according to claim 8, characterized in that, The optimal strategy combination is translated into specific device control commands, including: Establish a policy-instruction mapping knowledge base to map policy elements to control parameters; For human-computer interaction strategies, a set of control instructions is generated, including display content, audio parameters, and haptic modes. For cabin environment strategies, generate environmental adjustment instructions including temperature setting, wind speed control, fragrance selection, and lighting parameters; For vehicle control strategies, vehicle control commands are generated, including dynamic parameters, control system activation status, and intervention intensity.

10. A cockpit safety monitoring system based on multimodal data fusion, characterized in that, The system includes: A multimodal data acquisition module is used to acquire data on personnel, vehicles, and the environment inside the cabin; the personnel data includes the driver and passengers, the vehicle data includes vehicle speed, acceleration, steering angle, braking pressure, and lighting status, and the cabin environment data includes cabin audio signals; The feature extraction and fusion module is used to perform deep feature extraction and cross-modal feature fusion on the acquired data of people, vehicles and environment in the cockpit to obtain cockpit scene feature vectors; The safety situation assessment module is used to perform deep semantic analysis and risk assessment on the cockpit scenario feature vector to obtain a safety situation profile. The strategy control module is used to generate corresponding intervention strategies based on the security situation profile through a hierarchical decision-making mechanism, and to transform the strategies into specific control commands, and then into actual cockpit and vehicle behavior.

Citation Information

Patent Citations

  • Vehicle cabin intra-domain monitoring method and system based on WiFi module

    CN120229260A

  • Vehicle safety situation assessment and steady state driving mode switching method and system

    CN116279572A

  • Data model processing method and system

    CN120048127A

  • Multi-screen voice interaction system and method applied to automobile cabin

    CN120299457A

  • Intelligent cabin interaction method and system combined with digital twinning

    CN120509113A

Cited By

  • Passenger car intelligent cabin dynamic interaction system based on multi-passenger behavior recognition

    CN121706034A