In-vehicle occupant conflict protection method, system, and vehicle
By acquiring in-vehicle data through multimodal perception technology, generating a conflict risk score, and triggering a vehicle response, the problem of insufficient accuracy in identifying in-vehicle occupant conflicts is solved, enabling accurate identification and safe intervention of occupant conflicts.
Patent Information
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- WUHAN JIANGXIA CHUNENG AUTOMOBILE TECHNOLOGY R&D CO LTD
- Filing Date
- 2026-04-02
- Publication Date
- 2026-06-02
AI Technical Summary
The accuracy of identifying occupant conflicts in existing technologies is insufficient, which affects vehicle driving safety.
By acquiring in-vehicle data through multimodal perception technology, including visual, audio, tactile, and driver state characteristics, a conflict risk score is generated. Based on the risk score, the vehicle response state is triggered to achieve accurate identification and intervention of occupant conflicts.
It improves the accuracy of identifying occupant conflicts. Through the comprehensive application of multimodal perception technology, it can more accurately judge the risk of conflict and trigger appropriate vehicle responses, thereby reducing the safety threat of conflict to vehicles and occupants.
Smart Images

Figure CN122126282A_ABST
Abstract
Description
Technical Field
[0001] This invention relates to the field of vehicles, and more particularly to a method, system, and vehicle for protecting against occupant conflicts. Background Technology
[0002] With the increasing popularity of ride-sharing, car-hailing, and long-distance family driving, sudden conflicts between vehicle occupants (especially passenger attacks on drivers) have become a serious safety hazard. To mitigate the impact of these conflicts on vehicle safety, it is necessary to identify and intervene in such conflicts; however, current conflict prevention methods lack sufficient accuracy in conflict identification. Summary of the Invention
[0003] This invention provides a method, system, and vehicle for preventing occupant conflicts in a vehicle, which addresses the technical problem of improving the accuracy of identifying occupant conflicts in a vehicle.
[0004] A first aspect of this invention provides a method for preventing conflicts between vehicle occupants. The method includes: obtaining a conflict vector based on acquired in-vehicle data, the conflict vector including a visual feature vector, an audio feature vector, a tactile feature vector, and a driver state feature vector; obtaining a visual risk score, an audio risk score, a tactile risk score, and a state risk score from the visual feature vector, the audio feature vector, the tactile feature vector, and the driver feature vector, respectively; obtaining a conflict risk based on the visual risk score, the audio risk score, the tactile risk score, and the state risk score; and triggering a vehicle response state based on the conflict risk.
[0005] In some implementations, obtaining the conflict vector based on the acquired in-vehicle data includes: acquiring motion data of the occupants using an in-vehicle visual perception unit, the motion data including the motion trajectory of the occupants' skeletal key points, occupant facial data, and hazardous material data; obtaining aggression confidence, motion amplitude intensity, and dynamic distance between occupants based on the motion trajectory of the occupants' skeletal key points; obtaining abnormal expression identifiers based on the occupants' facial data, the abnormal expressions including anger and fear; extracting the outlines of objects inside the vehicle using the in-vehicle visual perception unit, and generating a hazardous material identifier when a suspected hazardous material is extracted that is not present in the initial state of the vehicle, wherein the initial state of the vehicle is the state of the objects inside the vehicle when the occupants board the vehicle; and obtaining the visual feature vector based on the aggression confidence, motion amplitude intensity, dynamic distance between occupants, abnormal expression identifier, and hazardous material identifier.
[0006] In some implementations, obtaining the conflict vector based on the acquired in-vehicle data includes: acquiring in-vehicle speech audio data and non-speech audio data from the in-vehicle audio perception unit; obtaining an emotional intensity score based on the prosody, volume, and frequency in the speech audio data, and generating a threat keyword identifier when the speech audio data contains threatening keywords, wherein the threatening keywords include personal threats, insults, and help-seeking phrases; extracting abnormal sounds based on the non-speech audio data and determining the source direction of the abnormal sounds, wherein the abnormal sounds include hitting sounds; and obtaining the audio feature vector based on the emotional intensity score, the keyword identifier, the abnormal sounds, and the source direction of the abnormal sounds.
[0007] In some implementations, obtaining the conflict vector based on the acquired in-vehicle data includes: acquiring interactive pressure data from pressure sensors within the driving interaction component, wherein the interactive pressure data includes: steering wheel grip force distribution, steering wheel impact indicator, center console pressure distribution data, and gear lever area pressure distribution data; acquiring non-interactive pressure data from pressure sensors within the seat back or cushion, wherein the non-interactive pressure data includes: driver's seat cushion pressure data, non-driver's seat cushion pressure data, and driver's backrest pressure data; obtaining a steering wheel grabbing indicator based on the steering wheel grip force distribution and the steering wheel impact indicator; obtaining a center console grabbing indicator based on the center console pressure distribution data and the gear lever area pressure distribution data; obtaining a seat impact indicator based on the driver's seat cushion pressure data, non-driver's seat cushion pressure data, and driver's backrest pressure data; and obtaining the tactile feature vector based on the steering wheel grabbing indicator, the center console grabbing indicator, and the seat impact indicator.
[0008] In some implementations, obtaining the conflict vector based on the acquired in-vehicle data includes: acquiring driver physiological state data from a driver physiological state sensor and obtaining a stress physiological index based on the physiological state data, wherein the driver physiological state data includes: driver's heart rate and heart rate variability, driver's respiratory rate, respiratory pattern, and driver's facial temperature data; acquiring the driver's posture, gaze focus distribution, and operating habits from an in-vehicle visual perception unit, and obtaining a tension score and attention separation degree based on the difference between the posture, gaze focus, and operating habits and the driver's habitual baseline; and obtaining the driver's state feature vector based on the stress physiological index, the tension score, and the attention separation degree.
[0009] In some implementations, obtaining visual risk scores, audio risk scores, tactile risk scores, and state risk scores from the visual feature vector, audio feature vector, tactile feature vector, and driver feature vector respectively includes: spatiotemporally synchronizing the visual feature vector, audio feature vector, tactile feature vector, and driver state feature vector; and inputting the spatiotemporally synchronized visual feature vector, audio feature vector, tactile feature vector, and driver state feature vector into corresponding risk assessment sub-models to obtain the visual risk score, audio risk score, tactile risk score, and state risk score.
[0010] In some implementations, obtaining the conflict risk based on the visual risk score, the audio risk score, the tactile risk score, and the state risk score includes: inputting the conflict vector, the visual risk score, the audio risk score, the tactile risk score, and the state risk score into a self-attention model to obtain weights for the visual risk score, the audio risk score, the tactile risk score, and the state risk score; and calculating a weighted sum of the visual risk score, the audio risk score, the tactile risk score, and the state risk score based on the weights to obtain the conflict risk.
[0011] In some implementations, the vehicle's response state triggered based on the conflict risk includes: when the conflict risk is greater than a first threshold and less than a second threshold, controlling the instrument panel to output a warning graphic and controlling the driver's seat to output a warning vibration, wherein the warning vibration is a vibration that will not affect the driver's normal driving; when the conflict risk is greater than the second threshold and less than a third threshold, controlling the vehicle's audio system to output a warning audio signal, controlling the in-vehicle ambient lighting system to output a warning visual signal, controlling the seat belt to generate a single-pulse contraction force, and controlling the vehicle's operating components to apply damping force to the vehicle's operating parts; when the conflict risk is greater than the third threshold, controlling the seat belt to generate a continuous tightening force, controlling the vehicle to decelerate and stop the vehicle at the roadside or emergency lane.
[0012] A second aspect of this invention provides an in-vehicle occupant conflict prevention system, comprising: a data processing module for obtaining a conflict vector based on acquired in-vehicle data, the conflict vector including: a visual feature vector, an audio feature vector, a tactile feature vector, and a driver state feature vector; a risk assessment module for obtaining a visual risk score, an audio risk score, a tactile risk score, and a state risk score from the visual feature vector, the audio feature vector, the tactile feature vector, and the driver feature vector, respectively, and for obtaining a conflict risk based on the visual risk score, the audio risk score, the tactile risk score, and the state risk score; and a risk response module for triggering a vehicle response state based on the conflict risk.
[0013] A third aspect of the present invention provides a vehicle including an occupant conflict prevention system, which is used to implement the occupant conflict prevention method provided in the first aspect of the foregoing embodiments.
[0014] This invention provides a method for preventing occupant conflicts in a vehicle. The method includes: obtaining a conflict vector based on acquired in-vehicle data, the conflict vector including a visual feature vector, an audio feature vector, a tactile feature vector, and a driver state feature vector; obtaining a visual risk score, an audio risk score, a tactile risk score, and a state risk score based on each feature vector; comprehensively determining the conflict risk based on the combined risk scores; and triggering a vehicle response state based on the conflict risk. In other words, the method determines the occupant conflict risk more accurately by perceiving the occupant's state through multimodal sensing and comprehensively considering various modalities. Attached Figure Description
[0015] Figure 1 A flowchart illustrating the first method for preventing occupant conflict in a vehicle, as provided in an embodiment of the present invention; Figure 2 A flowchart illustrating the second method for preventing occupant conflict in a vehicle, as provided in an embodiment of the present invention; Figure 3 A flowchart illustrating the third method for preventing occupant conflict in a vehicle, provided in an embodiment of the present invention. Figure 4 A flowchart illustrating the fourth method for preventing occupant conflict in a vehicle, provided in an embodiment of the present invention; Figure 5 A flowchart illustrating the fifth method for protecting against occupant conflicts in a vehicle, provided in an embodiment of the present invention; Figure 6 A flowchart illustrating the sixth method for preventing occupant conflict in a vehicle, as provided in an embodiment of the present invention. Figure 7A flowchart illustrating the seventh method for preventing occupant conflict in a vehicle, as provided in an embodiment of the present invention. Figure 8 This is a schematic diagram of the architecture of an in-vehicle occupant conflict protection system provided in an embodiment of the present invention. Detailed Implementation
[0016] To make the objectives, technical solutions, and advantages of this invention clearer, the invention will be described in detail below with reference to the accompanying drawings and specific embodiments.
[0017] The specific technical features described in the various embodiments in the detailed implementation can be combined in various ways without contradiction. For example, different implementation methods can be formed by combining different specific technical features. In order to avoid unnecessary repetition, the various possible combinations of the specific technical features in this invention will not be described separately.
[0018] It should also be noted that, in order to avoid obscuring the present invention with unnecessary details, only the structures and / or processing steps closely related to the present invention are shown in the accompanying drawings, while other details that are not closely related to the present invention are omitted.
[0019] Additionally, it should be noted that the terms "comprising," "including," or any other variations thereof are intended to cover non-exclusive inclusion, such that a process, method, article, or apparatus that comprises a list of elements includes not only those elements but also other elements not expressly listed, or elements inherent to such a process, method, article, or apparatus. In the following description, the terms "first," "second," etc., are used merely to distinguish different objects and do not indicate any similarity or connection between them. It should be understood that the directional descriptions such as "above," "below," "inside," and "outside" refer to the orientation under normal use conditions.
[0020] In the following specific embodiments, the in-vehicle occupant conflict protection method can be applied to any vehicle. For example, the conflict protection method can be applied to a sedan, a bus, or a truck. The in-vehicle occupant conflict protection method will be described by way of example below with reference to various embodiments.
[0021] In some embodiments, such as Figure 1 As shown, the main steps of the vehicle occupant conflict prevention method include: Step S101: Obtain the conflict vector based on the acquired in-vehicle data.
[0022] The conflict vector includes visual feature vectors, audio feature vectors, tactile feature vectors, and driver feature vectors. Visual feature vectors represent the movement and relative positions of occupants, allowing for the determination of physical conflict. Audio feature vectors represent conversational patterns between occupants and the presence of abnormal sounds. Tactile feature vectors can determine if arguments or threats are present in conversations, and if abnormal impact sounds are heard. Tactile feature vectors represent the force applied to the driver's seat and vehicle controls, allowing for the determination of impacts or blows to the driver's seat and whether unintentional forces are applied to vehicle controls. Driver state feature vectors represent the driver's physiological and psychological state, allowing for the determination of anger or fear, and abnormal behavioral patterns.
[0023] It should be noted that conflict vectors can be obtained in any way. For example, image data of occupants can be acquired through an in-vehicle camera, and the outer contours of each occupant can be extracted and their roles can be labeled using a visual neural network. These roles include the driver and non-driver. Visual feature vectors can be extracted based on the insulation, movement patterns, and relative positional relationships of each occupant. For example, audio data inside the vehicle can be acquired through an audio acquisition array, and audio feature vectors can be extracted based on the temporal and spectral data in the audio data. For example, pressure data of the driver's seat and non-driver's seat can be acquired through pressure sensors installed in the seat cushion, and tactile feature vectors can be extracted from the pressure data. For example, micro-expression data of the driver can be acquired through a driver's facial camera, and driver state feature vectors can be extracted based on the micro-expression data.
[0024] Step S102: Obtain risk scores, audio risk scores, tactile risk scores, and state risk scores from the visual feature vector, audio feature vector, tactile feature vector, and driver state feature vector, respectively.
[0025] That is, risk scores corresponding to the feature vector type are obtained through each feature vector, thereby obtaining risk scores for each type. For example, each feature vector is substituted into a preset expert scoring table, and risk scores for each type are obtained through each expert scoring table.
[0026] Step S103: Obtain the conflict risk based on the visual risk score, audio risk score, tactile risk score and state risk score, and trigger the vehicle's response state based on the conflict risk.
[0027] This can be understood as combining multiple types of conflict risks to obtain an overall conflict risk, which represents the risk of conflict between occupants in the vehicle. That is, the conflict risk is obtained based on a combination of factors from multiple modalities, fully considering various types of parameters related to occupant conflict, thus enabling a more accurate judgment of occupant conflict. At the same time, the vehicle will also trigger corresponding response states based on the degree of conflict risk, for example, limiting the maximum driving speed of the vehicle when the conflict risk is greater than a preset threshold.
[0028] This invention provides a method for preventing occupant conflicts in a vehicle. The method includes: obtaining a conflict vector based on acquired in-vehicle data, the conflict vector including a visual feature vector, an audio feature vector, a tactile feature vector, and a driver state feature vector; obtaining a visual risk score, an audio risk score, a tactile risk score, and a state risk score based on each feature vector; comprehensively determining the conflict risk based on the combined risk scores; and triggering a vehicle response state based on the conflict risk. In other words, the method determines the occupant conflict risk more accurately by perceiving the occupant's state through multimodal sensing and comprehensively considering various modalities.
[0029] The following describes, in conjunction with various embodiments, Figure 1 In step S101, the steps of acquiring in-vehicle data and obtaining various types of feature vectors based on the in-vehicle data are illustrated by way of example.
[0030] In some embodiments, such as Figure 2 As shown, Figure 1 The steps in S101 of the process, which involve acquiring visual data inside the vehicle and obtaining visual feature vectors based on the visual data, include: Step S201: Obtain motion data of the occupants inside the vehicle using the in-vehicle visual perception unit.
[0031] The visual perception unit is located in key positions inside the vehicle (such as the center of the roof, A-pillar, and rearview mirror), deploying at least three wide-angle (at least 140 degrees wide-angle) infrared camera arrays. The cameras use 940-nanometer invisible infrared light sources for illumination, ensuring clear imaging even in complete darkness without interfering with occupants; motion data includes: skeletal key point motion trajectories, occupant facial data, and hazardous material data.
[0032] Step S202: Based on the motion trajectory of the key points of the occupant skeleton, obtain the confidence level of the aggressive action, the amplitude and intensity of the action, and the dynamic distance between the occupants.
[0033] Specifically, a lightweight HRNet model is used to extract the 3D coordinates of multiple key points of all occupants in real time (30 frames per second) based on the motion trajectory of skeletal key points, accurately describing the posture of the head, torso, and limbs. Based on the continuous sequence of skeletal key points, a spatiotemporal graph convolutional network (ST-GCN) model is used to identify aggressive action patterns and the confidence of each aggressive action pattern. For example, the spatiotemporal sequence of "rapid abduction of the upper arm - accelerated swing of the forearm - hand approaching the driver's head area" is identified as "punching attack" and the confidence of this action is determined. At the same time, the spatiotemporal graph convolutional network can also obtain the amplitude and intensity of the action and the dynamic distance between occupants.
[0034] Step S203: Obtain abnormal expression identifiers based on occupant facial data.
[0035] Abnormal facial expressions include anger and fear. For example, when abnormal audio or violent movements are detected, an auxiliary convolutional neural network (CNN) is activated to identify facial micro-expressions of extreme anger or fear. That is, this recognition function is only activated when it is determined that there may be a conflict through audio data, thereby saving the system's computing power.
[0036] Step S204: Extract the outline of items inside the vehicle based on the visual perception unit inside the vehicle. If a suspected dangerous item is found that does not exist in the initial state of the vehicle, generate a dangerous item label.
[0037] The initial state inside the vehicle refers to the state of the items inside the vehicle when the passenger gets in. That is, the items that were in the vehicle when the passenger first got in. If a suspected dangerous item suddenly appears during the journey, such as a hard water bottle or a hard tool, it is considered that the passenger has taken out a suspected dangerous item, and a dangerous item label needs to be generated.
[0038] Step S205: Obtain visual feature vectors based on the confidence level of aggressive actions, the amplitude and intensity of actions, the dynamic distance between occupants, abnormal facial expressions, and dangerous items.
[0039] The visual feature vector can be a vector containing the parameters mentioned above, or it can be a feature vector extracted by the feature extraction layer based on these parameters.
[0040] In some embodiments, such as Figure 3 As shown, Figure 1 The steps in step S101, which involve acquiring audio data and obtaining audio feature vectors based on the audio data, include: Step S301: The in-vehicle audio sensing unit acquires the speech audio data and non-speech audio data in the vehicle.
[0041] Specifically, a multi-microphone linear array is used, precisely positioned inside the roof lining, for beamforming and sound source localization. At the same time, the detection layer extracts the speech audio data and non-speech audio data of the occupants from the audio data. The speech audio data is used to represent the conversation content of the occupants, and the non-speech audio is used to represent the background sound inside the vehicle.
[0042] Step S302: Obtain an emotional intensity score based on the prosody, volume, and frequency in the speech audio data, and include threatening words, insulting words, and help-seeking phrases in the speech audio.
[0043] This can be understood as follows: by using a pre-trained model to extract emotions based on the prosody, volume, and frequency spectrum of the passenger conversation in the speech audio data to obtain an emotion intensity score, and at the same time, by using a keyword recognition module to detect whether there are explicit threatening, insulting words or distress phrases.
[0044] Optionally, the context can be determined by the emotional intensity score. If the passenger's conversation is in a state of strong anger, fear or other emotional states, the current conversation is determined to be in a high-risk context. In a high-risk context, a lightweight keyword recognition module is activated to detect whether there are explicit threatening, insulting words or distress phrases, so as to save the system's computing power and improve the recognition speed of dangerous keywords.
[0045] Step S303: Extract abnormal sounds based on non-speech audio data and determine the source direction of the abnormal sounds.
[0046] This can be understood as extracting the impact sound from the volume and frequency of non-verbal audio in the car, and then determining the direction of the abnormal impact sound by combining the time delay difference of the microphone array.
[0047] Step S304: Obtain audio feature vectors based on emotional intensity scores, keyword identifiers, abnormal sounds, and the sources of abnormal sounds.
[0048] The audio feature vector can be a vector containing the parameters mentioned above, or it can be a feature vector extracted by the feature extraction layer based on these parameters.
[0049] In some embodiments, such as Figure 4 As shown, Figure 1 The step S101 in the process of obtaining the tactile feature vector based on the in-vehicle pressure data includes: Step S401: Acquire interactive pressure data from the pressure sensor in the driver interaction component, and acquire non-interactive pressure data from the pressure sensor in the seat back or seat cushion.
[0050] The interactive pressure data includes steering wheel grip force distribution, steering wheel impact indicators, center console pressure distribution data, and gear lever area pressure distribution data. The interactive pressure data can determine whether the vehicle's interactive components used to operate the vehicle have been subjected to external forces not intended by the driver, and whether the relevant operating components are in a state of being seized. The non-interactive pressure data includes driver's seat cushion pressure data, non-driver's seat cushion pressure data, and driver's backrest pressure data. The non-interactive pressure data can determine whether the seat has been struck.
[0051] Step S402: Obtain a steering wheel grabbing indicator based on steering wheel grip force distribution and steering wheel impact indicator; obtain a center console grabbing indicator based on center console pressure distribution data and gear lever area pressure distribution data; and obtain a seat impact indicator based on driver's seat cushion pressure data, non-driver's seat cushion pressure data and driver's backrest pressure data.
[0052] Specifically, a thin-film pressure sensor matrix is integrated inside the 3 o'clock and 9 o'clock grip areas to monitor grip force distribution and abnormal impacts. If the grip force distribution and impact of the steering wheel are different from the force state of the steering wheel during normal driving, it is considered that the steering wheel is being contested and a steering wheel grabbing indicator is generated. Capacitive proximity sensors and pressure sensors are arranged below the surface of the center console and gear lever area to detect abnormal intrusion and grabbing actions by non-driver's hands. If abnormal intrusion and grabbing actions are present, a center console grabbing indicator is generated.
[0053] Meanwhile, a highly sensitive piezoelectric film is embedded in the backrest and seat cushion to sense pushing or hitting of the seat, especially pushing or hitting forces from behind on the driver's seat.
[0054] Step S403: Obtain tactile feature vectors based on steering wheel grabbing markers, hollow grabbing markers, and seat striking markers.
[0055] The tactile feature vector can be a vector containing the above-mentioned identifiers, or it can be a feature vector extracted by the feature extraction layer based on these identifiers.
[0056] In some embodiments, such as Figure 5 As shown, Figure 1 The steps in S101 of the process, which involve acquiring driver physiological state data and obtaining driver state feature vectors, include: Step S501: Obtain driver physiological state data from the driver physiological state sensor and obtain the stress physiological index based on the physiological state data.
[0057] The driver's physiological data includes: driver's heart rate and heart rate variability, driver's respiratory rate, breathing pattern, and driver's facial temperature data. Specifically, the driver's heart rate (HR) and heart rate variability (HRV) are continuously monitored through a photoelectric pulse wave (PPG) sensor integrated into the steering wheel. A significant short-term drop in HRV (such as a decrease in RMSSD value of more than 50%) is a sensitive indicator of stress response. The driver's chest micro-movements are detected non-contactly through clothing using a bio-radar (60GHz millimeter-wave radar) built into the seat, accurately extracting respiratory rate and pattern. Rapid and disordered breathing are physiological manifestations of panic or anger. Skin temperature changes in specific areas of the driver's face (such as around the nose and forehead) are monitored using a miniature thermal imager to help determine the state of emotional agitation. By comprehensively analyzing the above driver status data, the driver's stress physiological index is determined to identify whether the driver is in a state of stress such as tension, anger, or fear.
[0058] Step S502: The driver's posture, eye focus distribution and operating habits are obtained by the in-vehicle visual perception unit. Based on the difference between the driving posture, eye focus distribution and operating habits and the driver's habit baseline, the tension score and attention separation degree are obtained.
[0059] Specifically, the system continuously learns the driver's normal driving posture, eye focus distribution, and operating habits to establish a personalized baseline. When the system detects that the driver suddenly tenses up, frequently and quickly glances at the passenger seat or back seat, or operates stiffly, it identifies the difference between these abnormal states and the driver's habitual baseline, thereby obtaining a stress score and attention dissociation score.
[0060] Step S503: Based on the stress physiological index, tension score and attention separation degree, obtain the driver state feature vector.
[0061] The driver state feature vector can be a vector containing the above parameters, or it can be a feature vector extracted by the feature extraction layer based on these parameters.
[0062] In some embodiments, such as Figure 6 As shown, Figure 1 Step S102 includes: Step S601: Spatiotemporally synchronize the visual feature vector, audio feature vector, tactile feature vector, and driver state feature vector.
[0063] For example, all sensor data are stamped with a high-precision unified timestamp (microsecond level) to ensure that cross-modal events are strictly aligned on the time axis; at the same time, for data involving space, if the spatial coordinates are identified by a local generalized coordinate system, it is also necessary to unify the data in each generalized coordinate system to the vehicle coordinate system through coordinate transformation.
[0064] Step S602: Input the spatiotemporally synchronized visual feature vector, audio feature vector, tactile feature vector, and driver state feature vector into the corresponding risk assessment sub-models to obtain the visual risk score, audio risk score, tactile risk score, and state risk score.
[0065] Specifically, the feature vectors of each perception module are first input into the corresponding risk assessment sub-model to calculate the sub-risk score.
[0066] In some embodiments, such as Figure 7 As shown, Figure 1 In step S103, the conflict risks derived from visual risk scores, audio risk scores, tactile risk scores, and state risk scores include: Step S701: Input the conflict vector, visual risk score, audio risk score, tactile risk score, and state risk score into the attention model to obtain the weights of the visual risk score, audio risk score, tactile risk score, and state risk score.
[0067] This can be understood as inputting all sub-risk scores, original feature vectors, and contextual information (vehicle speed, whether it is traveling at high speed) into a self-attention fusion neural network (such as a lightweight version of Transformer based on attention mechanism) to obtain the weights of each risk score. This weight calculation method not only considers the correlation between different modalities but also the temporal correlation between different data, thereby more accurately calculating the correlation between each risk score and the conflict between occupants in the current state.
[0068] Step S702: Calculate the conflict risk by weighting the visual risk score, audio risk score, tactile risk score and state risk score.
[0069] That is, the conflict risk calculated by weighting and summing the risk scores based on the weights obtained from the self-attention model not only takes into account the influence of multimodal factors, but also the degree of influence of each modality on occupant conflict, thus enabling the conflict risk to more accurately reflect the conflict risk of occupants in the vehicle.
[0070] In some embodiments, Figure 1 The step S101, which involves determining the vehicle's response status based on conflict risk, includes: When the risk of conflict is greater than the first threshold but less than the second threshold, the instrument panel outputs a warning graphic, and the driver's seat outputs a warning vibration. When the risk of conflict is greater than the second threshold but less than the third threshold, the vehicle's audio system outputs a warning audio signal, the ambient lighting system outputs a warning visual signal, the seat belt generates a single-pulse contraction force, and the vehicle's operating components apply damping force to the vehicle's operating parts. When the risk of conflict is greater than the third threshold, the seat belt generates a continuous tightening force, the vehicle decelerates, and the vehicle stops at the roadside or emergency lane.
[0071] This can be understood as triggering different levels of response states based on different conflict risks. In a low-risk state, only prompts and deterrent signals are output. As the conflict risk increases, the vehicle response will tend to be more coercive, reducing the harm caused by the conflict to the occupants or vehicle safety through proactive intervention. The following are exemplary descriptions of the responses at each stage.
[0072] When the conflict risk is low (e.g., greater than 30 but less than 60), the vehicle's warning and deterrence response is triggered. A minimalist warning icon (such as a red exclamation mark on the side) is displayed to the driver via the head-up display or instrument panel, making it difficult for passengers to notice. The driver's seat initiates a warning vibration—a vibration that will not affect the driver's normal driving, such as a brief, slight vibration. Optionally, in this state, multimodal data is recorded cyclically (covering 30 seconds before and after) to preserve evidence for potentially escalating events, and the vehicle control system enters a "standby" state.
[0073] When the risk of conflict is high (e.g., greater than 60 and less than 75, or a clear act of aggression is detected), the vehicle's active intervention response is triggered. This includes playing a pre-recorded, neutral, authoritative voice prompt, such as "Please drive safely, please remain calm"; activating a warning sound at a specific frequency (3-5 Hz), a frequency proven effective in inducing unease and suppressing aggressive impulses; the interior ambient lighting turning red and flashing at a high frequency (approximately 7 Hz) to create a sense of urgency through visual stimulation and deter the attacker (the light will not cause physical harm to the attacker); targeting only the identified attacker, controlling the seatbelt motor for a rapid, forceful, short contraction (higher than the collision pretension but lower than the injury threshold), pressing their body firmly against the seatback and limiting their ability to move their arms significantly, for the duration of the risk; and applying a counter-damping torque through the electronic power steering (EPS) system to counteract abnormal steering wheel movements not input by the driver, making the attempt to seize the vehicle extremely difficult. Simultaneously, gear shifting or non-critical vehicle control functions can be temporarily locked. Optionally, the hazard lights will automatically activate to alert surrounding vehicles of a potential abnormal situation.
[0074] When the risk level is high (e.g., greater than 85, or detection of a fatal attack or risk of loss of vehicle control), the system triggers the vehicle's emergency avoidance and distress response, immediately tightening all seat belts to secure occupants in their seats. All doors and windows are automatically locked (except for the driver's side window, which retains its emergency descent function). If the vehicle is in motion (speed > 30 km / h), the system smoothly takes over longitudinal control: hazard lights are activated, automatic deceleration occurs, and after determining it is safe (e.g., through blind spot monitoring), the system smoothly moves into or stops in the rightmost lane or emergency lane. If the vehicle is stationary or moving at low speed, the system remains locked. Through the in-vehicle T-Box, the system automatically calls pre-set emergency contacts (e.g., family members, fleet safety center) and sends a distress message containing precise geographical location, vehicle identification, and real-time risk level. Simultaneously, a 60-second multimodal data summary (key video clips, audio clips, sensor logs) before and after the event is encrypted and uploaded to a cloud security server to prevent data corruption and serve as evidence in subsequent legal proceedings.
[0075] This invention also provides an in-vehicle occupant conflict prevention system, which is used to achieve the following: Figures 1 to 7 Any of the images shown depicts methods for protecting against occupant conflicts within a vehicle.
[0076] In some embodiments, such as Figure 8 As shown, the in-vehicle occupant conflict prevention system includes: a data processing module 100, a risk assessment module 200, and a risk response module 300. The data processing module 100 is used to obtain conflict vectors based on acquired in-vehicle data. These conflict vectors include: visual feature vectors, audio feature vectors, tactile feature vectors, and driver state feature vectors. The risk assessment module 200 is used to obtain visual risk scores, audio risk scores, tactile risk scores, and state risk scores from the visual feature vectors, audio risk scores, tactile risk scores, and driver feature vectors, respectively. It is also used to determine the conflict risk based on these scores. The risk response module 300 is used to trigger the vehicle's response state based on the conflict risk.
[0077] This invention also provides a vehicle including an occupant conflict prevention system, the conflict prevention system being used to achieve, for example... Figures 1 to 7 Any of the images shown depicts methods for protecting against occupant conflicts within a vehicle.
[0078] The above embodiments are only used to illustrate the technical solutions of the present invention and are not intended to limit it. Although the present invention has been described in detail with reference to preferred embodiments, those skilled in the art should understand that modifications or equivalent substitutions can be made to the technical solutions of the present invention without departing from the spirit and scope of the technical solutions of the present invention.
Claims
1. A method for protecting against occupant conflict in a vehicle, characterized in that, The methods for preventing conflicts between vehicle occupants include: The conflict vector is obtained based on the acquired in-vehicle data. The conflict vector includes: visual feature vector, audio feature vector, tactile feature vector and driver state feature vector. The visual feature vector, the audio feature vector, the tactile feature vector, and the driver feature vector are used to obtain the visual risk score, audio risk score, tactile risk score, and state risk score, respectively. The conflict risk is obtained based on the visual risk score, the audio risk score, the tactile risk score, and the state risk score, and the vehicle's response state is triggered based on the conflict risk.
2. The method for protecting against occupant conflict in a vehicle according to claim 1, characterized in that, The conflict vector obtained based on the acquired in-vehicle data includes: The vehicle's visual perception unit acquires motion data of the occupants, including the motion trajectory of key points in the occupants' skeletons, facial data of the occupants, and data on hazardous materials. Based on the motion trajectory of the key points of the occupant skeleton, the confidence level of the aggressive action, the amplitude and intensity of the action, and the dynamic distance between the occupants are obtained. Abnormal expression identifiers are obtained based on the occupant facial data, and the abnormal expressions include anger and fear; Based on the visual perception unit inside the vehicle, the outlines of objects inside the vehicle are extracted. When a suspected dangerous object that is not present in the initial state of the vehicle is extracted, a dangerous object label is generated. The initial state of the vehicle is the state of the objects inside the vehicle when the occupant gets in. The visual feature vector is obtained based on the confidence level of the aggressive action, the amplitude and intensity of the action, the dynamic distance between occupants, the abnormal facial expression markers, and the dangerous item markers.
3. The method for protecting against occupant conflict in a vehicle according to claim 1, characterized in that, The conflict vector obtained based on the acquired in-vehicle data includes: The in-vehicle audio sensing unit acquires speech audio data and non-speech audio data from inside the vehicle. An emotional intensity score is obtained based on the prosody, volume, and frequency in the language audio data. When the language audio data contains threatening keywords, a threatening keyword identifier is generated. The threatening keywords include words that threaten personal safety, insulting words, and help-seeking phrases. Abnormal sounds are extracted based on the non-speech audio data and the direction of the source of the abnormal sounds is determined. The abnormal sounds include striking sounds. The audio feature vector is obtained based on the emotional intensity score, the keyword identifier, the abnormal sound, and the source direction of the abnormal sound.
4. The method for protecting against occupant conflict in a vehicle according to claim 1, characterized in that, The conflict vector obtained based on the acquired in-vehicle data includes: Interactive pressure data is acquired by pressure sensors within the driving interaction component, including: steering wheel grip force distribution, steering wheel impact indicators, pressure distribution data of the center console, and pressure distribution data of the gear lever area. Non-interactive pressure data is acquired by pressure sensors within the seat back or cushion, including: driver's seat cushion pressure data, non-driver's seat cushion pressure data, and driver's backrest pressure data. A steering wheel grabbing indicator is obtained based on the steering wheel grip force distribution and the steering wheel impact indicator; a center console grabbing indicator is obtained based on the pressure distribution data of the center console and the pressure distribution data of the gear lever area; and a seat impact indicator is obtained based on the driver's seat cushion pressure data, the non-driver's seat cushion pressure data and the driver's backrest pressure data. The tactile feature vector is obtained based on the steering wheel grabbing mark, the center console grabbing mark, and the seat hitting mark.
5. The method for protecting against occupant conflict in a vehicle according to claim 1, characterized in that, The conflict vector obtained based on the acquired in-vehicle data includes: The driver's physiological state data is acquired by a driver physiological state sensor and a stress physiological index is obtained based on the physiological state data. The driver's physiological state data includes: driver's heart rate and heart rate variability, driver's respiratory rate, respiratory pattern and driver's facial temperature data. The driver's posture, eye focus distribution, and operating habits are obtained by the in-vehicle visual perception unit. Based on the difference between the posture, eye focus, and operating habits and the driver's habitual baseline, a tension score and attention separation degree are obtained. The driver's state feature vector is obtained based on the stress physiological index, the tension score, and the attention separation degree.
6. The method for protecting against occupant conflict in a vehicle according to any one of claims 1 to 5, characterized in that, The step of obtaining visual risk scores, audio risk scores, tactile risk scores, and state risk scores from the visual feature vector, the audio feature vector, the tactile feature vector, and the driver feature vector, respectively, includes: The visual feature vector, the audio feature vector, the tactile feature vector, and the driver state feature vector are spatiotemporally synchronized. The spatiotemporally synchronized visual feature vector, audio feature vector, tactile feature vector, and driver state feature vector are input into the corresponding risk assessment sub-models to obtain the visual risk score, audio risk score, tactile risk score, and state risk score.
7. The method for protecting against occupant conflict in a vehicle according to any one of claims 1 to 5, characterized in that, The conflict risk derived from the visual risk score, the audio risk score, the tactile risk score, and the state risk score includes: The conflict vector, the visual risk score, the audio risk score, the tactile risk score, and the state risk score are input into the self-attention model to obtain the weights of the visual risk score, the audio risk score, the tactile risk score, and the state risk score; The conflict risk is obtained by calculating the weighted sum of the visual risk score, the audio risk score, the tactile risk score, and the state risk score based on the weights.
8. The method for protecting against occupant conflict in a vehicle according to any one of claims 1 to 5, characterized in that, The response status of the vehicle triggered based on the conflict risk includes: When the conflict risk is greater than a first threshold and less than a second threshold, the instrument panel outputs a warning graphic and the driver's seat outputs a warning vibration, which is a vibration that will not affect the driver's normal driving. When the conflict risk is greater than the second threshold and less than the third threshold, the vehicle's audio system is controlled to output a prompt audio signal, the vehicle's ambient lighting system is controlled to output a prompt visual signal, the seat belt is controlled to generate a single pulse contraction force, and the vehicle's operating components are controlled to apply damping force to the vehicle's operating parts. When the risk of conflict exceeds the third threshold, the seat belt is controlled to generate a continuous tightening force, thereby controlling the vehicle to decelerate and stop at the roadside or emergency lane.
9. A vehicle occupant conflict protection system, characterized in that, The in-vehicle occupant conflict protection system includes: The data processing module is used to obtain a conflict vector based on the acquired in-vehicle data. The conflict vector includes: visual feature vector, audio feature vector, tactile feature vector, and driver state feature vector. The risk assessment module is used to obtain a visual risk score, an audio risk score, a tactile risk score, and a state risk score from the visual feature vector, the audio feature vector, the tactile feature vector, and the driver feature vector, respectively. It is also used to obtain a conflict risk based on the visual risk score, the audio risk score, the tactile risk score, and the state risk score. The risk response module is used to trigger the vehicle's response status based on the conflict risk.
10. A vehicle, characterized in that, The vehicle includes an in-vehicle occupant conflict prevention system, which is used to implement the in-vehicle occupant conflict prevention method as described in any one of claims 1 to 8.