Vehicle sentry mode threat early warning method and system and vehicle
By combining multimodal perception data processing and adaptive feature extraction with decision-making algorithms based on causal graphs and symbolic rules, the accuracy of intent recognition and the multi-level nature of response strategies in the vehicle sentry mode threat warning system are improved, solving the problems of high false alarm rate and single response in existing systems.
Patent Information
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- BEI DOU ZHI LIAN KE JI YOU XIAN GONG SI
- Filing Date
- 2026-03-03
- Publication Date
- 2026-05-01
AI Technical Summary
Existing vehicle sentry-based threat warning systems suffer from high false alarm rates, lack intent understanding capabilities, and have limited response strategies, making it difficult to achieve precise deterrence and evidence preservation.
It employs multimodal perception data acquisition and adaptive backbone network to extract key features. It generates high-level semantic feature sequences through cross-modal feature interaction fusion, and uses a causal graph-based propagation algorithm and a symbolic behavior rule knowledge base for hybrid decision-making. It outputs multidimensional behavioral intent results, performs quantitative threat assessment, and triggers differentiated responses.
It improves the accuracy of threat intent identification, enables multi-level response and continuous evolution capabilities, and enhances the system's environmental adaptability and the refinement of response strategies.
Smart Images

Figure CN121963440A_ABST
Abstract
Description
Technical Field
[0001] This application relates to the field of vehicle technology, and in particular to a vehicle sentinel mode threat warning method, system and vehicle. Background Technology
[0002] With the development of intelligent electric vehicles, "sentinel mode" has become a standard feature in high-end models. Its basic principle is: when the vehicle is locked and parked, the onboard camera continuously monitors the surrounding environment, and automatically starts recording and issues an alarm when an abnormal event is detected (such as a moving object approaching or the vehicle vibrating).
[0003] However, the existing Sentinel mode has the following significant drawbacks: The false alarm rate is extremely high: relying solely on visual motion detection, it cannot distinguish between harmless events such as pedestrians passing by, small animals jumping, and tree shadows swaying in the wind, and real threatening behaviors. This results in users receiving frequent invalid notifications, leading to "alarm fatigue," and ultimately causing them to choose to disable the feature.
[0004] Lack of intent understanding: Traditional systems can only determine "whether someone is approaching," but not "what this person intends to do." For example, they cannot distinguish between a brief stop to check tire pressure and malicious scratching of a car.
[0005] The response strategy is too simplistic: most systems only have a two-level response of "record / not record", lacking a refined and hierarchical early warning mechanism, making it difficult to achieve precise deterrence and evidence preservation. Summary of the Invention
[0006] In view of this, embodiments of this application provide a vehicle sentinel mode threat warning method, system, and vehicle, which can effectively solve the problem of low accuracy in identifying potential threat intent in existing methods.
[0007] In a first aspect, embodiments of this application provide a vehicle sentry mode threat warning method, including: Acquire multimodal perception data collected from the area around the vehicle in Sentry mode; The key features of each modality in the multimodal sensing data are extracted by an adaptive backbone network to obtain a multimodal feature sequence; The multimodal feature sequence is input into the cross-modal feature interaction fusion module to generate a high-level semantic feature sequence containing threat behavior information; The high-level semantic feature sequence is input into the high-level semantic reasoning module to perform hybrid decision-making using a causal graph-based propagation algorithm and a symbolic behavior rule knowledge base, and outputs a multi-dimensional behavioral intent result. A quantitative threat assessment is performed based on the multidimensional behavioral intent results, outputting multiple levels of threat levels, and triggering differentiated early warning and response strategies based on the threat levels.
[0008] Secondly, embodiments of this application provide a vehicle sentry mode threat warning system, including: The multimodal data acquisition module is used to acquire multimodal perception data collected from the area around the vehicle in sentry mode; An adaptive backbone network module is used to extract key features of each modality in the multimodal sensing data through an adaptive backbone network to obtain a multimodal feature sequence. A cross-modal feature interaction fusion module is used to receive the multimodal feature sequence and generate a high-level semantic feature sequence containing threat behavior information; The high-level semantic reasoning module is used to receive the high-level semantic feature sequence, and to perform hybrid decision-making using a causal graph-based propagation algorithm and a symbolic behavior rule knowledge base, and output a multi-dimensional behavior intent result. The threat assessment and response module is used to perform quantitative threat assessment based on the multidimensional behavioral intent results, output multiple levels of threat levels, and trigger differentiated early warning and response strategies based on the threat levels.
[0009] Thirdly, embodiments of this application provide a vehicle, the vehicle comprising: a vehicle body and a vehicle controller; the vehicle controller is used to implement a vehicle sentinel mode threat warning method provided in the first aspect of this application.
[0010] The embodiments of this application have the following beneficial effects: This application acquires multimodal perception data collected from around the vehicle in sentry mode; extracts key features of each modality within the multimodal perception data using an adaptive backbone network to obtain a multimodal feature sequence; inputs the multimodal feature sequence into a cross-modal feature interaction fusion module to generate a high-level semantic feature sequence containing threat behavior information; inputs the high-level semantic feature sequence into a high-level semantic reasoning module to perform hybrid decision-making using a causal graph-based propagation algorithm and a symbolic behavior rule knowledge base, outputting a multidimensional behavioral intent result; performs a quantitative threat assessment based on the multidimensional behavioral intent result, outputting multiple levels of threat levels, and triggering differentiated warning and response strategies based on the threat level. This application uses multimodal perception data and employs key feature extraction, feature interaction processing, and high-level semantic reasoning to finally obtain a multidimensional behavioral intent result. Threat levels are assessed based on the intent result, and corresponding response strategies are provided. Therefore, it can effectively solve the problem of low accuracy in identifying potential threat intent in existing methods. Attached Figure Description
[0011] To more clearly illustrate the technical solutions of the embodiments of this application, the accompanying drawings used in the embodiments will be briefly introduced below. It should be understood that the following drawings only show some embodiments of this application and should not be regarded as a limitation of the scope. For those skilled in the art, other related drawings can be obtained based on these drawings without creative effort.
[0012] Figure 1 A flowchart of a vehicle sentinel mode threat warning method according to an embodiment of this application is shown; Figure 2 This paper shows a structural block diagram of a vehicle sentinel mode threat warning system according to an embodiment of this application; Figure 3 This paper presents a flowchart illustrating the generation of multimodal feature sequences in the vehicle sentinel mode threat warning method according to an embodiment of this application. Figure 4 A flowchart illustrating the extraction of semantic features in the spatiotemporal domain from the vehicle sentinel mode threat warning method according to an embodiment of this application is shown. Figure 5 This paper illustrates a flowchart of a vehicle sentinel mode threat warning method for generating high-level semantic feature sequences according to an embodiment of this application. Figure 6 This paper shows a structural block diagram of a cross-modal feature interaction fusion module in a vehicle sentinel mode threat warning system according to an embodiment of this application; Figure 7 This paper illustrates a flowchart of a vehicle sentinel mode threat warning method according to an embodiment of the present application, in which multimodal feature sequences perform cross-modal semantic interaction processing. Figure 8 This paper shows a structural block diagram of a context-aware fusion unit within a cross-modal feature interaction fusion module in a vehicle sentinel mode threat warning system according to an embodiment of this application. Figure 9 This paper presents a flowchart illustrating the output of multi-dimensional behavioral intent results by the high-level semantic reasoning module in the vehicle sentinel mode threat warning method according to an embodiment of this application. Figure 10 This paper presents a flowchart illustrating a method for evaluating behavioral intent based on a causal graph in a vehicle sentinel mode threat warning method according to an embodiment of this application. Figure 11 This paper presents a flowchart illustrating a method for evaluating behavioral intent based on symbolic behavioral rules in a vehicle sentinel mode threat warning method according to an embodiment of this application. Figure 12 This paper shows a structural block diagram of a high-level semantic reasoning module in a vehicle sentinel mode threat warning system according to an embodiment of this application; Figure 13This paper illustrates a flowchart of a vehicle sentinel mode threat warning method for assessing threat levels according to an embodiment of this application. Figure 14 This paper illustrates a flowchart of a vehicle sentinel mode threat warning method for triggering a warning and response strategy according to an embodiment of this application. Figure 15 Another structural block diagram of the vehicle sentinel mode threat warning system according to an embodiment of this application is shown. Detailed Implementation
[0013] The technical solutions in the embodiments of this application will be clearly and completely described below with reference to the accompanying drawings. Obviously, the described embodiments are only some embodiments of this application, and not all embodiments.
[0014] The components of the embodiments of this application described and illustrated in the accompanying drawings can be arranged and designed in a variety of different configurations. Therefore, the following detailed description of the embodiments of this application provided in the drawings is not intended to limit the scope of the claimed application, but merely to illustrate selected embodiments of the application. All other embodiments obtained by those skilled in the art based on the embodiments of this application without inventive effort are within the scope of protection of this application.
[0015] In the following text, the terms "comprising," "having," and their cognates, which may be used in various embodiments of this application, are intended only to indicate a particular feature, number, step, operation, element, component, or combination thereof, and should not be construed as primarily excluding the presence of one or more other features, numbers, steps, operations, elements, components, or combinations thereof, or adding the possibility of one or more combinations thereof. Furthermore, the terms "first," "second," "third," etc., are used only for distinguishing descriptions and should not be construed as indicating or implying relative importance.
[0016] Unless otherwise specified, all terms used herein (including technical and scientific terms) shall have the same meaning as commonly understood by one of ordinary skill in the art to which the various embodiments of this application pertain. Terms (such as those defined in commonly used dictionaries) shall be interpreted as having the same meaning as in their contextual meaning in the relevant technical field and shall not be construed as having an idealized or overly formal meaning, unless clearly defined in the various embodiments of this application.
[0017] The following detailed description of some embodiments of this application is provided in conjunction with the accompanying drawings. Unless otherwise specified, the following embodiments and features can be combined with each other.
[0018] The existing Sentinel mode still has the following shortcomings: The response strategy is too simplistic: most systems only have a two-level response of "record / not record", lacking a refined and hierarchical early warning mechanism, making it difficult to achieve precise deterrence and evidence preservation.
[0019] Post-event retrieval is difficult: a massive amount of unlabeled video footage requires manual screening of key segments, which is inefficient and impractical.
[0020] Poor environmental adaptability: In low light conditions such as nighttime, rain, snow, fog, or severe weather, the camera performance deteriorates, and the system reliability is significantly reduced.
[0021] Therefore, this application provides a vehicle sentinel mode threat warning method, system and vehicle that can proactively identify potential threats, accurately judge behavioral intentions, support multi-level response and have continuous evolution capabilities.
[0022] The following examples illustrate the threat warning method for the vehicle's sentry mode.
[0023] Figure 1 A flowchart of a vehicle sentinel mode threat warning method according to an embodiment of this application is shown. Exemplarily, the vehicle sentinel mode threat warning method includes the following steps: S100 acquires multimodal perception data collected from the area around the vehicle in Sentry mode.
[0024] In this embodiment, after the vehicle enters the parking and locking state, the sentry mode is automatically activated. At this time, multiple sensors deployed around the vehicle simultaneously collect environmental information, forming multimodal perception data. Exemplarily, the multimodal perception data includes, but is not limited to, video streams, audio streams, and millimeter-wave radar point cloud streams.
[0025] Video stream: Video data input from vision sensors. Multiple high-resolution cameras cover a 360-degree field of view of the vehicle.
[0026] Point cloud flow: Input data from millimeter-wave radar and / or lidar is used to accurately detect the distance, velocity, angle, and trajectory of a target.
[0027] Audio stream: Input data from a high-sensitivity microphone array, used to collect ambient sounds and unusual noises.
[0028] S200, the key features of each modality in the multimodal sensing data are extracted through an adaptive backbone network to obtain a multimodal feature sequence.
[0029] The adaptive backbone network can dynamically adjust the processing strategy of multimodal feature sequences or the depth of the adaptive backbone network itself according to changes in computer resources or environment. It receives raw data from the multimodal data acquisition module and performs real-time analysis using built-in algorithm models to extract key features.
[0030] Multimodal perception data is fed into an adaptive backbone network for feature extraction to obtain a discriminative multimodal feature sequence that can effectively support threat intent identification. The multimodal feature sequence is a compact representation (embedding) with physical meaning or behavioral interpretability extracted from the original multimodal signal by a dedicated deep learning model and organized in time series.
[0031] For example, key features for video streams include shallow textures, deep semantics, behavioral actions, and spatiotemporal trajectories; such as raising a hand, holding a tool, body posture, and wandering trajectories.
[0032] For audio streams, key features include time-frequency characteristics, sound events, timing patterns, and sound source direction; for example, the sound of breaking glass, knocking rhythms, unusual dialogue, and azimuth. For point cloud flow, key features include geometry, motion parameters, micro-motions, and point cloud quality; for example, the shape of the handheld object, approach speed, dwell time, and point density.
[0033] S300, the multimodal feature sequence is input into the cross-modal feature interaction fusion module to generate a high-level semantic feature sequence containing threat behavior information.
[0034] The multimodal feature sequence is input into the cross-modal feature interaction and fusion module. A dynamic weighting mechanism and a cross-modal attention structure are used to realize multi-level information interaction and fusion to generate a high-level semantic feature sequence containing threat behavior information.
[0035] S400, the high-level semantic feature sequence is input into the high-level semantic reasoning module to perform hybrid decision-making using a causal graph-based propagation algorithm and a symbolic behavior rule knowledge base, and outputs a multi-dimensional behavioral intent result.
[0036] The high-level semantic reasoning module transforms the feature sequences from the previous module, the "cross-modal feature interaction fusion module," into advanced, interpretable threat semantic understanding. It outputs the probability distribution of threat intent, behavioral semantic labels, and key interpretable evidence chains, providing a direct and reliable basis for the final "risk assessment."
[0037] S500 performs a quantitative threat assessment based on the multidimensional behavioral intent results, outputs multiple levels of threat levels, and triggers differentiated early warning and response strategies based on the threat levels.
[0038] In other words, such as Figure 2As shown, the vehicle sentry mode threat warning system (hereinafter referred to as the threat warning system) includes a multimodal data acquisition module (for executing step S100), an adaptive backbone network module (for executing step S200), a cross-modal spatiotemporal feature fusion module (for performing time alignment, etc.), a cross-modal feature interaction fusion module (for executing step S400), and a threat assessment and response module (for executing step S500).
[0039] The purpose of this application is to overcome the shortcomings of the prior art and provide a vehicle sentinel mode threat warning method based on multimodal perception and intent recognition, so as to achieve an upgrade from "passive recording" to "proactive intelligent defense and collaboration".
[0040] In one implementation, such as Figure 3 As shown, in step S200, the key features of each modality within the multimodal sensing data are extracted through an adaptive backbone network to obtain a multimodal feature sequence, including: S210, For different modalities, corresponding deep neural network models are used to extract semantic features in the spatiotemporal domain from the multimodal perception data.
[0041] S220, during the process of extracting key features of each modality in the multimodal perception data, the complexity or feature processing strategy of the deep neural network model corresponding to each modality is dynamically adjusted according to the load status of the vehicle computing resources or changes in the external environment, so as to achieve a balance between efficiency and performance.
[0042] For different perception modalities, corresponding deep neural network models are used to extract semantic features in their respective spatiotemporal domains. The feature extraction process has runtime adaptive capability, which can dynamically adjust the complexity of the model or the signal processing strategy according to the load status of on-board computing resources or changes in the external environment, so as to achieve a balance between efficiency and performance while ensuring recognition accuracy.
[0043] The adaptive backbone network module runs dedicated neural network sub-models for different modalities, extracting key semantic features from each domain to form a multimodal feature sequence.
[0044] Furthermore, the multimodal sensing data includes at least two of the following: video stream, audio stream, and point cloud stream.
[0045] like Figure 4 As shown, the step of using corresponding deep neural network models for different modalities to extract semantic features in the spatiotemporal domain from the multimodal perception data includes: S211, for the video stream, the first backbone sub-network is used to extract multi-scale features of the image and identify spatial semantic information. Based on the load status of the on-board computing resources, the network depth, width or input image resolution of the first backbone sub-network is dynamically adjusted through composite model scaling technology.
[0046] Video processing: Based on the EfficientNet-B4 backbone network, multi-scale features of images are extracted, and video streams are analyzed in real time to extract spatial semantic features (people, actions, objects) in the video, resulting in shallow texture features and deep semantic features. Through composite model scaling technology (coordinated adjustment of depth, width, and resolution), the resolution of the input image or the network depth is dynamically adjusted. While ensuring accuracy, the computing resources are adaptively allocated according to the load of the vehicle controller.
[0047] S212, for the audio stream, a second backbone sub-network is used to extract local time-frequency features and model the temporal dependencies of sound events to identify at least one acoustic threat among screams, explosions, impact sounds or abnormal dialogues, and the noise reduction intensity is adjusted according to the environmental noise level in the changing external environment.
[0048] Audio processing: Analyzes audio signals, extracts local time-frequency features through convolutional recurrent neural networks (CRNNs), effectively models the temporal dependencies of sound events, and identifies acoustic threats (screams, explosions, impact sounds, abnormal conversations, etc.). Adaptively adjusts the noise reduction intensity of the audio front-end processing for different environmental noise levels.
[0049] S213, for the point cloud flow, a third backbone sub-network is used to construct a local dynamic map for each point in the point cloud space to capture the fine geometric structure and micro-motion features of the target, and the neighborhood range is dynamically adjusted according to the changes in the point cloud density in the vehicle computing resources.
[0050] Point cloud processing: Dynamic Graph Convolutional Neural Network (DGCNN) constructs a local dynamic graph for each point in the point cloud space, effectively capturing the fine geometric structure and local details of targets (such as people and tools). For identifying micro-movements such as handheld objects, it adapts to changes in point cloud density (e.g., radar point clouds may be sparser in rainy weather) by dynamically adjusting the graph neighborhood to maintain feature stability. By dynamically adjusting the neighborhood range to adapt to changes in point cloud density under different weather conditions, it maintains the stability of feature extraction.
[0051] In one embodiment, the method of this application further includes: The multimodal feature sequence is timestamped and semantically aligned using a cross-modal spatiotemporal feature fusion module. Exemplarily, the multimodal feature sequence is timestamped and normalized using the cross-modal spatiotemporal feature fusion module. Specifically, this includes the following: 1) Timestamp Alignment: This addresses the issue of inconsistent acquisition times in multimodal sensing data. For example, in a 30fps video stream, the time interval between frames is approximately 33.3ms; in an audio stream with a sampling rate of 16kHz, there is one sampling point every 1 / 16000 seconds; and in a radar stream with a 10Hz output, there is one frame every 100ms. This application's embodiment resolves latency differences in hardware-driven data source acquisition, including missing frames, out-of-order transmissions, batch transmissions, and clock drift caused by different timestamp sources for different modalities, by aligning timestamps, ensuring they are on the same timeline. The specific implementation steps for timestamp alignment are as follows: A unified clock reference, such as the system monotonic clock, is used as a benchmark, and the device clocks of all modes are converted to the GTMS internal time through a linear mapping. Inter-frame correction addresses the issue of unequal sampling rates between data sources. For example, it can solve the problem of dropped frames in audio and video by generating virtual time points through linear interpolation, and solve the problem of missing frames in radar data by copying the previous frame and marking it with a missing status.
[0052] Timestamp drift correction is based on the alignment of change points of the same event in multiple modes. For example, if a radar suddenly approaches and an object is detected accelerating towards it in the video, the timestamps in two modes are calculated and timestamp correction is performed.
[0053] 2) Uniform Sampling Rate: A uniform time step is used, interpolation is performed using repeated frames, and multimodal window synchronization is implemented to prevent information loss when the single-modal temporal density is insufficient. Adaptive resampling is also performed, automatically increasing the system's local sampling rate when critical events occur. This ensures higher temporal resolution for the next-level GMT transformer over a period of time, used to identify sudden acceleration of people or objects and instantaneous impacts.
[0054] 3) Feature Normalization: Maps multimodal features to a consistent scale and statistical distribution, enabling the attention mechanism of the subsequent GMT Transformer to be executed effectively.
[0055] In one implementation, such as Figure 5 As shown, in step S300, inputting the multimodal feature sequence into the cross-modal feature interaction fusion module to generate a high-level semantic feature sequence containing threat behavior information includes: S310, based on the quality index of each modal feature sequence, dynamically adjust its fusion weight, and preliminarily weight and fuse each modal feature sequence based on the fusion weight.
[0056] In such Figure 6The cross-modal feature interaction fusion module includes a dynamic weight attention unit. This dynamic weight attention unit dynamically adjusts the fusion weights based on the quality metrics of each modal feature sequence.
[0057] Quality indicators refer to the data quality of each modality feature sequence for assessing threatening behavioral intent, i.e., whether it can be used.
[0058] Reliability metrics refer to the reliability of each modal feature sequence in assessing the intent of threatening behavior. For example, they can be based on time (daytime, nighttime, duration of a certain behavior), scenario (residential area, commercial parking lot, remote road, etc.), and environment (weather, etc.).
[0059] S320 performs cross-modal semantic interaction processing on the multimodal feature sequences to achieve bidirectional interaction between the multimodal feature sequences, and performs feature enhancement processing on the multimodal feature sequences by combining time, space and environmental context information.
[0060] like Figure 6 Within the threat warning system, the cross-modal feature interaction fusion module also includes a feature interaction enhancement unit. This unit performs cross-modal semantic interaction processing on the multimodal feature sequences to achieve bidirectional interaction between them, and performs feature enhancement processing on the multimodal feature sequences by combining temporal, spatial, and environmental context information. S340, contextual semantic information is added to each enhanced modal feature sequence to generate a high-level semantic feature sequence that is highly adaptive to the environment.
[0061] The cross-modal feature interaction fusion module also includes a context-aware fusion unit. The context-aware fusion unit is used to add contextual semantic information to the fused modal feature sequences to generate high-level semantic feature sequences that are highly adaptive to the environment.
[0062] In step S310, the step of dynamically adjusting the fusion weights based on the quality indicators of each modal feature sequence, and initially weighting and fusing the modal feature sequences based on the fusion weights, includes: S311, analyze the feature sequence of each modality, extract at least one internal attribute from image sharpness, audio signal-to-noise ratio or radar point cloud density, and output the corresponding real-time confidence score by combining time, scene and environment information; normalize all the real-time confidence scores to obtain multiple fusion weights.
[0063] like Figure 6 As shown, within the threat warning system, the dynamic weighted attention unit is used to evaluate the confidence level of the modality, calculate the dynamic weights, and generate preliminary weighted fusion features.
[0064] Each modality is assigned a fixed weight. Compared to the commonly used fixed weight allocation, this application dynamically adjusts the weights based on the quality and reliability of real-time data, realizing data-driven modal weight allocation. This enables the system to adapt to harsh environments or partial sensor failures, thereby improving the system's robustness.
[0065] Within the dynamic weighted attention unit, a confidence evaluation network is used to feed the feature sequences of each modality in parallel into a lightweight sub-network. This sub-network analyzes the internal properties of the features (such as image sharpness, audio signal-to-noise ratio, and radar point cloud density) and outputs a real-time confidence score.
[0066] The three confidence scores are normalized through a Softmax layer to obtain a dynamic fusion weight that sums to 1. For example, in the middle of a rainy night, the weight of the visual modality may automatically drop to 0.2, the weight of the radar modality may rise to 0.6, and the weight of the audio modality may remain at 0.2.
[0067] The feature sequences of each modality are multiplied by the calculated dynamic fusion weights, and then concatenated or added together to form preliminary weighted fusion features (corresponding to the preliminary weighted fusion multimodal feature sequences).
[0068] Furthermore, such as Figure 7 As shown, in step S320, performing cross-modal semantic interaction processing on the initially weighted and fused multimodal feature sequences to achieve bidirectional interaction between the multimodal feature sequences, and performing feature enhancement processing on the multimodal feature sequences by combining temporal, spatial, and environmental context information, includes: S321, the first modal feature sequence in the multimodal feature sequence is used as the query vector, and the second modal temporal feature sequence therein is used as the key vector and value vector. The attention response is calculated and relevant information is aggregated to obtain the first enhanced feature sequence, so as to realize active information retrieval between modalities.
[0069] Feature interaction is enhanced through feature interaction units to generate new semantic features not found in a single modality, thereby improving the accuracy of subsequent intent recognition. Specifically, cross-modal cross-attention allows a feature from one modality to act as a "query" to actively retrieve the most relevant information from features from another modality (as the key and value). For example, when vision detects someone raising their hand to smash a window, the visual feature is used as the query to retrieve possible sounds of breaking glass from the audio features (key and value).
[0070] S322, the multimodal feature sequence is used as a node in the graph neural network to construct a fully connected graph structure. Through the message passing mechanism, each node aggregates information from other nodes and updates its own feature representation to obtain the second enhanced feature sequence.
[0071] By using graph neural network interactions, features from different modalities are treated as nodes in a graph, and a fully connected graph is constructed. Each modal feature node aggregates information from nodes of other modalities, updates its own feature representation, and captures complex cross-modal associations.
[0072] S323, integrate the first enhanced feature sequence and the second enhanced feature sequence.
[0073] The multimodal feature information obtained after interaction and compensation in steps S321 and S322 is integrated into a unified, high-dimensional feature vector, providing an information foundation for subsequent "context-aware fusion" and "high-level semantic reasoning".
[0074] In other words, such as Figure 6 As shown, within the threat warning system, the feature interaction enhancement unit performs cross-modal cross-attention, graph neural network interaction, and generates deep fusion features. Specifically, it integrates the first enhanced feature sequence and the second enhanced feature sequence to obtain deep fusion features.
[0075] Furthermore, the step of adding contextual semantic information to the enhanced modal feature sequences to generate high-level semantic feature sequences that are highly adaptive to the environment includes: encoding temporal, spatial, and environmental information for each modal feature sequence, aggregating the encoded modal feature sequences, and outputting an adaptive high-level semantic feature sequence.
[0076] In other words, such as Figure 8 As shown, within the threat warning system, the context-aware fusion unit receives deep fusion features and performs spatiotemporal context encoding, global information aggregation, and outputs environment-adaptive features. The spatiotemporal context encoding includes: Temporal encoding is performed on each deep fusion feature to obtain temporal context, facilitating the extraction of behavior duration. The temporal context, through sinusoidal positional encoding or learnable positional embedding, labels the absolute and relative time information of each time step, used to distinguish between day and night differences or identify persistent behavior. The temporal context encoding is also used to identify behavior duration and trigger a threat level increase when the lingering exceeds a preset threshold.
[0077] Spatial encoding is performed on each deep fusion feature to obtain spatial context, which facilitates the extraction of geographic location / scene type; the spatial context encodes GPS geographic location and IMU vehicle attitude data into feature vectors, which are then injected into the system as spatial prior knowledge. Environmental information is encoded for each deep fusion feature to facilitate the extraction of information such as weather and lighting.
[0078] Subsequently, global information aggregation is performed on the encoded features. The global information aggregation process includes attention mechanism processing, cross-modal attention fusion, and context information injection. Finally, after dynamic weighting and adaptive adjustment of feature scale, the output environment-adaptive features are produced. That is, the output environment-adaptive features are used to output adaptive features (corresponding to the high-level semantic feature sequence) after dynamic weighting and adaptive feature scale processing.
[0079] In other words, deeply fused features are further integrated with broader environmental, scene, and spatiotemporal context information to generate a final feature representation that is highly adaptive to the environment. Specifically, spatiotemporal context encoding injects temporal, spatial, and environmental information into the features, enabling the system to understand the "scene" in which the event occurred. This includes three sub-modules: Temporal context: using sinusoidal position encoding or learnable position embedding, it labels each time step in the feature sequence with its absolute and relative time information. For example, it encodes the difference between "11:00 PM" and "3:00 PM," or that a "loitering" behavior has lasted for 3 minutes. Spatial context: it encodes GPS geographic locations (such as underground parking lots, roadside parking spaces, and highway service areas) and IMU data (vehicle attitude) into feature vectors, injecting them into the system as spatial prior knowledge.
[0080] This application embodiment includes a cross-modal feature interaction fusion module, which not only achieves weighted fusion of multimodal features but also introduces a context-aware mechanism. First, the weights are dynamically adjusted based on the real-time confidence of each modality (e.g., whether the visual image is blurry or the radar image is sparse). Then, deep semantic interaction is achieved through cross-attention and graph neural networks. Finally, the fusion result is jointly modeled with contextual information such as 'time' and 'location', enabling the system to understand complex parking scenarios.
[0081] In one implementation, such as Figure 9 As shown, in step S400, the high-level semantic feature sequence is input into the high-level semantic reasoning module to perform hybrid decision-making using a causal graph-based propagation algorithm and a symbolic behavior rule knowledge base, outputting a multi-dimensional behavioral intent result, including: S410, the high-level semantic feature sequence is input into the causal graph-based propagation algorithm to obtain the first behavioral intent result.
[0082] S420, the various behavior judgment rules in the symbolic behavior rule knowledge base are transformed into continuous differentiable function expressions, and the function expressions are used to perform logical reasoning on the high-level semantic feature sequence to obtain the second behavior intention result; S430, perform a hybrid decision-making and arbitration based on the first behavioral intention result and the second behavioral intention result to obtain the multidimensional behavioral intention result.
[0083] The high-level semantic reasoning module is based on a probabilistic reasoning engine using causal graphs to calculate the posterior probability of "malicious intent." It also uses a symbolic knowledge base based on IF-THEN rules, which is transformed into a differentiable function and then used in end-to-end training. The outputs of both are fused through arbitration to form a multi-dimensional behavioral intent result that includes a threat intent probability distribution, behavioral semantic labels, confidence scores, and an interpretable evidence chain.
[0084] Furthermore, the advanced semantic feature sequence includes a probability distribution of threatening intent, behavioral semantic labels, and an interpretable chain of evidence.
[0085] like Figure 10 As shown, in step S410, inputting the high-level semantic feature sequence into a causal graph-based propagation algorithm to obtain the first behavioral intent result includes: S411, Construct a causal graph model, wherein the causal graph model is a directed acyclic graph, where nodes represent behavioral semantic concepts and edges represent causal relationships; Configure a conditional probability table for each non-root node to quantify the causal strength of the parent node to the corresponding non-root node.
[0086] Semantic concepts include at least one of the following: handheld tool, proximity to glass, knocking sound, and malicious intent. For example, the logic of "handheld tool detected" → "proximity to car window" + "knocking sound" → "determined as a high-risk attack" needs to be expressed and quantified using a formal causal graph model.
[0087] In this embodiment of the application, at least one of the following tags is used as a node: Behavioral semantic tags, such as: "loitering", "approaching a vehicle", "holding an object"; Contextual tags, such as: "nighttime", "underground parking", "rainy day"; Multimodal evidence labels, such as: "Audio: tapping sound", "Radar: distance <0.5m", "Video: hand raising action"); High-level inferences include: "suspected possession of a tool" and "constantly staring at the car window".
[0088] Specifically, the nodes include: handheld tools, approaching glass, knocking sounds, dwell time > 120 seconds, time = late at night, location = underground parking lot, malicious intent (root node), etc.
[0089] Exemplary examples show causal edges including: handheld tool → malicious intent, approaching glass → malicious intent, knocking sound → malicious intent, dwell time > 120s → malicious intent, late at night → malicious intent, handheld tool ∧ approaching glass → attack preparation phase, attack preparation phase ∧ knocking sound → malicious intent. In this embodiment, the causal strength is determined based on a conditional probability table. For each non-root node, a conditional probability table is constructed, representing the probability that the non-root node is "true" when its parent nodes take different combinations. For example, for the node "malicious intent", assuming its parent nodes are: A - handheld tool (Yes / No), B - approaching glass (Yes / No), C - knocking sound (Yes / No), the conditional probability table is shown in Table 1: Table 1 Conditional Probability Table
[0090] S412, receive the probability distribution of threat intent, behavioral semantic labels, and key interpretability evidence chains contained in the high-level semantic feature sequence as observational evidence and input them into the causal graph model.
[0091] High-level semantic feature sequences are received in real time and used as observational evidence, inputting them into the causal graph model. For example, probability distributions (e.g., 85% confidence level for "holding a tool"), behavioral semantic labels (e.g., "detected loitering behavior"), and interpretable evidence chains (e.g., "approaching in the 1st minute → lingering in the 2nd minute → raising hand in the 3rd minute") are input into the causal graph model. The interpretable evidence chain includes the detected key behavior, the triggered behavioral judgment logic rules, the contribution of the causal path, and multimodal verification.
[0092] S413, use the Bayesian probability propagation algorithm to calculate the posterior probability of all nodes regarding the malicious behavior intent, and finally obtain the first probability that the root node's malicious behavior intent is true. Based on the first probability and the interpretable evidence chain corresponding to the malicious behavior intent, obtain the first behavior intent result.
[0093] In other words, the first behavioral intent result includes the evidence chain corresponding to the malicious behavioral intent and the corresponding probability of the malicious behavioral intent, that is, the probability of the malicious behavioral intent P (malicious intent | evidence).
[0094] Evidence is mapped to the corresponding leaf nodes, triggering the Bayesian probability propagation algorithm to derive the posterior probability of the root node `P(malicious intent|all evidence)`.
[0095] Furthermore, the symbolic behavior rule knowledge base stores multiple behavior judgment rules expressed in the form of "IF-THEN". IF (condition 1 AND condition 2 … AND condition n) THEN (conclusion). Enabling the symbolic behavior rule knowledge base, which contains expert behavior judgment rules encoded in the form of "IF-THEN", is used to express high-order behavior judgment logic, for example: IF (behavior = loitering AND duration > 120 seconds AND environment = late at night) THEN (threat level increased by one level).
[0096] like Figure 11 As shown, in step S420, the step of converting each behavior judgment rule in the symbolic behavior rule knowledge base into a continuously differentiable function expression, and using the function expression to perform logical reasoning on the high-level semantic feature sequence to obtain the second behavior intent result, includes: S421, decode semantic propositions containing behavioral semantic labels, environmental context, or spatiotemporal state from the high-level semantic feature sequence.
[0097] S422, the differentiable logic reasoning layer maps each behavior judgment rule into a continuously differentiable function.
[0098] The behavior judgment rules are expressed in IF-THEN format, which is then converted into a mathematical expression. For example, IF (behavior = loitering) AND (duration ≥ 120s) AND (time is late at night) THEN → threat level + 1, in differentiable function form: RuleOutput = Sigma(w1f(“loitering”) + w2f(“duration”) + w3f(“night”) - b).
[0099] Where f(“loitering”): converts the original observation of “whether loitering behavior occurs” into a numerical feature (such as binarization or continuous scoring); f(“duration”): Maps “stay time” to a numerical value that reflects the degree of abnormality (such as a normalized time score). f(“Nighttime”): Determines whether the event occurred at night. Abnormal behavior is usually more likely to occur at night and can be used as an enhancing factor. w1, w2, w3: represent the importance coefficients of each feature, which are set by model training or manually; b: Threshold offset (similar to bias in a neuron), controls the overall trigger sensitivity; sigma: A sigmoid activation function used to compress a weighted sum into the (0, 1) interval, which can be interpreted as a probability or confidence level; RuleOutput: The final output is a soft decision result, which represents the probability that the system determines that there is an anomaly in the current scenario (such as suspicious lingering).
[0100] S423, calculate the threat intent probability distribution of the semantic proposition using each of the differentiable functions to obtain the second probability.
[0101] S424, Determine the multidimensional behavioral intent result based on the first probability and the second probability.
[0102] If both the first probability and the second probability represent a high threat probability, the highest threat outcome data is used directly. If the first probability and the second probability outputs are inconsistent, a weighted decision is made based on the confidence level of the current single outcome.
[0103] like Figure 12 As shown, within the threat warning system, the high-level semantic reasoning module decodes feature semantics to obtain semantic propositions containing behavioral semantic labels, environmental context, or spatiotemporal state. Then, it extracts high-level semantics from these propositions and performs hybrid decision-making using a causal graph propagation algorithm and a symbolic behavioral rule knowledge base. Specifically, the causal graph propagation algorithm includes constructing the causal graph, probability propagation and derivation processing, and outputting the intent probability. The processing based on the symbolic behavioral rule knowledge base includes symbolizing the behavioral judgment rules, performing differentiable logical reasoning, and outputting a logical conclusion. Finally, the multidimensional behavioral intent result is determined based on the intent probability and the logical conclusion.
[0104] During the training phase, the differentiable logic inference layer receives labeled real threat event samples as supervision signals and uses the gradient descent algorithm to jointly optimize rule weights and underlying deep model parameters to achieve end-to-end learning.
[0105] During the reasoning phase, the differentiable logic reasoning layer activates the matching logic rules based on the semantic propositions generated by the input multimodal fusion features, outputs the weighted logic conclusions, and participates in the final calculation of the threat intent probability distribution.
[0106] Through causal graphs and symbolic behavioral rules, the high-level semantic reasoning module can organize how to make a certain judgment for user trust, judicial evidence collection and system debugging. By using hybrid decision-making and arbitration to deal with uncertainties caused by real-world environments such as sensor noise and occlusion, it avoids arbitrary decision-making.
[0107] This application's embodiments include a high-level semantic reasoning module for transforming low-level features into threat judgments that are understandable to humans. This module integrates a causal graph model and a symbolic rule engine; the former models causal relationships such as 'handheld tool → attack intent,' while the latter encodes expert knowledge such as 'loitering at night escalates risk.' Through the collaborative work of Bayesian inference and differentiable logic, and by outputting the final decision through a hybrid arbitration mechanism, the system can not only 'make a judgment' but also 'explain the reasons,' greatly enhancing user trust and the likelihood of judicial acceptance.
[0108] In one implementation, structured results are received from the high-level semantic reasoning module, including the probability distribution of threat intent, behavioral semantic labels, confidence scores, and interpretable evidence chains. The threat is then precisely quantified and classified, and a quantified threat level is output.
[0109] Exemplary, such as Figure 13 As shown, in step S500, the step of performing a quantitative threat assessment based on the multi-dimensional behavioral intent results, outputting multiple levels of threat levels, and triggering differentiated early warning and response strategies based on the threat levels includes: S511, if it is determined from the multi-dimensional behavioral intent result that the target passes by the vehicle quickly without stopping, making contact, or making any suspicious movements or unusual sounds, then the threat level is determined to be Level 1 (harmless); for example, a pedestrian walks by normally.
[0110] S512, if it is determined from the multidimensional behavioral intent result that the target briefly stays near the vehicle (e.g., leans against the vehicle) without direct contact, then the threat level is determined to be Level 2 (attention / low risk); for example, someone is looking at their phone next to the car.
[0111] S513, if it is determined from the results of the multidimensional behavioral intent that the target has non-violent contact with the vehicle (such as touching the vehicle body with a hand), repeatedly lingers, or identifies a low-level abnormal sound, then the threat level is determined to be level three (warning / medium risk); for example, someone leans against the car door for a long time.
[0112] S514, if, based on the multi-dimensional behavioral intent results, at least two of the following are identified: the target is holding a suspicious tool, is within a preset distance from the vehicle and is stationary, and abnormal sounds are detected, then the threat level is determined to be Level 4 (high risk / high danger); a clear malicious intent or behavior is identified. For example, visual identification reveals a target holding a suspicious tool and approaching the vehicle window / door, while radar data confirms an extremely close distance and stationary position, and the audio module identifies the sound of breaking glass or violent knocking. Evidence from multiple modalities corroborates each other, forming a high-risk determination.
[0113] Furthermore, such as Figure 14As shown, the differentiated early warning and response strategy triggered based on the threat level includes: S521, if the threat level is level one, then it is determined not to send an alarm notification to the user and to selectively save the short video log locally; S522, if the threat level is level two, then the event will be recorded locally and marked in the log, and no alarm will be pushed proactively; S523, if the threat level is level three, then determine to start high-definition video recording and encrypt the storage, and at the same time send a mild prompt notification to the user terminal, such as "There is suspicious activity near your vehicle".
[0114] S524, if the threat level is level four, then it is determined to immediately push the highest priority alarm corresponding to the real-time video stream to the user terminal, activate the vehicle's sound and light alarm system (such as horn and flashing lights), or automatically upload key video clips to the cloud server for security storage.
[0115] Based on the results of high-level semantic reasoning, a quantitative threat level is generated, and corresponding local processing, user notification, vehicle device activation, and cloud-based linkage operations are performed based on the level, forming a complete intelligent defense closed loop from perception to response.
[0116] The embodiments of this application offer a variety of response strategies, enabling precise deterrence and evidence preservation. These embodiments upload key videos and other evidence to the cloud for convenient post-event retrieval.
[0117] Furthermore, such as Figure 15 As shown, the method in this application embodiment further includes: continuously optimizing learning based on incremental knowledge distillation, end-to-end cloud aggregation updates, and model self-repair.
[0118] Demonstratingly, the system achieves continuous learning and self-optimization through continuous optimization of the learning module. Edge-cloud collaboration addresses the problems of model rigidity, poor environmental adaptability, and inability to cope with new threats faced by traditional systems. Specifically, continuous learning optimization includes: Incremental knowledge distillation: Employs a "teacher-student" network architecture. The teacher model is the latest stable version currently deployed in a large number of vehicles. New threat knowledge is learned in the cloud, and training is performed jointly using distillation loss and new task loss.
[0119] Edge-to-cloud aggregation update: Through data upload from the edge, the cloud uses a federated averaging algorithm to aggregate model updates from tens of thousands of vehicles, generating a more powerful and universal global model. The latest global model is then encrypted (e.g., based on MD5+AES model signature verification) and distributed to all vehicles, enabling the evolution of collective vehicle intelligence.
[0120] Model self-healing: By tracking key metrics in real time, self-healing is initiated when performance degradation is detected. For example, parameters of the "video stream feature extraction" submodule with degraded performance can be rolled back or incrementally updated to improve efficiency and stability.
[0121] This application also provides a vehicle sentry mode threat warning system. Exemplarily, such as... Figure 15 As shown, the vehicle's sentry mode threat warning system includes: The multimodal data acquisition module is used to acquire multimodal perception data collected from the area around the vehicle in sentry mode; An adaptive backbone network module is used to extract key features of each modality in the multimodal sensing data through an adaptive backbone network to obtain a multimodal feature sequence. A cross-modal feature interaction fusion module is used to receive the multimodal feature sequence and generate a high-level semantic feature sequence containing threat behavior information; The high-level semantic reasoning module is used to receive the high-level semantic feature sequence, and to perform hybrid decision-making using a causal graph-based propagation algorithm and a symbolic behavior rule knowledge base, and output a multi-dimensional behavior intent result. The threat assessment and response module is used to perform quantitative threat assessment based on the multidimensional behavioral intent results, output multiple levels of threat levels, and trigger differentiated early warning and response strategies based on the threat levels.
[0122] It is understood that the device in this embodiment corresponds to the vehicle sentinel mode threat warning method in the above embodiment, and the options in the above embodiment are also applicable to this embodiment, so they will not be described again here.
[0123] This application also provides a vehicle, exemplary, including a vehicle body and a vehicle controller; the vehicle controller is used to implement a vehicle sentinel mode threat warning method as provided in this application.
[0124] It is understood that the device in this embodiment corresponds to the vehicle sentinel mode threat warning method in the above embodiment, and the options in the above embodiment are also applicable to this embodiment, so they will not be described again here.
[0125] This application also provides a terminal device, exemplary of which includes a processor and a memory, wherein the memory stores a computer program, and the processor executes the computer program to enable the terminal device to perform the functions of the aforementioned vehicle sentinel mode threat warning method or the various modules in the aforementioned vehicle sentinel mode threat warning method system. Exemplarily, the terminal device includes, but is not limited to, a vehicle controller.
[0126] The processor can be an integrated circuit chip with signal processing capabilities. The processor can be a general-purpose processor, including at least one of a Central Processing Unit (CPU), Graphics Processing Unit (GPU), Network Processor (NP), Digital Signal Processor (DSP), Application-Specific Integrated Circuit (ASIC), Field-Programmable Gate Array (FPGA), or other programmable logic devices, discrete gate or transistor logic devices, or discrete hardware components. The general-purpose processor can be a microprocessor or any conventional processor, capable of implementing or executing the methods, steps, and logic block diagrams disclosed in the embodiments of this application.
[0127] The memory can be, but is not limited to, Random Access Memory (RAM), Read Only Memory (ROM), Programmable Read-Only Memory (PROM), Erasable Programmable Read-Only Memory (EPROM), Electrically Erasable Programmable Read-Only Memory (EEPROM), etc. The memory is used to store computer programs, and the processor can execute the computer programs accordingly after receiving execution instructions.
[0128] This application also provides a computer-readable storage medium for storing the computer program used in the aforementioned terminal device. For example, the computer-readable storage medium may include, but is not limited to, various media capable of storing program code, such as a USB flash drive, a portable hard drive, a read-only memory (ROM), a random access memory (RAM), a magnetic disk, or an optical disk.
[0129] In the several embodiments provided in this application, it should be understood that the disclosed apparatus and methods can also be implemented in other ways. The apparatus embodiments described above are merely illustrative. For example, the flowcharts and block diagrams in the accompanying drawings show the architecture, functionality, and operation of possible implementations of apparatus, methods, and computer program products according to various embodiments of this application. In this regard, each block in a flowchart or block diagram may represent a module, segment, or portion of code containing one or more executable instructions for implementing a specified logical function. It should also be noted that, in alternative implementations, the functions marked in the blocks may occur in a different order than those marked in the drawings. For example, two consecutive blocks may actually be executed substantially in parallel, and they may sometimes be executed in reverse order, depending on the functions involved. It should also be noted that each block in the block diagram and / or flowchart, and combinations of blocks in the block diagram and / or flowchart, can be implemented using a dedicated hardware-based system that performs the specified function or action, or using a combination of dedicated hardware and computer instructions.
[0130] In addition, the functional modules or units in the various embodiments of this application can be integrated together to form an independent part, or each module can exist independently, or two or more modules can be integrated to form an independent part.
[0131] If the aforementioned functions are implemented as software functional modules and sold or used as independent products, they can be stored in a computer-readable storage medium. Based on this understanding, the technical solution of this application, in essence, or the part that contributes to the prior art, or a part of the technical solution, can be embodied in the form of a software product. This computer software product is stored in a storage medium and includes several instructions to cause a computer device (which may be a smartphone, personal computer, server, or network device, etc.) to execute all or part of the steps of the methods described in the various embodiments of this application.
[0132] The above description is merely a specific embodiment of this application, but the scope of protection of this application is not limited thereto. Any changes or substitutions that can be easily conceived by those skilled in the art within the scope of the technology disclosed in this application should be included within the scope of protection of this application.
Claims
1. A vehicle sentry mode threat early warning method, characterized in that, include: Acquire multimodal perception data collected from the area around the vehicle in Sentry mode; The key features of each modality in the multimodal sensing data are extracted by an adaptive backbone network to obtain a multimodal feature sequence; The multimodal feature sequence is input into the cross-modal feature interaction fusion module to generate a high-level semantic feature sequence containing threat behavior information; The high-level semantic feature sequence is input into the high-level semantic reasoning module to perform hybrid decision-making using a causal graph-based propagation algorithm and a symbolic behavior rule knowledge base, and outputs a multi-dimensional behavioral intent result. A quantitative threat assessment is performed based on the multidimensional behavioral intent results, outputting multiple levels of threat levels, and triggering differentiated early warning and response strategies based on the threat levels.
2. The vehicle sentry mode threat warning method according to claim 1, characterized in that, The step involves extracting key features of each modality within the multimodal sensing data using an adaptive backbone network to obtain a multimodal feature sequence, including: For different modalities, corresponding deep neural network models are used to extract semantic features in the spatiotemporal domain from the multimodal perception data. And / or, in the process of extracting key features of each modality within the multimodal perception data, the complexity of the deep neural network model or feature processing strategy corresponding to each modality is dynamically adjusted according to the load status of onboard computing resources or changes in the external environment.
3. The vehicle sentry mode threat warning method according to claim 2, characterized in that, The multimodal sensing data includes at least two of the following: video stream, audio stream, and point cloud stream; The method involves employing corresponding deep neural network models for different modalities to extract semantic features in the spatiotemporal domain from the multimodal perception data, and dynamically adjusting the complexity or feature processing strategy of each deep neural network model based on the load status of onboard computing resources or changes in the external environment, including: For the video stream, a first backbone sub-network is used to extract multi-scale features of the image and identify spatial semantic information. Based on the load status of the on-board computing resources, the network depth, width or input image resolution of the first backbone sub-network is dynamically adjusted through composite model scaling technology. For the audio stream, a second backbone subnetwork is used to extract local time-frequency features and model the temporal dependencies of sound events to identify at least one acoustic threat among screams, explosions, impact sounds, or abnormal conversations, and the noise reduction intensity is adjusted according to the environmental noise level in the changing external environment. For the point cloud flow, a third backbone sub-network is used to construct a local dynamic map for each point in the point cloud space to capture the fine geometric structure and micro-motion features of the target. The neighborhood range is dynamically adjusted according to the changes in the point cloud density in the vehicle computing resources.
4. The vehicle sentry mode threat warning method according to claim 1, characterized in that, The step of inputting the multimodal feature sequence into the cross-modal feature interaction fusion module to generate a high-level semantic feature sequence containing threat behavior information includes: Based on the quality index of each modal feature sequence, its fusion weight is dynamically adjusted, and the modal feature sequences are initially weighted and fused based on the fusion weight. Cross-modal semantic interaction processing is performed on the preliminarily weighted and fused multimodal feature sequences to achieve bidirectional interaction between the multimodal feature sequences, and feature enhancement processing is performed on the multimodal feature sequences in combination with time, space and environmental context information; Contextual semantic information is added to the enhanced modal feature sequences to generate high-level semantic feature sequences that are highly adaptive to the environment.
5. The vehicle sentry mode threat warning method according to claim 4, characterized in that, The quality index based on each modal feature sequence, dynamically adjusting its fusion weight, includes: The feature sequence of each modality is analyzed to extract at least one internal attribute from image sharpness, audio signal-to-noise ratio, or radar point cloud density, and combined with time, scene, and environmental information to output the corresponding real-time confidence score; all the real-time confidence scores are normalized to obtain multiple fusion weights; And / or, the step of performing cross-modal semantic interaction processing on the preliminarily weighted fused multimodal feature sequences to achieve bidirectional interaction between the multimodal feature sequences, and performing feature enhancement processing on the multimodal feature sequences in combination with temporal, spatial, and environmental context information, includes: The first modality feature sequence in the multimodal feature sequence is used as the query vector, and the second modality temporal feature sequence therein is used as the key vector and value vector. The attention response is calculated and relevant information is aggregated to obtain the first enhanced feature sequence. The multimodal feature sequence is used as a node in a graph neural network to construct a fully connected graph structure. Through a message passing mechanism, each node aggregates information from other nodes and updates its own feature representation to obtain the second enhanced feature sequence. Integrate the first enhanced feature sequence and the second enhanced feature sequence; And / or, the addition of contextual semantic information to the enhanced modal feature sequences to generate high-level semantic feature sequences that are highly adaptive to the environment includes: Each modal feature sequence is encoded with temporal, spatial, and environmental information. The encoded modal feature sequences are then aggregated to output an adaptive high-level semantic feature sequence.
6. The vehicle sentry mode threat warning method according to claim 1, characterized in that, The step of inputting the high-level semantic feature sequence into the high-level semantic reasoning module to perform hybrid decision-making using a causal graph-based propagation algorithm and a symbolic behavior rule knowledge base, and outputting a multi-dimensional behavioral intent result, includes: The high-level semantic feature sequence is input into a causal graph-based propagation algorithm to obtain the first behavioral intent result; Each behavior judgment rule in the symbolic behavior rule knowledge base is transformed into a continuously differentiable function expression, and the function expression is used to perform logical reasoning on the high-level semantic feature sequence to obtain the second behavior intent result; The multidimensional behavioral intention result is obtained by performing a hybrid decision-making and arbitration based on the first behavioral intention result and the second behavioral intention result.
7. The vehicle sentry mode threat warning method according to claim 6, characterized in that, The advanced semantic feature sequence includes a probability distribution of threatening intent, behavioral semantic labels, and an interpretable chain of evidence. The step of inputting the high-level semantic feature sequence into a causal graph-based propagation algorithm to obtain the first behavioral intent result includes: A causal graph model is constructed, which is a directed acyclic graph, where nodes represent behavioral semantic concepts and edges represent causal relationships; a conditional probability table is configured for each non-root node to quantify the causal strength of the parent node to the corresponding non-root node; The probability distribution of threat intent, behavioral semantic labels, and key interpretable evidence chains contained in the high-level semantic feature sequence are received as observational evidence and input into the causal graph model. The posterior probability of all nodes regarding the malicious intent is calculated using the Bayesian probability propagation algorithm. Finally, the first probability that the root node's malicious intent is true is obtained. Based on the first probability and the interpretable evidence chain corresponding to the malicious intent, the first intent result is obtained. And / or, the symbolic behavior rule knowledge base stores multiple behavior judgment rules represented in the form of "IF-THEN"; the step of converting each behavior judgment rule in the symbolic behavior rule knowledge base into a continuously differentiable function expression, and using the function expression to perform logical reasoning on the high-level semantic feature sequence to obtain the second behavior intent result includes: Decode semantic propositions containing behavioral semantic labels, environmental context, or spatiotemporal state from the high-level semantic feature sequence; Each behavior judgment rule is mapped to a continuously differentiable function through a differentiable logic reasoning layer; The second probability is obtained by calculating the threat intent probability distribution of the semantic proposition through each of the differentiable functions.
8. The vehicle sentry mode threat warning method according to claim 1, characterized in that, The multidimensional behavioral intent results include the probability distribution of threatening intent, behavioral semantic labels, confidence scores, and interpretable evidence chains; The process of performing a quantitative threat assessment based on the multidimensional behavioral intent results, outputting multiple levels of threat levels, and triggering differentiated early warning and response strategies based on the threat levels includes: If, based on the multidimensional behavioral intent results, it is determined that the target passes by the vehicle quickly without stopping, making contact, engaging in any suspicious actions, or making any abnormal sounds, then the threat level is determined to be Level 1. If, based on the results of the multidimensional behavioral intent, it is determined that the target briefly stayed near the vehicle without direct contact, then the threat level is determined to be level two. If, based on the results of the multidimensional behavioral intent, it is determined that the target has engaged in non-violent contact with the vehicle, repeatedly loitered, or detected low-level abnormal sounds, then the threat level is determined to be level three. If, based on the multidimensional behavioral intent results, at least two of the following are identified: the target is holding a suspicious tool, is at a distance less than a preset distance from the vehicle and is stationary, and an abnormal sound is detected, then the threat level is determined to be level four. And / or, the differentiated early warning and response strategy triggered based on the threat level includes: If the threat level is Level 1, then no alarm notification will be sent to the user, and short video logs will be selectively saved locally. If the threat level is level two, the event will be recorded locally and marked in the log, but no alarm will be pushed proactively. If the threat level is level three, then high-definition video recording and encrypted storage will be initiated, and a mild alert notification will be sent to the user terminal. If the threat level is level four, then it will immediately push the highest priority alarm containing the real-time video stream to the user terminal, activate the vehicle's audio and visual alarm system, or automatically upload key video clips to the cloud server for security storage.
9. A vehicle sentry mode threat warning system, characterized in that, include: The multimodal data acquisition module is used to acquire multimodal perception data collected from the area around the vehicle in sentry mode; An adaptive backbone network module is used to extract key features of each modality in the multimodal sensing data through an adaptive backbone network to obtain a multimodal feature sequence. A cross-modal feature interaction fusion module is used to receive the multimodal feature sequence and generate a high-level semantic feature sequence containing threat behavior information; The high-level semantic reasoning module is used to receive the high-level semantic feature sequence, and to perform hybrid decision-making using a causal graph-based propagation algorithm and a symbolic behavior rule knowledge base, and output a multi-dimensional behavior intent result. The threat assessment and response module is used to perform quantitative threat assessment based on the multidimensional behavioral intent results, output multiple levels of threat levels, and trigger differentiated early warning and response strategies based on the threat levels.
10. A vehicle, characterized in that, The vehicle includes: a vehicle body and a vehicle controller; the vehicle controller is used to implement the vehicle sentry mode threat warning method as described in any one of claims 1-8.