A car security method and device based on multi-modal perception and distributed AI

The automotive security system, which utilizes multimodal perception and distributed AI, achieves hierarchical perception, multimodal data acquisition, and cloud-based recognition. This solves the problems of insufficient computing power and slow response in traditional automotive security systems, thereby improving the accuracy and timeliness of automotive security.

CN120902673BActive Publication Date: 2026-07-21BEI DOU ZHI LIAN KE JI YOU XIAN GONG SI +1
View PDF 2 Cites 0 Cited by

Patent Information

Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
BEI DOU ZHI LIAN KE JI YOU XIAN GONG SI
Filing Date
2025-09-15
Publication Date
2026-07-21

AI Technical Summary

Technical Problem

Traditional automotive security systems have limited computing power, are passive in identification, have slow response times, cannot effectively protect vehicle safety, and lack proactive and differentiated response mechanisms.

Method used

Employing a multimodal perception and distributed AI approach, the system utilizes millimeter-wave radar and inertial measurement units for initial monitoring through hierarchical perception, multimodal data acquisition, and distributed computing power. It then activates panoramic cameras and vehicle-mounted microphones to record audio and video data, which is uploaded to a cloud computing center for threat identification. Based on the identification results, a hierarchical response is initiated.

Benefits of technology

It enables low-power sensors to operate on demand, multimodal data acquisition to provide a foundation for accurate identification, distributed computing power to reduce the transmission pressure between the vehicle and the cloud, and high-precision identification and matching of risk levels and warning actions in the cloud. It solves the problems of insufficient computing power and delayed response in the traditional mode, and improves the accuracy and timeliness of security.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120902673B_ABST
    Figure CN120902673B_ABST
Patent Text Reader

Abstract

The application provides a vehicle security method and device based on multi-modal perception and distributed AI, which comprises the following steps: after the vehicle is powered off, a primary sensor comprising a millimeter wave radar and an inertial measurement unit (IMU) is started, the millimeter wave radar monitors the approaching of objects around the vehicle within a preset detection distance, and the IMU monitors the displacement of the vehicle in X and Y directions; when the millimeter wave radar detects that a target is approaching, a secondary sensor comprising a panoramic camera and a vehicle-end microphone is woken up; if the IMU detects that the vehicle is vibrating, the audio and video recording data of the secondary sensor within a specific time period before and after the vibration is acquired; feature extraction is performed on the recording data to generate a feature vector file, which is uploaded to a cloud computing center, threat behavior recognition is completed by the cloud based on AI, and a threat recognition result is output; after the result is received, a matched risk level and a corresponding warning action are determined; and the application can improve the accuracy, timeliness and initiative of vehicle security.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] This application relates to the field of automotive security technology, and in particular to an automotive security method and device based on multimodal perception and distributed AI. Background Technology

[0002] In the field of automotive security, the traditional car sentry mode usually relies on a limited number of sensors on the vehicle side (such as cameras, vibration sensors, etc.) for monitoring, which has many shortcomings.

[0003] On the one hand, the limited computing power of vehicles makes it difficult to perform high-precision threat identification on the large amounts of audio and video data collected, easily leading to false alarms and missed alarms. For example, non-threatening behaviors such as wind blowing the vehicle or animals approaching may be misjudged as threats, or real acts of violence, vandalism, or theft may be not identified in a timely manner. On the other hand, traditional models are mostly passive recording methods, lacking proactive and differentiated response mechanisms. They cannot take corresponding warning measures based on the severity of the threat, either causing excessive alarms that disturb users and the surrounding environment, or failing to respond adequately when serious threats occur, thus failing to effectively protect vehicle safety. In addition, traditional models lack distributed computing power, requiring vehicles to handle high-computation identification tasks, resulting in high resource consumption and an inability to utilize the powerful computing power of the cloud for more accurate model training and iteration, making it difficult to adapt to constantly emerging new threats. Therefore, it is urgent to improve the accuracy, timeliness, and proactivity of automotive security systems. Summary of the Invention

[0004] In view of this, the embodiments of this application provide a vehicle security method and device based on multimodal perception and distributed AI, which can solve the pain points of limited computing power, passive identification and delayed response of traditional vehicle sentry mode through hierarchical perception, multimodal acquisition, distributed computing power division and hierarchical response, and realize the upgrade from "passive recording" to "active security".

[0005] The technical solution of this application embodiment is implemented as follows: In a first aspect, embodiments of this application provide a vehicle security method based on multimodal perception and distributed AI, the method comprising: In response to the vehicle power failure, the vehicle's primary sensors are activated; wherein, the primary sensors include a millimeter-wave radar and an inertial measurement unit (IMU), the millimeter-wave radar continuously monitors whether any objects are approaching the vehicle's vicinity according to a preset detection distance, and the IMU continuously monitors the vehicle's displacement in the X and Y directions; In response to the millimeter-wave radar detecting an approaching target, the vehicle's secondary sensors are activated. If the IMU detects vibration in the vehicle, the recorded data from the secondary sensors within a specific time period before and after the vibration is acquired. The secondary sensors include a panoramic camera and a vehicle-mounted microphone. The panoramic camera records video data, and the vehicle-mounted microphone records audio data. Feature extraction is performed on the recorded data to generate a feature vector file, and the feature vector file is uploaded to a cloud computing center so that the cloud computing center can perform threat behavior identification on the feature vector file and obtain threat identification results; wherein, the cloud computing center performs threat behavior identification based on AI; Receive the threat identification results issued by the cloud computing center, and determine the risk level and corresponding warning action that match the threat identification results based on the threat identification results.

[0006] Secondly, embodiments of this application also provide an automotive security device based on multimodal perception and distributed AI, the device comprising: A startup module is used to activate the vehicle's primary sensors in response to the vehicle being powered off. The primary sensors include a millimeter-wave radar and an inertial measurement unit (IMU). The millimeter-wave radar continuously monitors whether any objects are approaching the vehicle's vicinity at a preset detection distance, and the IMU continuously monitors the vehicle's displacement in the X and Y directions. The wake-up module is used to wake up the secondary sensors on the vehicle in response to the millimeter-wave radar detecting the approach of a target. If the IMU detects that the vehicle is vibrating, it acquires the recorded data of the secondary sensors within a specific time period before and after the vibration. The secondary sensors include a panoramic camera and a vehicle-end microphone. The panoramic camera records video data, and the vehicle-end microphone records audio data. The upload module is used to extract features from the recorded data, generate a feature vector file, and upload the feature vector file to the cloud computing center, so that the cloud computing center can perform threat behavior identification on the feature vector file and obtain threat identification results; wherein, the cloud computing center performs threat behavior identification based on AI; The receiving module is used to receive the threat identification results issued by the cloud computing center, and determine the risk level and corresponding warning action that match the threat identification results based on the threat identification results.

[0007] Thirdly, embodiments of this application also provide an electronic device, including: a processor, a storage medium, and a bus, wherein the storage medium stores machine-readable instructions executable by the processor, and when the electronic device is running, the processor communicates with the storage medium via the bus, and the processor executes the machine-readable instructions to perform the automotive security method based on multimodal perception and distributed AI as described in any of the first aspects.

[0008] Fourthly, embodiments of this application also provide a computer-readable storage medium storing a computer program, which, when executed by a processor, performs the vehicle security method based on multimodal perception and distributed AI as described in any one of the first aspects.

[0009] The embodiments of this application have the following beneficial effects: Through a tiered sensor wake-up mechanism, low-power, on-demand operation of vehicle-side sensors is achieved, avoiding resource waste. Multimodal data acquisition preserves a complete "visual + auditory" evidence chain of threat events, providing a foundation for accurate identification. Distributed computing power division of labor allows the vehicle-side to perform only lightweight feature extraction, reducing the transmission pressure between the vehicle and the cloud, while the cloud utilizes high computing power to achieve high-precision identification. After receiving the cloud identification results, risk levels and warning actions are matched, providing a basis for subsequent differentiated deterrence. This effectively solves the pain points of limited computing power, passive identification, and delayed response in the traditional vehicle sentry mode, achieving an upgrade from "passive recording" to "proactive security." Attached Figure Description

[0010] To more clearly illustrate the technical solutions of the embodiments of this application, the accompanying drawings used in the embodiments will be briefly introduced below. It should be understood that the following drawings only show some embodiments of this application and should not be regarded as a limitation of the scope. For those skilled in the art, other related drawings can be obtained based on these drawings without creative effort.

[0011] Figure 1 This is a flowchart illustrating steps S101-S104 provided in the embodiments of this application; Figure 2 This is a flowchart illustrating steps S201-S204 provided in the embodiments of this application; Figure 3 This is a flowchart illustrating steps S301-S305 provided in the embodiments of this application; Figure 4 This is a flowchart illustrating steps S401-S402 provided in the embodiments of this application; Figure 5 This is a flowchart illustrating steps S501-S502 provided in the embodiments of this application; Figure 6This is a schematic diagram provided in the embodiments of this application; Figure 7 This is a system architecture diagram provided in the embodiments of this application; Figure 8 This is a schematic diagram of the structure of an automotive security device based on multimodal perception and distributed AI provided in an embodiment of this application; Figure 9 This is a schematic diagram of the composition structure of the electronic device provided in the embodiments of this application. Detailed Implementation

[0012] To make the objectives, technical solutions, and advantages of the embodiments of this application clearer, the technical solutions of the embodiments of this application will be clearly and completely described below with reference to the accompanying drawings. It should be understood that the accompanying drawings in this application are for illustrative and descriptive purposes only and are not intended to limit the scope of protection of this application. Furthermore, it should be understood that the schematic drawings are not drawn to scale. The flowcharts used in this application illustrate operations implemented according to some embodiments of this application. It should be understood that the operations in the flowcharts may not be implemented in sequence, and steps without logical contextual relationships may be reversed or implemented simultaneously. In addition, those skilled in the art, guided by the content of this application, may add one or more other operations to the flowcharts, or remove one or more operations from the flowcharts.

[0013] In the following description, references are made to “some embodiments,” which describe a subset of all possible embodiments. However, it is understood that “some embodiments” may be the same subset or different subsets of all possible embodiments and may be combined with each other without conflict.

[0014] Furthermore, the described embodiments are merely some, not all, of the embodiments of this application. The components of the embodiments of this application described and illustrated herein can typically be arranged and designed in various different configurations. Therefore, the following detailed description of the embodiments of this application provided in the accompanying drawings is not intended to limit the scope of the claimed application, but merely to illustrate selected embodiments of the application. All other embodiments obtained by those skilled in the art based on the embodiments of this application without inventive effort are within the scope of protection of this application.

[0015] In the following description, the terms "first, second, third" are used merely to distinguish similar objects and do not represent a specific ordering of objects. It is understood that "first, second, third" may be interchanged in a specific order or sequence where permitted, so that the embodiments of this application described herein can be implemented in an order other than that illustrated or described herein.

[0016] It should be noted that the term "comprising" will be used in the embodiments of this application to indicate the presence of the features declared thereafter, but does not exclude the addition of other features.

[0017] Unless otherwise defined, all technical and scientific terms used herein have the same meaning as commonly understood by one of ordinary skill in the art to which this application belongs. The terminology used herein is for the purpose of describing embodiments of this application and is not intended to limit this application.

[0018] See Figure 1 , Figure 1 This is a flowchart illustrating steps S101-S104 of the vehicle security method based on multimodal perception and distributed AI provided in this application embodiment, which will be combined with... Figure 1 Steps S101-S104 are explained below.

[0019] The embodiments of this application aim to solve the pain points of the traditional car sentry mode, namely "limited computing power, passive identification, and delayed response". The core principle revolves around the closed-loop logic of "hierarchical perception - data collection - cloud identification - hierarchical response".

[0020] In step S101, in response to the vehicle being powered off, the vehicle's primary sensors are activated; wherein, the primary sensors include a millimeter-wave radar and an inertial measurement unit (IMU), the millimeter-wave radar continuously monitors whether any objects are approaching the vehicle's vicinity according to a preset detection distance, and the IMU continuously monitors the vehicle's displacement in the X and Y directions.

[0021] As an example of "hierarchical perception," the primary sensors are millimeter-wave radar and IMU (Inertial Measurement Unit), while the secondary sensors are panoramic cameras. This embodiment continues this hierarchical design. Only the primary sensors are activated after the vehicle is powered off. The core principle is low-power continuous monitoring: the millimeter-wave radar can scan the vehicle's surroundings in real time at a preset detection distance of 3-5 meters, avoiding the waste of storage resources caused by the continuous operation of the secondary sensors (panoramic cameras); the IMU focuses on monitoring the vehicle's displacement in the X and Y directions (such as vibrations caused by collisions or prying), achieving a preliminary warning combining static and dynamic elements. This ensures that the high-power secondary sensors are only activated when a potential threat exists, balancing security needs with vehicle range.

[0022] In step S102, in response to the millimeter-wave radar detecting the approach of a target, the secondary sensor on the vehicle is activated. If the IMU detects vibration of the vehicle, the recorded data of the secondary sensor within a specific time period before and after the vibration is acquired. The secondary sensor includes a panoramic camera and a vehicle-mounted microphone. The panoramic camera records video data, and the vehicle-mounted microphone records audio data.

[0023] As an example of "data acquisition," when a primary sensor is triggered (an object approaches / vibrates), a secondary sensor (panoramic camera + vehicle microphone) is activated to simultaneously collect audio and video data. Relying solely on the camera and vibration sensor can easily lead to misjudgments due to the limited data dimensions (such as a false alarm caused by wind blowing on the vehicle). However, simultaneous audio and video acquisition can retain dual evidence of "visual + auditory": video records behavioral patterns (such as whether a person is holding tools or making a lock-picking motion), and audio captures characteristic sounds (such as the high-frequency sound of breaking glass or the low-frequency sound of metal rubbing), providing a complete data foundation for subsequent feature extraction and threat identification, and reducing misjudgments caused by missing data.

[0024] In step S103, feature extraction is performed on the recorded data to generate a feature vector file, and the feature vector file is uploaded to the cloud computing center so that the cloud computing center can perform threat behavior identification on the feature vector file and obtain threat identification results; wherein, the cloud computing center performs threat behavior identification based on AI.

[0025] As an example of "cloud-based recognition," the distributed design of "edge computing center (ADAS) + cloud computing center" enables the rational allocation of computing resources. The vehicle end (ADAS or CDC) only extracts features from the audio and video data, rather than directly uploading the original audio and video, which greatly reduces the bandwidth pressure of vehicle-cloud communication and solves the problem of "limited computing power on the vehicle end makes it impossible to make accurate judgments" in the traditional mode. The cloud computing center deploys distributed AI and uses the powerful computing power of the cloud to perform high-precision threat identification on feature vectors, realizing efficient collaboration between "lightweight processing at the edge end + high-computing power recognition in the cloud".

[0026] In step S104, the threat identification result issued by the cloud computing center is received, and the risk level and corresponding warning action matching the threat identification result are determined based on the threat identification result.

[0027] As an example of "tiered response", this application's embodiments divide threat levels into multiple levels, with different levels corresponding to different warning actions. By describing "matching risk levels with warning actions", it provides a framework for subsequent specific execution steps, ensuring that the response is upgraded from "passive recording" to "proactive security", avoiding excessive alarms caused by minor behaviors (such as copying signs) and responding quickly to violent damage (such as smashing windows).

[0028] In some embodiments, see Figure 2 , Figure 2 This is a flowchart illustrating steps S201-S204 provided in the embodiments of this application. Features can be extracted from the audio data in the recorded data through steps S201-S204, and the following will be explained in conjunction with each step.

[0029] In step S201, the audio file is read to obtain the audio digital signal and sampling rate.

[0030] In step S202, based on the audio digital signal and sampling rate, 20 MFCC coefficients are extracted using the Mel Frequency Cepstral Coefficient (MFCC) algorithm to obtain MFCC feature data.

[0031] In step S203, the MFCC feature data is converted into decibel units and normalized to obtain normalized MFCC features.

[0032] In step S204, the normalized MFCC features are used to generate feature vectors, forming the audio feature vector portion of the feature vector file.

[0033] Here, the core of this application's embodiment lies in addressing audio-related threats (such as glass breakage and metal friction). The principle is to transform "analog sound signals that cannot be directly calculated" into "standardized digital features that can be recognized by machines," thus solving the problem of "low accuracy in audio threat recognition" in traditional modes.

[0034] First, the audio file is read using the librosa library (the core algorithm library for the document), converting the analog signal into a discrete digital sequence (y is the sound intensity data, sr is the number of samples per second, such as 44100Hz). This is a prerequisite for audio feature extraction, ensuring that subsequent processing is based on a data format that the computer can recognize.

[0035] The human ear has different sensitivities to different frequencies of sound (more sensitive to frequencies of 2-5kHz, where the sounds of breaking glass and metal friction are mostly concentrated). The MFCC algorithm simulates this characteristic through the Mel filter bank, which can accurately capture the "frequency distribution characteristics" of threatening audio (such as the high-frequency peak of the sound of breaking glass and the low-frequency oscillation of the collision sound). In this embodiment, 20 MFCC coefficients are selected because 20 coefficients can cover the core features of threatening audio while avoiding computational redundancy caused by too many coefficients (too many coefficients will increase the pressure on cloud model training and inference).

[0036] The volume of the same threatening audio (such as the sound of breaking a window) varies greatly in different environments (quiet neighborhood vs. noisy street). If the original MFCC features are used directly, the model may misjudge "a soft sound of breaking a window in a noisy environment" as normal sound, or misjudge "loud talking in a quiet environment" as a threat. By converting to decibels (quantifying sound intensity) and normalizing (mapping feature values ​​to a uniform range), volume and environmental interference can be eliminated, ensuring that the feature scale of the same type of threatening audio is consistent, and providing standardized features for cloud identification.

[0037] Cloud-based AI models have specific requirements for the input data format (usually vectors or matrices). By integrating the normalized MFCC features into a vector file, we can ensure that the data format is compatible with the cloud model input, avoid recognition failures due to format mismatch, and make the vector file smaller in size, which is convenient for vehicle-to-cloud transmission and further optimizes communication efficiency.

[0038] In some embodiments, see Figure 3 , Figure 3 This is a flowchart illustrating steps S301-S305 provided in the embodiments of this application. The cloud computing center identifies threat behaviors based on a CNN+Transformer deep learning model. The CNN+Transformer deep learning model can be trained through steps S301-S305, and will be explained in conjunction with each step.

[0039] In step S301, an MFCC dataset class is constructed, and a CSV file containing MFCC features and corresponding labels is loaded. The labels are used to identify the behavior category corresponding to the audio.

[0040] In step S302, the data in the MFCC dataset is preprocessed, including reshaping the MFCC features into a preset shape and performing standardization.

[0041] In step S303, a CNN+Transformer hybrid model is built. The model includes a CNN module, a fully connected layer, a Transformer module, and a classification layer. The CNN module is used to extract local features. The fully connected layer is used to convert the features output by the CNN module into dimensions that are compatible with the Transformer module. The Transformer module is used to capture temporal correlation features. The classification layer is used to output the behavior category prediction results.

[0042] In step S304, training parameters are set, including batch size, learning rate, and training epochs. The prediction loss is calculated using the cross-entropy loss function, and the model parameters are updated using the Adam optimizer. The CNN+Transformer hybrid model is then trained.

[0043] In step S305, the trained model parameters are saved to obtain a CNN+Transformer deep learning model that can be used for threat behavior recognition.

[0044] Here, model training requires "supervised samples" (i.e., behavior categories corresponding to known features, such as "a certain MFCC feature - glass breakage" or "a certain MFCC feature - normal speech"). By reading CSV files through the MFCC dataset class (a document-customized Dataset class), features are associated with labels to provide "learning materials" for the model, ensuring that the model can master the characteristic patterns of different threat behaviors through sample training.

[0045] On the one hand, CNN models require multi-channel features as input, and reshaping the feature shape can ensure adaptation to the CNN input format; on the other hand, the MFCC feature scale of different samples may differ (e.g., the MFCC coefficient range of different audios is different). By standardization (preprocess function, (data - mean) / standard deviation), the features can be mapped to a uniform scale, avoiding model training bias towards a certain type of sample due to feature bias (e.g., large-scale features dominate training), and ensuring training fairness.

[0046] One of the core innovations of this application is the "CNN+Transformer hybrid model," which is based on the principle of complementing the advantages of the two models: (1) CNN module: Extracting local detail features: CNN excels at extracting local details from data (such as the frequency peak of a certain time period in MFCC features, corresponding to the instantaneous high-frequency features of the window smashing sound and the low-frequency oscillation features of the collision sound). It extracts basic local features through the first convolutional layer, reduces the feature size (reducing the amount of computation) through the pooling layer, and further extracts more complex local features through the second convolutional layer, thus solving the problem of "insufficient capture of local threat features" in traditional models (such as the inability to distinguish the differences in local features between "wind blowing" and "minor collision").

[0047] (2) Transformer module: Capturing temporal correlation features: Threatening behaviors often exhibit temporal correlations (e.g., "lingering for more than 30 seconds" is a continuous temporal behavior, while "multiple collisions" have time interval characteristics). CNNs cannot effectively capture this type of temporal information, while Transformers, through multi-head attention mechanisms, can focus on the correlation of MFCC features at different time steps (e.g., if a certain frequency feature appears multiple times within 10 seconds, it may correspond to continuous lock-picking actions), making up for the shortcomings of CNNs in recognizing "temporal dimension features" and solving the problem that "the nature of the behavior cannot be determined by local features alone" (e.g., a single slight vibration may be due to wind, while multiple vibrations may be intentional sabotage).

[0048] (3) Fully connected layer and classification layer: to achieve feature dimension adaptation and category output: The fully connected layer converts the high-dimensional local features output by the CNN into features consistent with the dimensions of the hidden layers of the Transformer (e.g., 256 dimensions for a document), ensuring that the data can be successfully input into the Transformer; the classification layer maps the temporal features processed by the Transformer into "behavior category probabilities" (e.g., 95% probability of glass breakage, 3% probability of normal vibration), directly outputting threat identification results and providing a basis for distributing identification results to the cloud.

[0049] (4) Training parameters and optimization: Ensure the model converges to the optimal state: The batch size is set to 32 because it ensures the diversity of training samples each time (avoiding model overfitting) without causing insufficient computing resources due to an excessively large batch size. The learning rate of 0.001 is a commonly used initial value in deep learning, which can avoid model oscillation caused by an excessively high learning rate and slow training caused by an excessively low learning rate. The cross-entropy loss function is suitable for multi-classification tasks (such as distinguishing between broken glass, collision, and normal sound), and can accurately calculate the difference between the "predicted category" and the "true label". The Adam optimizer can dynamically adjust the learning rate, accelerate model convergence, and ensure that the trained model has high recognition accuracy.

[0050] In some embodiments, the AI ​​training data import method of the cloud computing center includes importing with the original settings and importing through an upgraded feature library. The upgraded feature library is used to supplement new threat behavior feature data to improve the accuracy of threat behavior identification.

[0051] Here, AI imports training data on common threat behaviors (such as common features like broken glass, vehicle collisions, and lock picking) through the original settings, ensuring that Sentinel Mode immediately has basic threat recognition capabilities after the vehicle is delivered to the user, meeting the needs of most daily scenarios (such as common threats in residential areas and parking lots), and avoiding the problem of "waiting for subsequent data imports before the solution can be used after implementation".

[0052] In real-world applications, new threat behaviors may emerge (such as the friction sound of new lock-picking tools, or threat behavior characteristics under special weather conditions (such as the sound of windows being smashed during a rainstorm)). If the model relies solely on the original factory data, it cannot identify these new threats, causing the solution to fail. By upgrading the feature library (e.g., automakers regularly push new threat feature CSV files to the cloud), these new threat features are added to the model's training data. This eliminates the need to retrain the entire model (only incremental training is required), enabling the model to identify new threats and achieving "lifetime model iteration," ensuring the long-term effectiveness of the solution.

[0053] In some embodiments, the cloud computing center pre-records and calibrates the sound features of collisions and glass breakage to form an initial sound feature library, and supports updating the initial sound feature library to optimize the recognition accuracy of audio-based threat behaviors.

[0054] Here, different types of audio threats (such as glass breaking and metal collisions) have their unique "frequency fingerprints." Before leaving the factory, standard samples of these audio types are collected in the cloud using professional equipment for feature calibration (such as determining the MFCC coefficient range of glass breaking sounds and the frequency peak range of collision sounds) to form an initial sound feature library. When the vehicle uploads the audio feature vector, the cloud can compare it with the "standard template" in the feature library to quickly determine whether it is a threatening audio, solving the problem of "misjudgment due to lack of standard reference" in the traditional mode (such as misjudging "bottle cap falling sound" as a collision sound).

[0055] On the one hand, audio interference varies greatly in different regions and environments (such as the sound of cold wind in winter in the north and the sound of rain in the rainy season in the south). The initial feature library may not be able to cover the audio threat features in these environments. By updating the feature library and supplementing it with "environment-adaptive features" (such as the feature of "the sound of glass breaking in the rain"), misjudgments caused by environmental interference can be reduced. On the other hand, new audio threats (such as the sound of new window-breaking tools) will continue to emerge. Updating the feature library can incorporate these new threat features into the "standard template" to ensure that the model can identify new types of audio threats and continuously optimize the recognition accuracy of audio threats.

[0056] In some embodiments, see Figure 4 , Figure 4 This is a flowchart illustrating steps S401-S402 provided in the embodiments of this application. The method further includes steps S401-S402, which will be explained in conjunction with each step.

[0057] In step S401, if the identified threat behavior is misjudged, the feature data corresponding to the misjudged case is collected.

[0058] In step S402, the calibration data is updated based on the feature data corresponding to the misjudged case, and the AI ​​is trained based on the updated calibration data to correct the misidentification.

[0059] In practical applications, models may misjudge situations due to insufficient training data (such as not covering normal behavior in certain special environments). For example, misjudging "a cleaner touching a car body with a broom" as a collision, or "a tree branch hitting a car window during a rainstorm" as breaking the window. By collecting feature data of these misjudgment cases (such as the MFCC audio features and video feature vectors of the behavior), it is possible to identify in which scenarios and under which features the model will misjudge, providing targeted goals for subsequent corrections and avoiding situations where "blind corrections cannot solve the problem."

[0060] The feature data of misclassified cases is relabeled (e.g., the feature of "broom touching the car body" is corrected from "threat" to "normal") and added to the model's training dataset. Then, the distributed AI (such as a CNN+Transformer model) is incrementally trained (without retraining the entire model, only training on the newly added labeled data). This allows the model to relearn the "correct feature-category correspondence" and eliminate previous blind spots in recognition. For example, if the model misclassifies "tree branches hitting the car window during a rainstorm" as window smashing, by adding feature data for this behavior and relabeling it as "normal," the trained model can accurately distinguish the feature differences between "tree branches hitting" and "window smashing," significantly reducing the misclassification rate for the same type of case.

[0061] In some embodiments, see Figure 5 , Figure 5 This is a flowchart illustrating steps S501-S502 provided in the embodiments of this application. The risk level and corresponding warning action that match the threat identification result are determined based on the threat identification result. This can be achieved through steps S501-S502, which will be explained in conjunction with each step.

[0062] In step S501, the risk level and the corresponding warning action instruction are obtained.

[0063] In step S502, the vehicle's lighting system and external voice system are controlled to perform warning actions according to the warning action command; if the risk level is Level 3, the risk event is pushed to the owner's APP; wherein, in the threat level correspondence table, the risk levels include Level 1, Level 2 and Level 3; when the behavior is to approach the vehicle and stay for a set time, the risk level is Level 1, and the warning action is hazard lights; when the behavior is to linger near the vehicle for more than a set time, the risk level is Level 2, and the warning action is hazard lights and playing a voice warning once; when the behavior is glass breakage or collision, the risk level is Level 3, and the warning action is hazard lights, horn, and event push to the owner's APP.

[0064] Here, minor behaviors (such as a passerby taking a picture or briefly taking a photo) do not require a high-intensity response. Level 1 only requires flashing the hazard lights to alert the perpetrator that they are being monitored, avoiding noise disturbances such as honking. Moderate behaviors (such as loitering for more than 30 seconds, which may indicate theft intent) require stronger deterrence. Level 2 uses a voice message, "You have entered the monitored area," to clearly inform the boundaries of the behavior and prevent escalation. Severe behaviors (such as smashing windows or collisions that have caused or are about to cause damage to the vehicle) require an emergency response. Level 3 uses honking to deter the perpetrator (such as forcing the perpetrator to stop) and simultaneously pushes a notification to the owner's app (as described in the document, "The owner's app receives a threat alert and takes necessary action") to ensure that the owner is aware in real time and can take measures (such as remotely viewing video, contacting property management, or calling the police). This avoids user resentment caused by excessive alarms and allows for rapid handling of serious threats.

[0065] The BDCU (Body Control Unit) is the direct control unit for the vehicle's lighting and audio systems. The CDC (Vehicle Control Unit) sends commands to the BDCU via the CAN bus (Automotive Internal Communication Bus) to ensure rapid execution of lights (hazard lights) and voice / horn (external audio). The CDC also undertakes the task of pushing messages to the owner's app ("The CDC pushes risk events to the owner's app" in the document). Because the CDC has communication links with the vehicle's infotainment system and the cloud, it can push threat events (including audio and video clips) to the app in real time, achieving dual protection of "vehicle-side deterrence + remote user interaction" and completely solving the passive problem of the traditional mode of "only recording and not responding".

[0066] Below, we will combine Figure 6 , Figure 7 The embodiments of this application will be fully explained and described. Figure 6 This is a schematic diagram provided in the embodiments of this application. Figure 7 This is a system architecture diagram provided in the embodiments of this application, such as... Figure 6 , Figure 7 As shown, the edge of the vehicle's infotainment system is the "perception front end," responsible for multi-source data acquisition and feature extraction. An AVM (Around View Camera), RADAR (Millimeter-Wave Radar), IMU (Inertial Measurement Unit), and microphone array work together. RADAR and IMU, as primary low-power sensors, continuously monitor the proximity of objects within a 3-5m radius around the vehicle (RADAR) and the vehicle's X / Y displacement (IMU, such as vibrations caused by collisions or prying). When a primary sensor is triggered (an object approaches or vibration is detected), the AVM and microphone array (secondary sensors) are activated, simultaneously acquiring audio and video data from the vehicle's surroundings, retaining multi-dimensional "visual + auditory" information about threat events.

[0067] ADAS (Advanced Driver Assistance Systems, acting as an edge computing center) extracts features from the collected audio and video data, generating a lightweight feature library (stored in CSV file format). Simultaneously, ADAS can preliminarily determine the threat level (e.g., Level 1 to Level 3 corresponding to slight approach, loitering, or violent destruction) based on preset rules. If a threat is detected, it directly triggers audio and visual deterrence (by controlling the external voice and lighting systems of the warning system via the BDCU to execute actions such as hazard lights, horn blaring, and voice warnings). For more precise identification, the feature library file is uploaded to the cloud computing center via T-BOX (vehicle-to-everything terminal). Furthermore, DVR (Dashcam) is responsible for storing the original audio and video data, providing evidence for subsequent review.

[0068] The cloud-based identification and model computing center serves as the system's "intelligent brain," undertaking high-computing-power-demand threat identification and model optimization tasks. After receiving the CSV feature library file uploaded from the vehicle's edge, it utilizes the deployed CNN+Transformer deep learning model to perform high-precision analysis of the feature data. The CNN module extracts local features (such as frequency peaks in audio and intra-frame object contours in video), while the Transformer module captures temporal correlations (such as the time pattern of people loitering and the rhythm of continuous collisions). Combining the advantages of both, it accurately identifies threatening behaviors (such as distinguishing between "animal approach" and "people picking locks").

[0069] The cloud-based maintenance system has an initial feature library (threat features such as collisions and glass breakage imported before leaving the factory) and supports updating the feature library via notifications from vehicle manufacturers to continuously optimize the model. After identifying a threat behavior, the system sends the threat type to the CDC (vehicle control unit) at the edge of the vehicle's infotainment system to assist in local threat level determination and response. On the other hand, it records threat events to achieve cross-regional and multi-vehicle joint defense and linkage (e.g., when multiple vehicles in a certain area encounter the same threat, the cloud can trigger a cluster warning).

[0070] The mobile terminal serves as the "remote security entry point" for car owners. Through the owner's app and system linkage, when a serious threat (such as glass breakage or violent collision) is detected by the cloud or the vehicle's edge detection system, a notification will be pushed to the owner's app via T-BOX. This notification includes the time of the threat event, preliminary characteristics, and audio / video clips, ensuring that the owner is aware of any vehicle abnormalities in real time. After receiving the notification, the owner can remotely view real-time or historical surveillance footage of the vehicle's surroundings through the app and take appropriate measures, such as contacting property management or calling the police, achieving dual security through "vehicle-side deterrence + remote user linkage."

[0071] This system architecture, through a distributed architecture of "lightweight processing at the vehicle's edge + high-performance cloud-based recognition + user interaction with mobile terminals," not only solves the problems of low recognition accuracy and high false alarm rate caused by insufficient vehicle-side computing power in the traditional model, but also achieves a security upgrade from "passive recording" to "active deterrence + user participation." At the same time, it supports dynamic updates of the feature library and model iteration, ensuring that the system can adapt to new threats and complex scenarios in the long term, providing vehicles with all-weather, high-precision intelligent security protection.

[0072] In summary, the embodiments of this application have the following beneficial effects: Compared to the traditional vehicle sentry mode, the embodiments of this application have significant and comprehensive technical advantages and application value: First, through the hierarchical perception design of "low-power continuous monitoring by primary sensors (millimeter-wave radar + IMU) and on-demand wake-up of secondary sensors (panoramic camera + microphone)," it avoids the power consumption and storage waste caused by continuous sensor operation in the traditional mode, and can retain a complete threat evidence chain through audio and video multimodal data collection, laying a data foundation for accurate identification; Second, relying on the distributed computing power architecture of "feature vector extraction by vehicle-side edge computing power center (ADAS or CDC) + threat identification by cloud computing power center CNN + Transformer hybrid model," it effectively solves the pain points of limited vehicle-side computing power and low identification accuracy in the traditional mode. Lightweight processing at the edge end significantly reduces the bandwidth pressure of vehicle-to-cloud transmission, and the cloud model combines CNN local feature capture and Transformer temporal correlation analysis. The system boasts several key features: First, it can accurately distinguish between non-threatening events such as rustling grass and approaching animals, and threatening behaviors such as people loitering and broken glass, significantly reducing false alarm rates. Second, it matches differentiated warning actions (from hazard lights to horn + vehicle APP push) based on threat levels (Level 1 to Level 3), upgrading from traditional "passive recording" to "active deterrence + human-vehicle joint defense." This avoids excessive alarms for minor behaviors that might disturb users, while also enabling rapid response to serious threats, ensuring vehicle safety. Third, it supports dynamic updates to the feature library and correction of misjudgments. By upgrading the feature library to supplement new threat behavior features and updating calibration data, it corrects misidentifications, ensuring that the system's recognition capabilities are continuously iterated and optimized with application scenarios. Furthermore, the solution is compatible with ADAS or CDC as edge computing centers, adapting to various vehicle types such as electric vehicles, high-end fuel vehicles, and shared cars, thus broadening its application scope and comprehensively improving the reliability, timeliness, and user experience of automotive security.

[0073] Based on the same inventive concept, this application also provides a vehicle security device based on multimodal perception and distributed AI, which corresponds to the vehicle security method based on multimodal perception and distributed AI in the first embodiment. Since the principle of the device in this application is similar to the above-mentioned vehicle security method based on multimodal perception and distributed AI, the implementation of the device can refer to the implementation of the method, and the repeated parts will not be described again.

[0074] like Figure 8 As shown, Figure 8 This is a schematic diagram of the structure of a car security device 800 based on multimodal perception and distributed AI provided in an embodiment of this application. The car security device 800 based on multimodal perception and distributed AI includes: The startup module 801 is used to start the vehicle's primary sensors in response to the vehicle being powered off; wherein, the primary sensors include a millimeter-wave radar and an inertial measurement unit (IMU), the millimeter-wave radar continuously monitors whether there are objects approaching the vehicle's vicinity according to a preset detection distance, and the IMU continuously monitors the vehicle's displacement in the X and Y directions; The wake-up module 802 is used to wake up the secondary sensors on the vehicle in response to the millimeter-wave radar detecting that a target is approaching. If the IMU detects that the vehicle is vibrating, the module acquires the recorded data of the secondary sensors within a specific time period before and after the vibration. The secondary sensors include a panoramic camera and a vehicle-end microphone. The panoramic camera records video data, and the vehicle-end microphone records audio data. The upload module 803 is used to extract features from the recorded data, generate a feature vector file, and upload the feature vector file to the cloud computing center, so that the cloud computing center can perform threat behavior identification on the feature vector file and obtain threat identification results; wherein, the cloud computing center performs threat behavior identification based on AI. The receiving module 804 is used to receive the threat identification results issued by the cloud computing center, and determine the risk level and corresponding warning action that match the threat identification results based on the threat identification results.

[0075] Those skilled in the art should understand that Figure 8 The functions of each unit in the multimodal perception and distributed AI-based automotive security device 800 shown can be understood by referring to the relevant description of the aforementioned multimodal perception and distributed AI-based automotive security method. Figure 8 The functions of each unit in the multimodal perception and distributed AI-based automotive security device 800 shown can be implemented through a program running on a processor or through specific logic circuits.

[0076] In one possible implementation, the upload module 803 extracts features from the audio data in the recorded data in the following manner: Read audio files to obtain audio digital signals and sampling rates; Based on the audio digital signal and sampling rate, 20 MFCC coefficients are extracted using the Mel frequency cepstral coefficient (MFCC) algorithm to obtain MFCC feature data. The MFCC feature data is converted into decibel units and normalized to obtain normalized MFCC features. The normalized MFCC features are used to generate feature vectors, forming the audio feature vector portion of the feature vector file.

[0077] In one possible implementation, the cloud computing center identifies threat behaviors based on a CNN+Transformer deep learning model; the CNN+Transformer deep learning model is trained in the following manner: Construct an MFCC dataset class and load a CSV file containing MFCC features and corresponding labels, whereby the labels are used to identify the behavior category corresponding to the audio. The data in the MFCC dataset are preprocessed, including reshaping the MFCC features into a preset shape and performing standardization. A CNN+Transformer hybrid model is constructed, which includes a CNN module, a fully connected layer, a Transformer module, and a classification layer. The CNN module is used to extract local features, the fully connected layer is used to convert the features output by the CNN module into dimensions that are compatible with the Transformer module, the Transformer module is used to capture temporal correlation features, and the classification layer is used to output the behavior category prediction results. Set training parameters, including batch size, learning rate, and training epochs; calculate the prediction loss using the cross-entropy loss function; update the model parameters using the Adam optimizer; and train the CNN+Transformer hybrid model. Save the trained model parameters to obtain a CNN+Transformer deep learning model that can be used for threat behavior recognition.

[0078] In one possible implementation, the AI ​​training data import method of the cloud computing center includes importing with the original settings and importing through an upgraded feature library. The upgraded feature library is used to supplement new threat behavior feature data to improve the accuracy of threat behavior identification.

[0079] In one possible implementation, the cloud computing center pre-records and calibrates the sound features of collisions and glass breakage to form an initial sound feature library, and supports updating the initial sound feature library to optimize the recognition accuracy of audio-based threat behaviors.

[0080] In one possible implementation, the upload module 803 further includes: If a threatening behavior is identified as a misjudgment, the feature data corresponding to that misjudgment case will be collected. The calibration data is updated based on the feature data corresponding to the misjudgment case, and the AI ​​is trained based on the updated calibration data to correct the misidentification.

[0081] In one possible implementation, the receiving module 804 determines a risk level and corresponding warning action matching the threat identification result based on the threat identification result, including: Obtain the risk level and the corresponding warning action instruction; The vehicle's lighting system and external voice system are controlled to perform warning actions according to the warning action command; if the risk level is Level 3, the risk event is pushed to the owner's APP; wherein, the risk level in the threat level correspondence table includes Level 1, Level 2 and Level 3; when the behavior is to approach the vehicle and stay for a set time, the risk level is Level 1, and the warning action is hazard lights; when the behavior is to linger near the vehicle for more than a set time, the risk level is Level 2, and the warning action is hazard lights and playing a voice warning once; when the behavior is glass breakage or collision, the risk level is Level 3, and the warning action is hazard lights, horn, and event push to the owner's APP.

[0082] The aforementioned automotive security device based on multimodal perception and distributed AI has the following beneficial effects: Compared to the traditional vehicle sentry mode, the embodiments of this application have significant and comprehensive technical advantages and application value: First, through the hierarchical perception design of "low-power continuous monitoring by primary sensors (millimeter-wave radar + IMU) and on-demand wake-up of secondary sensors (panoramic camera + microphone)," it avoids the power consumption and storage waste caused by continuous sensor operation in the traditional mode, and can retain a complete threat evidence chain through audio and video multimodal data collection, laying a data foundation for accurate identification; Second, relying on the distributed computing power architecture of "feature vector extraction by vehicle-side edge computing power center (ADAS or CDC) + threat identification by cloud computing power center CNN + Transformer hybrid model," it effectively solves the pain points of limited vehicle-side computing power and low identification accuracy in the traditional mode. Lightweight processing at the edge end significantly reduces the bandwidth pressure of vehicle-to-cloud transmission, and the cloud model combines CNN local feature capture and Transformer temporal correlation analysis. The system boasts several key features: First, it can accurately distinguish between non-threatening events such as rustling grass and approaching animals, and threatening behaviors such as people loitering and broken glass, significantly reducing false alarm rates. Second, it matches differentiated warning actions (from hazard lights to horn + vehicle APP push) based on threat levels (Level 1 to Level 3), upgrading from traditional "passive recording" to "active deterrence + human-vehicle joint defense." This avoids excessive alarms for minor behaviors that might disturb users, while also enabling rapid response to serious threats, ensuring vehicle safety. Third, it supports dynamic updates to the feature library and correction of misjudgments. By upgrading the feature library to supplement new threat behavior features and updating calibration data, it corrects misidentifications, ensuring that the system's recognition capabilities are continuously iterated and optimized with application scenarios. Furthermore, the solution is compatible with ADAS or CDC as edge computing centers, adapting to various vehicle types such as electric vehicles, high-end fuel vehicles, and shared cars, thus broadening its application scope and comprehensively improving the reliability, timeliness, and user experience of automotive security.

[0083] like Figure 9 As shown, Figure 9 This is a schematic diagram of the composition structure of the electronic device 900 provided in the embodiments of this application. The electronic device 900 includes: The device 900 includes a processor 901, a storage medium 902, and a bus 903. The storage medium 902 stores machine-readable instructions executable by the processor 901. When the electronic device 900 is running, the processor 901 communicates with the storage medium 902 via the bus 903. The processor 901 executes the machine-readable instructions to perform the steps of the vehicle security method based on multimodal perception and distributed AI described in the embodiments of this application.

[0084] In practical applications, the various components in the electronic device 900 are coupled together via a bus 903. It is understood that the bus 903 is used to achieve communication between these components. In addition to a data bus, the bus 903 also includes a power bus, a control bus, and a status signal bus. However, for clarity, in... Figure 9 The general designated all buses as Bus 903.

[0085] The above-mentioned electronic devices have the following beneficial effects: Compared to the traditional vehicle sentry mode, the embodiments of this application have significant and comprehensive technical advantages and application value: First, through the hierarchical perception design of "low-power continuous monitoring by primary sensors (millimeter-wave radar + IMU) and on-demand wake-up of secondary sensors (panoramic camera + microphone)," it avoids the power consumption and storage waste caused by continuous sensor operation in the traditional mode, and can retain a complete threat evidence chain through audio and video multimodal data collection, laying a data foundation for accurate identification; Second, relying on the distributed computing power architecture of "feature vector extraction by vehicle-side edge computing power center (ADAS or CDC) + threat identification by cloud computing power center CNN + Transformer hybrid model," it effectively solves the pain points of limited vehicle-side computing power and low identification accuracy in the traditional mode. Lightweight processing at the edge end significantly reduces the bandwidth pressure of vehicle-to-cloud transmission, and the cloud model combines CNN local feature capture and Transformer temporal correlation analysis. The system boasts several key features: First, it can accurately distinguish between non-threatening events such as rustling grass and approaching animals, and threatening behaviors such as people loitering and broken glass, significantly reducing false alarm rates. Second, it matches differentiated warning actions (from hazard lights to horn + vehicle APP push) based on threat levels (Level 1 to Level 3), upgrading from traditional "passive recording" to "active deterrence + human-vehicle joint defense." This avoids excessive alarms for minor behaviors that might disturb users, while also enabling rapid response to serious threats, ensuring vehicle safety. Third, it supports dynamic updates to the feature library and correction of misjudgments. By upgrading the feature library to supplement new threat behavior features and updating calibration data, it corrects misidentifications, ensuring that the system's recognition capabilities are continuously iterated and optimized with application scenarios. Furthermore, the solution is compatible with ADAS or CDC as edge computing centers, adapting to various vehicle types such as electric vehicles, high-end fuel vehicles, and shared cars, thus broadening its application scope and comprehensively improving the reliability, timeliness, and user experience of automotive security.

[0086] This application also provides a computer-readable storage medium storing executable instructions. When the executable instructions are executed by at least one processor 901, the vehicle security method based on multimodal perception and distributed AI described in this application is implemented.

[0087] In some embodiments, the storage medium may be a magnetic random access memory (FRAM), a read-only memory (ROM), a programmable read-only memory (PROM), an erasable programmable read-only memory (EPROM), an electrically erasable programmable read-only memory (EEPROM), a flash memory, a magnetic surface memory, an optical disc, or a compact disc read-only memory (CD-ROM), etc.; or it may be a device that includes one or any combination of the above-mentioned memories.

[0088] In some embodiments, executable instructions may take the form of a program, software, software module, script, or code, written in any form of programming language (including compiled or interpreted languages, or declarative or procedural languages), and may be deployed in any form, including as a standalone program or as a module, component, subroutine, or other unit suitable for use in a computing environment.

[0089] As an example, executable instructions may, but do not necessarily, correspond to files in the file system. They may be stored as part of a file that holds other programs or data, for example, in one or more scripts in a HyperText Markup Language (HTML) document, in a single file dedicated to the program in question, or in multiple collaborating files (e.g., a file that stores one or more modules, subroutines, or code sections).

[0090] As an example, executable instructions can be deployed to execute on a single computing device, or on multiple computing devices located in one location, or on multiple computing devices distributed across multiple locations and interconnected via a communication network.

[0091] The aforementioned computer-readable storage media have the following beneficial effects: Compared to the traditional vehicle sentry mode, the embodiments of this application have significant and comprehensive technical advantages and application value: First, through the hierarchical perception design of "low-power continuous monitoring by primary sensors (millimeter-wave radar + IMU) and on-demand wake-up of secondary sensors (panoramic camera + microphone)," it avoids the power consumption and storage waste caused by continuous sensor operation in the traditional mode, and can retain a complete threat evidence chain through audio and video multimodal data collection, laying a data foundation for accurate identification; Second, relying on the distributed computing power architecture of "feature vector extraction by vehicle-side edge computing power center (ADAS or CDC) + threat identification by cloud computing power center CNN + Transformer hybrid model," it effectively solves the pain points of limited vehicle-side computing power and low identification accuracy in the traditional mode. Lightweight processing at the edge end significantly reduces the bandwidth pressure of vehicle-to-cloud transmission, and the cloud model combines CNN local feature capture and Transformer temporal correlation analysis. The system boasts several key features: First, it can accurately distinguish between non-threatening events such as rustling grass and approaching animals, and threatening behaviors such as people loitering and broken glass, significantly reducing false alarm rates. Second, it matches differentiated warning actions (from hazard lights to horn + vehicle APP push) based on threat levels (Level 1 to Level 3), upgrading from traditional "passive recording" to "active deterrence + human-vehicle joint defense." This avoids excessive alarms for minor behaviors that might disturb users, while also enabling rapid response to serious threats, ensuring vehicle safety. Third, it supports dynamic updates to the feature library and correction of misjudgments. By upgrading the feature library to supplement new threat behavior features and updating calibration data, it corrects misidentifications, ensuring that the system's recognition capabilities are continuously iterated and optimized with application scenarios. Furthermore, the solution is compatible with ADAS or CDC as edge computing centers, adapting to various vehicle types such as electric vehicles, high-end fuel vehicles, and shared cars, thus broadening its application scope and comprehensively improving the reliability, timeliness, and user experience of automotive security.

[0092] In the several embodiments provided in this application, it should be understood that the disclosed methods and electronic devices can be implemented in other ways. The device embodiments described above are merely illustrative. For example, the division of units is only a logical functional division, and in actual implementation, there may be other division methods, such as: multiple units or components may be combined, or integrated into another system, or some features may be ignored or not executed. In addition, the coupling, direct coupling, or communication connection between the various components shown or discussed may be through some interfaces, and the indirect coupling or communication connection between devices or units may be electrical, mechanical, or other forms.

[0093] The modules described as separate components may or may not be physically separate. The components shown as modules may or may not be physical units; that is, they may be located in one place or distributed across multiple network units. Some or all of the units can be selected to achieve the purpose of this embodiment according to actual needs.

[0094] In addition, the functional units in the various embodiments of this application can be integrated into one processing unit, or each unit can exist physically separately, or two or more units can be integrated into one unit.

[0095] If the aforementioned functions are implemented as software functional units and sold or used as independent products, they can be stored in a processor-executable, non-volatile, computer-readable storage medium. Based on this understanding, the technical solution of this application, in essence, or the part that contributes to the prior art, or a portion of the technical solution, can be embodied in the form of a software product. This computer software product is stored in a storage medium and includes several instructions to cause a computer device (which may be a personal computer, a platform server, or a network device, etc.) to execute all or part of the steps of the methods described in the various embodiments of this application. The aforementioned storage medium includes various media capable of storing program code, such as USB flash drives, portable hard drives, ROM, RAM, magnetic disks, or optical disks.

[0096] The above are merely specific embodiments of this application, but the scope of protection of this application is not limited thereto. Any variations or substitutions that can be easily conceived by those skilled in the art within the scope of the technology disclosed in this application should be included within the scope of protection of this application. Therefore, the scope of protection of this application should be determined by the scope of the claims.

Claims

1. A vehicle security method based on multimodal perception and distributed AI, characterized in that, The method includes: In response to the vehicle power failure, the vehicle's primary sensors are activated; wherein, the primary sensors include a millimeter-wave radar and an inertial measurement unit (IMU), the millimeter-wave radar continuously monitors whether any objects are approaching the vehicle's vicinity according to a preset detection distance, and the IMU continuously monitors the vehicle's displacement in the X and Y directions; In response to the millimeter-wave radar detecting an approaching target, the vehicle's secondary sensors are activated. If the IMU detects vibration in the vehicle, the recorded data from the secondary sensors within a specific time period before and after the vibration is acquired. The secondary sensors include a panoramic camera and a vehicle-mounted microphone. The panoramic camera records video data, and the vehicle-mounted microphone records audio data. Feature extraction is performed on the recorded data to generate a feature vector file, and the feature vector file is uploaded to a cloud computing center so that the cloud computing center can perform threat behavior identification on the feature vector file and obtain threat identification results; wherein, the cloud computing center performs threat behavior identification based on AI; Receive threat identification results from the cloud computing center, and determine the risk level and corresponding warning action that match the threat identification results based on the threat identification results; The cloud computing center uses a CNN+Transformer deep learning model for threat behavior identification; the CNN+Transformer deep learning model is trained in the following way: Construct an MFCC dataset class and load a CSV file containing MFCC features and corresponding labels, whereby the labels are used to identify the behavior category corresponding to the audio. The data in the MFCC dataset are preprocessed, including reshaping the MFCC features into a preset shape and performing standardization. A CNN+Transformer hybrid model is constructed, which includes a CNN module, a fully connected layer, a Transformer module, and a classification layer. The CNN module is used to extract local features, the fully connected layer is used to convert the features output by the CNN module into dimensions that are compatible with the Transformer module, the Transformer module is used to capture temporal correlation features, and the classification layer is used to output the behavior category prediction results. Set training parameters, including batch size, learning rate, and training epochs; calculate the prediction loss using the cross-entropy loss function; update the model parameters using the Adam optimizer; and train the CNN+Transformer hybrid model. Save the trained model parameters to obtain a CNN+Transformer deep learning model that can be used for threat behavior recognition.

2. The method according to claim 1, characterized in that, The audio data in the recorded data is feature extracted using the following method: Read audio files to obtain audio digital signals and sampling rates; Based on the audio digital signal and sampling rate, 20 MFCC coefficients are extracted using the Mel frequency cepstral coefficient (MFCC) algorithm to obtain MFCC feature data. The MFCC feature data is converted into decibel units and normalized to obtain normalized MFCC features. The normalized MFCC features are used to generate feature vectors, forming the audio feature vector portion of the feature vector file.

3. The method according to claim 1, characterized in that, The AI ​​training data import methods of the cloud computing center include importing with original settings and importing through an upgraded feature library. The upgraded feature library is used to supplement new threat behavior feature data to improve the accuracy of threat behavior identification.

4. The method according to claim 1, characterized in that, The cloud computing center pre-records and labels the sound features of collisions and glass breakage to form an initial sound feature library, and supports updating the initial sound feature library to optimize the recognition accuracy of audio-based threat behaviors.

5. The method according to claim 1, characterized in that, The method further includes: If a threatening behavior is identified as a misjudgment, feature data corresponding to the misjudgment case will be collected. The calibration data is updated based on the feature data corresponding to the misjudgment case, and the AI ​​is trained based on the updated calibration data to correct the misidentification.

6. The method according to claim 1, characterized in that, Based on the threat identification results, determine the risk level and corresponding warning action that match the threat identification results, including: Obtain the risk level and the corresponding warning action instruction; The vehicle's lighting system and external voice system are controlled to perform warning actions according to the warning action command; if the risk level is Level 3, the risk event is pushed to the owner's APP at the same time; wherein, the threat level correspondence table includes risk levels Level 1, Level 2 and Level 3; when the behavior is to approach the vehicle and stay for a set time, the risk level is Level 1, and the warning action is hazard lights; when the behavior is to linger near the vehicle for more than a set time, the risk level is Level 2, and the warning action is hazard lights and playing a voice warning once; when the behavior is glass breakage or collision, the risk level is Level 3, and the warning action is hazard lights, horn, and event push to the owner's APP.

7. A vehicle security device based on multimodal perception and distributed AI, characterized in that, The apparatus, applied to the method of claim 1, comprises: A startup module is used to activate the vehicle's primary sensors in response to the vehicle being powered off. The primary sensors include a millimeter-wave radar and an inertial measurement unit (IMU). The millimeter-wave radar continuously monitors whether any objects are approaching the vehicle's vicinity at a preset detection distance, and the IMU continuously monitors the vehicle's displacement in the X and Y directions. The wake-up module is used to wake up the secondary sensors on the vehicle in response to the millimeter-wave radar detecting the approach of a target. If the IMU detects that the vehicle is vibrating, it acquires the recorded data of the secondary sensors within a specific time period before and after the vibration. The secondary sensors include a panoramic camera and a vehicle-end microphone. The panoramic camera records video data, and the vehicle-end microphone records audio data. The upload module is used to extract features from the recorded data, generate a feature vector file, and upload the feature vector file to the cloud computing center, so that the cloud computing center can perform threat behavior identification on the feature vector file and obtain threat identification results; wherein, the cloud computing center performs threat behavior identification based on AI; The receiving module is used to receive the threat identification results sent by the cloud computing center, and determine the risk level and corresponding warning action that match the threat identification results based on the threat identification results; The cloud computing center uses a CNN+Transformer deep learning model for threat behavior identification; the CNN+Transformer deep learning model is trained in the following way: Construct an MFCC dataset class and load a CSV file containing MFCC features and corresponding labels, whereby the labels are used to identify the behavior category corresponding to the audio. The data in the MFCC dataset are preprocessed, including reshaping the MFCC features into a preset shape and performing standardization. A CNN+Transformer hybrid model is constructed, which includes a CNN module, a fully connected layer, a Transformer module, and a classification layer. The CNN module is used to extract local features, the fully connected layer is used to convert the features output by the CNN module into dimensions that are compatible with the Transformer module, the Transformer module is used to capture temporal correlation features, and the classification layer is used to output the behavior category prediction results. Set training parameters, including batch size, learning rate, and training epochs; calculate the prediction loss using the cross-entropy loss function; update the model parameters using the Adam optimizer; and train the CNN+Transformer hybrid model. Save the trained model parameters to obtain a CNN+Transformer deep learning model that can be used for threat behavior recognition.

8. An electronic device, characterized in that, include: The device includes a processor, a storage medium, and a bus, wherein the storage medium stores machine-readable instructions executable by the processor, and when the electronic device is running, the processor communicates with the storage medium via the bus, and the processor executes the machine-readable instructions to perform the automotive security method based on multimodal perception and distributed AI as described in any one of claims 1 to 6.

9. A computer-readable storage medium, characterized in that, The computer-readable storage medium stores a computer program, which, when executed by a processor, performs the vehicle security method based on multimodal perception and distributed AI as described in any one of claims 1 to 6.