Vehicle collision detection method based on elastic wave and multi-modal large model

By employing a three-layer intelligent collision detection architecture based on elastic waves and a multimodal large model, the problems of response delay and high false positive rate in traditional vehicle collision detection methods are solved. This architecture achieves instant response and accurate classification, provides interpretable judgment results, and improves the reliability and refined analysis capabilities of collision detection.

CN122379460APending Publication Date: 2026-07-14SHANGHAI YOUKA NETWORK TECH CO LTD
View PDF 0 Cites 0 Cited by

Patent Information

Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
SHANGHAI YOUKA NETWORK TECH CO LTD
Filing Date
2026-04-14
Publication Date
2026-07-14

AI Technical Summary

Technical Problem

Traditional vehicle collision detection methods suffer from problems such as response delay, high false positive rate, lack of multi-source information fusion, and insufficient interpretability, making it difficult to simultaneously meet the requirements of immediate response and accurate classification.

Method used

A three-layer intelligent collision detection architecture based on elastic waves and a multimodal large model is adopted, including an instant response layer, an edge intelligence layer, and a cloud-based deep analysis layer. The elastic wave sensor collects signals, the EWCN model of the edge intelligence layer performs fine-grained classification, and the VLM model in the cloud performs multimodal data fusion analysis to provide interpretable judgment results.

Benefits of technology

It significantly reduces the false alarm rate, improves the reliability and interpretability of collision detection, and achieves a balance between immediate response and accurate classification, meeting the needs of rapid airbag triggering and refined collision analysis.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN122379460A_ABST
    Figure CN122379460A_ABST
Patent Text Reader

Abstract

The application discloses a vehicle collision detection method based on elastic waves and a multimodal large model, which comprises the following steps: collecting elastic wave signals by a vehicle end MCU, performing threshold judgment through an instant response layer, triggering a safety system immediately if it is a high-intensity collision, inputting the signals into an EWCN model of an edge intelligent layer by the vehicle end MCU if it is not a high-intensity collision, performing fine-grained classification on the non-high-intensity collision signals, directly adopting the classification result if the EWCN model outputs a high confidence or low confidence result, uploading multimodal data packets to a cloud deep analysis layer if the EWCN model outputs a to-be-confirmed scene, receiving data by a cloud server, performing fusion analysis on the multimodal input data by calling a VLM model, outputting an event type judgment result and corresponding explanation and returning to the vehicle end. The application solves the problem that a traditional scheme cannot simultaneously satisfy instant response and accurate classification, and meets the dual demands of rapid triggering of a safety airbag and fine collision analysis.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] This invention belongs to the field of collision detection technology, specifically a vehicle collision detection method based on elastic waves and a multimodal large model. Background Technology

[0002] Vehicle collision detection is one of the core technologies of automotive safety systems, directly affecting the reliability of critical functions such as airbag deployment, emergency calls, and accident recording. Traditional collision detection mainly relies on acceleration sensors and pressure sensors to determine whether a collision has occurred by detecting changes in acceleration or pressure caused by a collision.

[0003] Elastic wave collision detection technology is an emerging collision detection method. Its principle is as follows: when a vehicle collides, the impact force generates elastic waves within the vehicle's structure. These elastic waves propagate at high speed along the frame, carrying a wealth of collision information. By placing elastic wave sensors at key locations on the vehicle body, these signals can be captured, allowing analysis of the collision's location, intensity, type, and other characteristics.

[0004] Existing technologies include: (1) Accelerometer sensor solution This is currently the most mainstream collision detection solution. It detects changes in acceleration caused by a collision by installing acceleration sensors at key locations on the vehicle body. When the acceleration exceeds a preset threshold, a collision is determined to have occurred. Figure 4 The diagram illustrates the working principle of collision detection using a traditional accelerometer.

[0005] However, there is a certain delay in the response of the acceleration sensor. Since acceleration is an "effect" signal, it takes a certain amount of time for it to be reflected in the overall acceleration of the vehicle after a collision. In high-speed collision scenarios, this delay may lead to a delayed response from the safety system.

[0006] Accelerometers struggle to distinguish between genuine collisions and interference signals. Actions such as road bumps, sudden braking, and door slamming can all cause significant changes in acceleration, easily leading to false alarms. While setting a higher threshold can reduce false alarms, it also lowers detection sensitivity and increases the risk of missed detections.

[0007] (2) Elastic wave sensor scheme Elastic wave signals generated by collisions are detected using piezoelectric sensors such as PVDF. Compared with traditional methods, elastic wave detection has advantages such as fast response and large information capacity, but currently, it mainly suffers from problems such as complex signal processing and a high false positive rate.

[0008] Current elastic wave collision detection schemes mainly employ signal processing methods based on thresholds or simple rules, which have the following problems: Insufficient signal classification capability: Simple threshold judgment is insufficient to distinguish between different types of collision and interference signals. For example, the elastic wave generated by closing a door may have similar amplitude to the elastic wave from a minor collision, but should be handled by completely different strategies.

[0009] Limited ability to handle complex scenarios: When encountering novel collision scenarios or boundary conditions not covered by training data, traditional methods cannot provide reliable judgments, which may lead to misjudgments or omissions.

[0010] Lack of multi-source information fusion: Existing solutions rely solely on elastic wave signals, failing to fully utilize visual information such as those from vehicle-mounted cameras. For scenarios where elastic wave signals are ambiguous, there is a lack of effective auxiliary judgment methods.

[0011] Insufficient interpretability: The judgment logic of traditional methods is relatively simple, making it difficult to provide explanations for the judgment basis, which is not conducive to accident analysis and system optimization.

[0012] Therefore, to address the above problems, a vehicle collision detection method based on elastic waves and a multimodal large model is proposed. Summary of the Invention

[0013] To address the aforementioned problems in existing technologies, this invention provides a vehicle collision detection method based on elastic waves and a multimodal large model. This method solves the problem that traditional solutions cannot simultaneously meet the requirements of immediate response and accurate classification, thus satisfying the dual needs of rapid airbag triggering and refined collision analysis.

[0014] The technical solution to achieve the above objectives is: A vehicle collision detection method based on elastic waves and a multimodal large model includes: Step S1: The vehicle-mounted MCU (microcontroller unit) collects the elastic wave signal and performs threshold judgment through the instant response layer. If it is a high-intensity collision, the safety system is immediately triggered. Step S2: If it is not a high-intensity collision, the vehicle-side MCU will input the signal into the EWCN model (Elastic Wave Classification Network) of the edge intelligent layer to perform fine-grained classification of the non-high-intensity collision signal; the classification results include: high-intensity collision, medium-intensity collision, minor collision, interference signal and scenario to be confirmed. Step S3: If the EWCN model outputs a high-confidence or low-confidence result, the classification result is used directly. Step S4: If the EWCN model outputs a scenario to be confirmed, the vehicle-side MCU will upload the multimodal data packet to the cloud deep analysis layer; wherein, the multimodal data packet includes the original elastic wave signal, the classification probability distribution output by the edge-side EWCN model, the image data collected by the vehicle camera, the vehicle status information, the GPS positioning information, and the corresponding timestamp; In step S5, the cloud server receives the data, performs fusion analysis on the multimodal input data by calling the VLM model (Visual Language Model), outputs the event type judgment results and corresponding explanations, and returns them to the vehicle.

[0015] Preferably, in step S1, the threshold determination process of the instantaneous response layer includes: The elastic wave signals of each channel are acquired in real time by an elastic wave sensor and digitized at a sampling rate of 200kHz. The original signal is bandpass filtered to remove low-frequency noise and high-frequency interference, while retaining the effective frequency band of 1kHz-50kHz. Using a sliding window, calculate short-time energy: ; In the formula, For the first Each channel at time Short-term energy, For the first Each channel at the sampling point The signal value, The length of the sliding window is 200 sampling points; If the short-time energy of any channel exceeds the preset high-intensity threshold If the collision occurs, it is determined to be a high-intensity collision, and the passive safety system is immediately triggered. If the short-time energy of any channel is lower than the preset high-intensity threshold If the collision is not high-intensity, it is determined to be a non-high-intensity collision, and the non-high-intensity collision signal is transferred to step S2.

[0016] Preferably, in step S2, the non-high-intensity collision signal is classified by the EWCN model of the edge intelligence layer. The structure of the EWCN model includes: an input layer, a multi-scale feature extraction module, a feature compression layer, a temporal modeling module, an attention enhancement module, and a classification output layer. in, Input layer: The input is an 8-channel elastic wave time-series signal, with a single input length of 4000 sampling points, corresponding to a signal window of approximately 20ms. Its input characteristics are as follows: ; Multi-scale feature extraction module: Three parallel one-dimensional convolutional branches are used to extract features at different time scales. The branches have the same structure, only the kernel size is different. Branch 1: The kernel size is 3, used to extract short-term impact features; Branch 2: The kernel size is 7, used to extract mesoscale variation features; Branch 3: The kernel size is 15, used to extract long-term trend features; For the Each branch has a convolution operation represented as follows: ; In the formula, , Represents a non-linear activation function. Indicates batch standardization. This represents a one-dimensional convolution operation; The outputs of each branch are obtained after max pooling: ; The outputs of the three branches are concatenated along the channel dimension: ; Feature compression layer: One-dimensional convolution is used to compress features, which reduces computational complexity while preserving key information. ; Temporal modeling module: A lightweight temporal convolutional network is used for modeling, consisting of two layers of dilated causal convolutions: ; ; In the formula, This represents dilated causal convolution. Indicates the expansion rate; Attention Enhancement Module: First, perform global average pooling on the time dimension: ; In the formula, Indicates the total number of time steps; Channel weights are then generated using a two-layer fully connected network: ; In the formula, , , Represents the ReLU function. Represents the Sigmoid function; Finally, the features are weighted: ; In the formula, This represents channel-by-channel multiplication; Classification output layer: Perform global average pooling on the weighted features: ; The classification results are output through a fully connected layer: ; In the formula, This represents the weight matrix of the classification layer. This represents the bias vector of the classification layer. Indicates the number of categories; in, ; There are five types of signals: high-intensity collision, medium-intensity collision, minor collision, interference signal, and scenario to be confirmed.

[0017] Preferably, in step S2, a confidence level determination mechanism is introduced based on the output probability of the EWCN model, and the model confidence level is: ; Classification decisions are made when the following conditions are met: like If so, the corresponding classification result will be used directly; like If so, the classification result is adopted and marked as low confidence. like If the signal is not confirmed, it will be considered a scenario to be confirmed and will be transferred to the subsequent processing flow. in, .

[0018] Preferably, in step S5, the cloud performs fusion analysis on the multimodal input data by calling the VLM model, including: The cloud performs further feature processing on the elastic wave signal to extract high-level feature information for decision support. The feature processing includes: performing time-frequency analysis on the signal to obtain a time-spectrum representation, performing energy distribution statistics on different frequency bands to obtain frequency band energy characteristics, and calculating the time difference of arrival based on multi-channel signals to estimate the location of the signal source, thereby forming a structured elastic wave feature description. Elastic wave features, vehicle status information, and visual information are uniformly formatted to construct multimodal input data; among them, visual information is input in the form of images, and elastic wave features and vehicle status information are encoded in the form of structured text or descriptive text. The VLM model is invoked to perform fusion analysis on multimodal input data. Inference results are obtained through remote interface calls. The model performs cross-modal understanding based on the input image and text information, and outputs event type judgment results and corresponding explanations. The VLM model service is based on vision-language joint modeling capabilities to align and fuse input images and text descriptions. Its processing includes: visual feature encoding of input images, semantic encoding of input text, and correlation modeling within a unified semantic space, thereby generating a comprehensive judgment result and interpretable description of the current event. The analysis results are returned to the vehicle for decision-making reference, while the relevant data and analysis results are stored in the cloud database to support subsequent data backtracking, model optimization and statistical analysis.

[0019] Preferred options also include: Step S6: The collision location is then located based on the TDOA localization method.

[0020] Preferably, in step S6, the positioning method based on TDOA specifically includes: When the vehicle-side MCU detects a collision signal, it records the sensor data. , and The time when the signal was detected; Computational Sensors With sensors ,sensor With sensors Signal arrival time difference: ; ; In the formula, This indicates that a collision signal has arrived at the sensor. absolute time, This indicates that a collision signal has arrived at the sensor. absolute time, This indicates that a collision signal has arrived at the sensor. absolute time, Indicates sensor With sensors The time delay difference of the detected collision signal Indicates sensor With sensors The time delay difference in detecting the collision signal; The speed of elastic wave propagation in the frame With time difference , The product is used as a distance difference constraint to construct a system based on collision positions. The system of hyperbolic equations with unknowns yields the collision location estimate: ; ; In the formula, , and These represent the sensors. , and The location.

[0021] Compared with the prior art, the beneficial effects of the present invention are: This invention deploys the EWCN model at the edge intelligence layer and combines it with a confidence judgment mechanism. Instead of relying solely on simple threshold judgment, it performs multi-class classification through a deep learning model and introduces a confidence threshold to transfer uncertain scenarios to the cloud for processing. Because the EWCN model can learn the deep feature differences between collision signals and interference signals, and the confidence mechanism avoids errors caused by "forced judgment", the false positive rate is significantly reduced by about 70%, solving the problem of high false positive rate of traditional threshold judgment methods and improving the reliability of the collision detection system. This invention introduces a VLM model to generate natural language judgment criteria, which can not only provide judgment results, but also explain the reasons for the judgment in a human-understandable way. Because the VLM model has visual perception and language reasoning capabilities, it can combine image and signal features for comprehensive analysis and output explanations. Therefore, the collision judgment results have good interpretability, which solves the problem of lack of evidence explanation in traditional black box judgments, and facilitates accident analysis, system optimization and user understanding. In summary, this invention adopts a three-layer intelligent collision detection architecture, which distributes collision detection tasks to three layers—the immediate response layer, the edge intelligence layer, and the cloud-based deep analysis layer—according to their urgency and complexity. High-intensity collisions are handled immediately by the first layer, while complex scenarios are analyzed in depth by the cloud. This achieves a balance between response speed and detection accuracy, solving the problem that traditional solutions cannot simultaneously meet the requirements of immediate response and accurate classification. It also satisfies the dual needs of rapid airbag triggering and refined collision analysis. Attached Figure Description

[0022] The accompanying drawings are provided to further illustrate the invention and form part of the specification. They are used in conjunction with embodiments of the invention to explain the invention and do not constitute a limitation thereof. In the drawings: Figure 1 This is a flowchart of a vehicle collision detection method based on elastic waves and a multimodal large model according to the present invention; Figure 2 This is another flowchart of the vehicle collision detection method based on elastic waves and multimodal large model in this invention; Figure 3 This is a diagram of the EWCN model architecture in this invention; Figure 4 This is the working principle of collision detection using traditional accelerometers. Detailed Implementation

[0023] The technical solutions of the embodiments of the present invention will be clearly and completely described below with reference to the accompanying drawings. Obviously, the described embodiments are only some embodiments of the present invention, and not all embodiments. Based on the embodiments of the present invention, all other embodiments obtained by those skilled in the art without creative effort are within the scope of protection of the present invention.

[0024] A vehicle collision detection method based on elastic waves and multimodal large models is proposed. It is a three-layer intelligent collision detection architecture that achieves efficient and accurate detection of various collision scenarios through a layered processing mechanism of "real-time response layer - edge intelligence layer - cloud deep analysis layer".

[0025] The architecture comprises three layers as described in Table 1: Table 1 like Figure 1 , 2 As shown, the specific solutions include: Step S1: The vehicle-side MCU collects the elastic wave signal and performs threshold judgment through the instant response layer. If it is a high-intensity collision, the safety system is immediately triggered. The vehicle-side MCU is an automotive-grade high-performance MCU, such as the ARM Cortex-M series or automotive-grade AI chips (such as NVIDIA Orin and Horizon Robotics series).

[0026] In this embodiment, the threshold determination process of the instant response layer includes: Elastic wave signals from each channel are acquired in real time using elastic wave sensors (such as PVDF piezoelectric films, PZT piezoelectric ceramics, fiber optic sensors, MEMS accelerometers, etc.) and digitized at a sampling rate of 200kHz. The original signal is bandpass filtered to remove low-frequency noise and high-frequency interference, while retaining the effective frequency band of 1kHz-50kHz. Using a sliding window, calculate short-time energy: ; In the formula, For the first Each channel at time Short-term energy, For the first Each channel at the sampling point The signal value, The length of the sliding window is 200 sampling points; If the short-time energy of any channel exceeds the preset high-intensity threshold If the collision occurs, it is determined to be a high-intensity collision, and the passive safety system is immediately triggered. If the short-time energy of any channel is lower than the preset high-intensity threshold If the collision is not high-intensity, it is determined to be a non-high-intensity collision, and the non-high-intensity collision signal is transferred to step S2.

[0027] Based on a large amount of collision experiment data, the signal energy distribution under different collision intensities was statistically analyzed. The energy value that can distinguish high-intensity collisions from other events was selected as the threshold typical value: the signal energy corresponding to a collision deceleration ≥30g.

[0028] Step S2: If it is not a high-intensity collision, the vehicle-side MCU will input the signal into the EWCN model of the edge intelligence layer to perform fine-grained classification of the non-high-intensity collision signal; among which, the classification results include: high-intensity collision, medium-intensity collision, minor collision, interference signal and scenario to be confirmed.

[0029] In this embodiment, based on the rapid judgment made by energy threshold in step S1, there is a possibility of missed judgment due to blind spots in sensor placement or signal propagation attenuation. Therefore, the EWCN model retains the ability to identify high-intensity collisions as a first-layer insurance mechanism.

[0030] In this embodiment, non-high-intensity collision signals are classified using the EWCN model of the edge intelligent layer, wherein, for example... Figure 3 As shown, the structure of the EWCN model includes: an input layer, a multi-scale feature extraction module, a feature compression layer, a temporal modeling module, an attention enhancement module, and a classification output layer; in, Input layer: The input is an 8-channel elastic wave time-series signal, with a single input length of 4000 sampling points, corresponding to a signal window of approximately 20ms. Its input characteristics are as follows: ; Multi-scale feature extraction module: Three parallel one-dimensional convolutional branches are used to extract features at different time scales. The branches have the same structure, only the kernel size is different. Branch 1: The kernel size is 3, used to extract short-term impact features; Branch 2: The kernel size is 7, used to extract mesoscale variation features; Branch 3: The kernel size is 15, used to extract long-term trend features; For the Each branch has a convolution operation represented as follows: ; In the formula, , Represents a non-linear activation function. Indicates batch standardization. This represents a one-dimensional convolution operation; The outputs of each branch are obtained after max pooling: ; The outputs of the three branches are concatenated along the channel dimension: ; Feature compression layer: One-dimensional convolution is used to compress features, which reduces computational complexity while preserving key information. ; Temporal modeling module: A lightweight temporal convolutional network is used for modeling, consisting of two layers of dilated causal convolutions: ; ; In the formula, This represents dilated causal convolution. Indicates the expansion rate; Attention Enhancement Module: First, perform global average pooling on the time dimension: ; In the formula, Indicates the total number of time steps; Channel weights are then generated using a two-layer fully connected network: ; In the formula, , , Represents the ReLU function. Represents the Sigmoid function; Finally, the features are weighted: ; In the formula, This represents channel-by-channel multiplication; Classification output layer: Perform global average pooling on the weighted features: ; The classification results are output through a fully connected layer: ; In the formula, This represents the weight matrix of the classification layer. This represents the bias vector of the classification layer. Indicates the number of categories; in, ; There are five types of signals: high-intensity collision, medium-intensity collision, minor collision, interference signal, and scenario to be confirmed.

[0031] In this embodiment, a confidence level determination mechanism is introduced based on the output probability of the EWCN model, and the model confidence level is: ; Classification decisions are made when the following conditions are met: like If so, the corresponding classification result will be used directly; like If so, the classification result is adopted and marked as low confidence. like If the signal is not confirmed, it will be considered a scenario to be confirmed and will be transferred to the subsequent processing flow. in, .

[0032] Through the above structural design and confidence determination mechanism, the model still has good classification performance and reliability under limited computing resources, and can meet the needs of edge side for real-time identification and hierarchical processing of elastic wave signals.

[0033] Step S3: If the EWCN model outputs a high-confidence or low-confidence result, the classification result is used directly.

[0034] Step S4: If the EWCN model outputs a scenario to be confirmed, the vehicle-side MCU will upload the multimodal data packet to the cloud-based deep analysis layer. The multimodal data packet includes the original elastic wave signal, the classification probability distribution output by the edge-side EWCN model, image data collected by the vehicle-mounted camera, vehicle status information, GPS positioning information, and corresponding timestamps.

[0035] In step S5, the cloud server receives the data and performs fusion analysis on the multimodal input data by calling the VLM model (which can be different visual language large models such as GPT-4V, LLaVA, Qwen-VL, etc.). The output includes the event type judgment result and the corresponding explanation, and is returned to the vehicle. The cloud server is a cloud server equipped with GPU acceleration.

[0036] In this embodiment, the cloud performs fusion analysis on multimodal input data by calling the VLM model, including: The cloud performs further feature processing on the elastic wave signal to extract high-level feature information for decision support. The feature processing includes: performing time-frequency analysis on the signal to obtain a time-spectrum representation, performing energy distribution statistics on different frequency bands to obtain frequency band energy characteristics, and calculating the time difference of arrival based on multi-channel signals to estimate the location of the signal source, thereby forming a structured elastic wave feature description. Elastic wave features, vehicle status information, and visual information are uniformly formatted to construct multimodal input data; among them, visual information is input in the form of images, and elastic wave features and vehicle status information are encoded in the form of structured text or descriptive text. The VLM model is invoked to perform fusion analysis on multimodal input data. Inference results are obtained through remote interface calls. The model performs cross-modal understanding based on the input image and text information, and outputs event type judgment results and corresponding explanations. The VLM model service is based on vision-language joint modeling capabilities to align and fuse input images and text descriptions. Its processing includes: visual feature encoding of input images, semantic encoding of input text, and correlation modeling within a unified semantic space, thereby generating a comprehensive judgment result and interpretable description of the current event. The analysis results are returned to the vehicle for decision-making reference, while the relevant data and analysis results are stored in the cloud database to support subsequent data backtracking, model optimization and statistical analysis.

[0037] Step S6: The collision location is then located based on the TDOA localization method.

[0038] In this embodiment, the TDOA-based positioning method specifically includes: When the vehicle-side MCU detects a collision signal, it records the sensor data. , and The time when the signal was detected; Computational Sensors With sensors ,sensor With sensors Signal arrival time difference: ; ; In the formula, This indicates that a collision signal has arrived at the sensor. absolute time, This indicates that a collision signal has arrived at the sensor. absolute time, This indicates that a collision signal has arrived at the sensor. absolute time, Indicates sensor With sensors The time delay difference of the detected collision signal Indicates sensor With sensors The time delay difference in detecting the collision signal; The speed of elastic wave propagation in the frame With time difference , The product is used as a distance difference constraint to construct a system based on collision positions. The system of hyperbolic equations with unknowns yields the collision location estimate: ; ; In the formula, , and These represent the sensors. , and The location.

[0039] Finally, it should be noted that the above are merely preferred embodiments of the present invention and are not intended to limit the present invention. Although the present invention has been described in detail with reference to the foregoing embodiments, those skilled in the art can still modify the technical solutions described in the foregoing embodiments or make equivalent substitutions for some of the technical features. Any modifications, equivalent substitutions, improvements, etc., made within the spirit and principles of the present invention should be included within the protection scope of the present invention.

Claims

1. A vehicle collision detection method based on elastic waves and a multimodal large model, characterized in that, include: Step S1: The vehicle-side MCU collects the elastic wave signal and performs threshold judgment through the instant response layer. If it is a high-intensity collision, the safety system is immediately triggered. Step S2: If it is not a high-intensity collision, the vehicle-side MCU will input the signal into the EWCN model of the edge intelligence layer to perform fine-grained classification of the non-high-intensity collision signal; the classification results include: high-intensity collision, medium-intensity collision, minor collision, interference signal and scenario to be confirmed. Step S3: If the EWCN model outputs a high-confidence or low-confidence result, the classification result is used directly. Step S4: If the EWCN model outputs a scenario to be confirmed, the vehicle-side MCU will upload the multimodal data packet to the cloud deep analysis layer; wherein, the multimodal data packet includes the original elastic wave signal, the classification probability distribution output by the edge-side EWCN model, the image data collected by the vehicle camera, the vehicle status information, the GPS positioning information, and the corresponding timestamp; In step S5, the cloud server receives the data, performs fusion analysis on the multimodal input data by calling the VLM model, outputs the event type judgment results and corresponding explanations, and returns them to the vehicle.

2. The vehicle collision detection method based on elastic waves and a multimodal large model according to claim 1, characterized in that, In step S1, the threshold determination process of the instantaneous response layer includes: The elastic wave signals of each channel are acquired in real time by an elastic wave sensor and digitized at a sampling rate of 200kHz. The original signal is bandpass filtered to remove low-frequency noise and high-frequency interference, while retaining the effective frequency band of 1kHz-50kHz. Using a sliding window, calculate short-time energy: ; In the formula, For the first Each channel at time Short-term energy, For the first Each channel at the sampling point The signal value, The length of the sliding window is 200 sampling points; If the short-time energy of any channel exceeds the preset high-intensity threshold If the collision occurs, it is determined to be a high-intensity collision, and the passive safety system is immediately triggered. If the short-time energy of any channel is lower than the preset high-intensity threshold If the collision is not high-intensity, it is determined to be a non-high-intensity collision, and the non-high-intensity collision signal is transferred to step S2.

3. The vehicle collision detection method based on elastic waves and a multimodal large model according to claim 1, characterized in that, In step S2, non-high-intensity collision signals are classified using the EWCN model of the edge intelligence layer. The structure of the EWCN model includes: an input layer, a multi-scale feature extraction module, a feature compression layer, a temporal modeling module, an attention enhancement module, and a classification output layer. in, Input layer: The input is an 8-channel elastic wave time-series signal, with a single input length of 4000 sampling points, corresponding to a signal window of approximately 20ms. Its input characteristics are as follows: ; Multi-scale feature extraction module: Three parallel one-dimensional convolutional branches are used to extract features at different time scales. The branches have the same structure, only the kernel size is different. Branch 1: The kernel size is 3, used to extract short-term impact features; Branch 2: The kernel size is 7, used to extract mesoscale variation features; Branch 3: The kernel size is 15, used to extract long-term trend features; For the Each branch has a convolution operation represented as follows: ; In the formula, , Represents a non-linear activation function. Indicates batch standardization. This represents a one-dimensional convolution operation; The outputs of each branch are obtained after max pooling: ; The outputs of the three branches are concatenated along the channel dimension: ; Feature compression layer: One-dimensional convolution is used to compress features, which reduces computational complexity while preserving key information. ; Temporal modeling module: A lightweight temporal convolutional network is used for modeling, consisting of two layers of dilated causal convolutions: ; ; In the formula, This represents dilated causal convolution. Indicates the expansion rate; Attention Enhancement Module: First, perform global average pooling on the time dimension: ; In the formula, Indicates the total number of time steps; Channel weights are then generated using a two-layer fully connected network: ; In the formula, , , Represents the ReLU function. Represents the Sigmoid function; Finally, the features are weighted: ; In the formula, This represents channel-by-channel multiplication; Classification output layer: Perform global average pooling on the weighted features: ; The classification results are output through a fully connected layer: ; In the formula, This represents the weight matrix of the classification layer. This represents the bias vector of the classification layer. Indicates the number of categories; in, ; There are five types of signals: high-intensity collision, medium-intensity collision, minor collision, interference signal, and scenario to be confirmed.

4. The vehicle collision detection method based on elastic waves and a multimodal large model according to claim 3, characterized in that, In step S2, a confidence level determination mechanism is introduced based on the output probability of the EWCN model. The model confidence level is: ; Classification decisions are made when the following conditions are met: like If so, the corresponding classification result will be used directly; like If so, the classification result is adopted and marked as low confidence. like If the signal is not confirmed, it will be considered a scenario to be confirmed and will be transferred to the subsequent processing flow. in, .

5. The vehicle collision detection method based on elastic waves and a multimodal large model according to claim 1, characterized in that, In step S5, the cloud platform performs fusion analysis on the multimodal input data by calling the VLM model, including: The cloud performs further feature processing on the elastic wave signal to extract high-level feature information for decision support. The feature processing includes: performing time-frequency analysis on the signal to obtain a time-spectrum representation, performing energy distribution statistics on different frequency bands to obtain frequency band energy characteristics, and calculating the time difference of arrival based on multi-channel signals to estimate the location of the signal source, thereby forming a structured elastic wave feature description. Elastic wave features, vehicle status information, and visual information are uniformly formatted to construct multimodal input data; among them, visual information is input in the form of images, and elastic wave features and vehicle status information are encoded in the form of structured text or descriptive text. The VLM model is invoked to perform fusion analysis on multimodal input data. Inference results are obtained through remote interface calls. The model performs cross-modal understanding based on the input image and text information, and outputs event type judgment results and corresponding explanations. The VLM model service is based on vision-language joint modeling capabilities to align and fuse input images and text descriptions. Its processing includes: visual feature encoding of input images, semantic encoding of input text, and correlation modeling within a unified semantic space, thereby generating a comprehensive judgment result and interpretable description of the current event. The analysis results are returned to the vehicle for decision-making reference, while the relevant data and analysis results are stored in the cloud database to support subsequent data backtracking, model optimization and statistical analysis.

6. The vehicle collision detection method based on elastic waves and a multimodal large model according to claim 1, characterized in that, Also includes: Step S6: The collision location is then located based on the TDOA localization method.

7. The vehicle collision detection method based on elastic waves and a multimodal large model according to claim 6, characterized in that, In step S6, the TDOA-based positioning method specifically includes: When the vehicle-side MCU detects a collision signal, it records the sensor data. , and The time when the signal was detected; Computational Sensors With sensors ,sensor With sensors Signal arrival time difference: ; ; In the formula, This indicates that a collision signal has arrived at the sensor. absolute time, This indicates that a collision signal has arrived at the sensor. absolute time, This indicates that a collision signal has arrived at the sensor. absolute time, Indicates sensor With sensors The time delay difference of the detected collision signal Indicates sensor With sensors The time delay difference in detecting the collision signal; The speed of elastic wave propagation in the frame With time difference , The product is used as a distance difference constraint to construct a system based on collision positions. The system of hyperbolic equations with unknowns yields the collision location estimate: ; ; In the formula, , and These represent the sensors. , and The location.