System and method for detecting and intervening illegal manned sharing of shared electric vehicle
By combining a multimodal fusion control module with pressure sensors, distance sensors, and a visual recognition module, the system can detect illegal passenger-carrying behavior in shared electric vehicles in real time. This solves the problems of limited regulatory coverage and slow response time in existing technologies, achieving efficient and accurate detection and intervention of illegal passenger-carrying and ensuring traffic safety.
Patent Information
- Application Number
- CN202510982252.7
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-07-16
- Publication Date
- 2025-11-14
AI Technical Summary
Existing technologies for detecting illegal passenger transport on shared electric bikes suffer from limited coverage, slow response time, high false positive and false negative rates, lack of real-time monitoring and closed-loop management, and inability to effectively respond to rapidly changing on-site situations, resulting in traffic safety hazards and poor regulatory effectiveness.
The system employs a multimodal fusion control module, which combines a pressure sensor array, a distance sensor, and a visual recognition module to collect real-time data on seat pressure distribution, passenger distance, and passenger images. The multimodal fusion control module determines whether passengers are violating regulations and disconnects the motor from the power supply when a violation is detected. Real-time intervention is achieved by combining helmet detection and cloud data analysis.
It enables millisecond-level real-time detection and dynamic intervention of shared electric vehicles carrying passengers illegally, reducing the false detection rate, improving the real-time nature and accuracy of supervision, forming a closed-loop safety management system, and reducing the occurrence of traffic accidents.
Smart Images

Figure CN120942471A_ABST
Abstract
Description
Technical Field
[0001] This invention relates to the field of safety monitoring and management technology for shared electric vehicles, and in particular to a system and method for detecting and intervening in the illegal carrying of passengers on shared electric vehicles. Background Technology
[0002] Currently, shared electric bikes, with their outstanding convenience and environmental friendliness, have rapidly become an important supplementary mode of short-distance urban travel. They provide urban residents with a flexible and efficient solution for short-distance travel, effectively alleviating urban traffic congestion, meeting diverse travel needs, and becoming an indispensable part of the urban transportation system.
[0003] Riding shared electric bikes requires adhering to the "one person, one bike" usage rule. However, in reality, violations such as riding with two people, side-sitting, and children standing on the footrests frequently occur. These violations disrupt the vehicle's center of gravity balance, significantly increasing braking distance and greatly raising the probability of rollovers and rear-end collisions, posing a serious threat to the safety of riders, pedestrians, and other vehicles.
[0004] Currently, the supervision of shared electric bikes carrying passengers illegally relies mainly on manual inspections and user reports. However, these two methods have limited coverage and cannot achieve comprehensive, real-time monitoring of shared electric bike usage. Moreover, there is a significant time lag between the occurrence of a violation and its discovery and intervention, making it difficult to respond promptly to rapidly changing on-site situations and failing to meet the needs of dynamic supervision, resulting in poor regulatory effectiveness. Summary of the Invention
[0005] The purpose of this invention is to provide a system and method for detecting and intervening in illegal passenger carrying on shared electric vehicles, in order to solve the problems of low detection rate and slow response time of illegal behavior when passengers ride electric vehicles.
[0006] To address the aforementioned technical problems, in a first aspect, the present invention provides a shared electric vehicle illegal passenger-carrying detection and intervention system, comprising: a pressure sensor array module, a distance sensor module, a visual recognition module, a multimodal fusion control module, and an execution module; The multimodal fusion control module is connected to the pressure sensor array module, the distance sensor module, the visual recognition module, and the execution module; The pressure sensor array modules are distributed below the saddle and are used to collect the pressure distribution on the seat surface; The distance sensor module is located at the rear of the vehicle and is used to detect the distance between the occupants and the vehicle body. The visual recognition module is used to acquire images of passengers in real time. The multimodal fusion control module is used to determine whether the passenger has violated regulations based on the seat pressure distribution, the riding distance, and the riding image. The execution module is used to disconnect the motor from the power supply when the occupant's violation is detected.
[0007] In one possible implementation, the pressure sensor array module includes: a plurality of piezoresistive sensing units; The piezoresistive sensing unit is arranged in a dot matrix under the seat cushion, and the piezoresistive sensing unit is connected to the multimodal fusion control module. The plurality of piezoresistive sensing units are used to collect sampling voltages when the occupant is riding, and to map the sampling voltages as seat pressure distributions. The multimodal fusion control module is used to receive the seat pressure distribution at preset intervals and to calibrate and compensate the seat pressure distribution to obtain a pressure heat map.
[0008] In one possible implementation, the multimodal fusion control module is further configured to perform weighted clustering on the pressure heatmap using a dual clustering algorithm, output the cluster center spacing and the intra-cluster variance ratio, and determine the occupant violation when the cluster center spacing is greater than a preset spacing or the intra-cluster variance ratio is greater than a preset ratio.
[0009] In one possible implementation, the distance sensor module includes: an infrared ranging sensor and an ultrasonic module; The infrared ranging sensor and the ultrasonic module are connected to the multimodal fusion control module; The infrared ranging sensor is installed at the rear end of the vehicle and is used to measure the distance between the sensor and the occupants as the first distance measurement. The ultrasonic module is installed at the rear end of the vehicle and is used to measure the distance between the vehicle and the occupants as a second distance measurement. The multimodal fusion control module is further configured to fuse the first ranging and the second ranging through Kalman filtering to output a fused distance, and to determine that the occupant violates the rules when the fused distance is less than a preset distance and the cluster center spacing is less than a preset spacing but greater than a minimum spacing.
[0010] In one possible implementation, the multimodal fusion control module is used to identify the passenger's riding posture based on the riding image, and to determine that the passenger has violated the rules when a high-risk posture is detected and the confidence level is greater than the minimum confidence level threshold.
[0011] In one possible implementation, the multimodal fusion control module uses the pressure heatmap as a reference to align the riding distance and the riding image through linear interpolation mapping. Based on the pressure heatmap, the fused distance, and the riding image, it uses a Bi-LSTM and self-attention fusion network to output the violation probability. When the violation probability is greater than a preset threshold, it determines that the occupant has violated the rules.
[0012] In one possible implementation, the system further includes: a helmet detection module; The helmet detection module is connected to the execution module; The helmet detection module is used to send a helmet violation signal to the execution module when it detects that the helmet has fallen off for more than a set time. The execution module is also used to issue an audible and visual alarm and limit the vehicle's speed when the helmet violation signal is detected.
[0013] In one possible implementation, the system further includes: a communication module and a cloud data analysis module; The communication module connects the multimodal fusion control module and the cloud data analysis module; The communication module is used to send violation event data to the cloud server and output voice prompts when the occupant's violation is detected. The cloud-based data analysis module is used to store the violation data, generate heat map reports of high-incidence areas and high-incidence periods through spatiotemporal clustering algorithms, and push safety education videos to passengers.
[0014] In one possible implementation, the system further includes: a model update module; The model update module is connected to the cloud data analysis module, the communication module, the pressure sensor array module, the distance sensor module, the visual recognition module, the multimodal fusion control module, and the execution module; The model update module is used to receive false alarm data from occupants and update and adjust each module through transfer learning.
[0015] On the other hand, the present invention provides a method for detecting and intervening in the illegal carrying of passengers in shared electric vehicles, applied to the aforementioned system for detecting and intervening in the illegal carrying of passengers in shared electric vehicles, the method comprising: Collect seat pressure distribution, passenger distance, and passenger images; The pressure heatmap, the travel distance, and the travel image are preprocessed and feature extracted. Multimodal fusion is used to determine whether the passenger violated the rules. When the occupant violates the rules, the vehicle's power will be cut off and an alarm will be triggered. When the occupant resumes safe driving, the violation is reported and safety education is sent to the occupant.
[0016] The beneficial effects of this invention are as follows: This invention provides a shared electric vehicle illegal passenger-carrying detection and intervention system and method. It collects seat pressure distribution data using a pressure sensor array module located under the saddle, generating a pressure heatmap. A distance sensor module detects the passenger's distance from the vehicle body. A visual recognition module collects passenger images. A multimodal fusion control module collects the outputs of these three modules to determine if the passenger is violating regulations. When a passenger violates regulations, the execution module cuts off the power. Compared to existing single-sensor or single-vision solutions, this system exhibits significant advantages in real-world road scenarios. Furthermore, this application incorporates a helmet status interlock mechanism, achieving millisecond-level violation judgment and dynamic speed limit control, and linking with cloud-based educational push notifications to form a closed-loop safety management system. Attached Figure Description
[0017] To more clearly illustrate the technical solutions in the embodiments of the present invention, the accompanying drawings used in the description of the embodiments will be briefly introduced below. Obviously, the accompanying drawings described below are only some embodiments of the present invention. For those skilled in the art, other drawings can be obtained based on these drawings without creative effort.
[0018] Figure 1 This is a schematic diagram of the first embodiment of the shared electric vehicle illegal passenger carrying detection and intervention system of the present invention; Figure 2 This is a schematic diagram of the pressure sensor array module structure of the first embodiment of the shared electric vehicle illegal passenger carrying detection and intervention system of the present invention; Figure 3 This is a schematic diagram of the distance sensor module structure of the first embodiment of the shared electric vehicle illegal passenger carrying detection and intervention system of the present invention; Figure 4 This is a schematic diagram of the structure of the visual recognition module in the first embodiment of the shared electric vehicle illegal passenger carrying detection and intervention system of the present invention; Figure 5 This is a schematic diagram of the third embodiment of the shared electric vehicle illegal passenger carrying detection and intervention system of the present invention; Figure 6 This is a schematic diagram of the modules in the third embodiment of the shared electric vehicle illegal passenger carrying detection and intervention system of the present invention; Figure 7 This is a flowchart illustrating one implementation method of the shared electric vehicle illegal passenger carrying detection and intervention method of the present invention.
[0019] In the diagram: 100 - visual recognition module, 200 - multimodal fusion control module, 300 - pressure sensor array module, 400 - distance sensor module, 500 - execution module. Detailed Implementation
[0020] The technical solutions of the embodiments of the present invention will be clearly and completely described below with reference to the accompanying drawings. Obviously, the described embodiments are only a part of the embodiments of the present invention, and not all of them. All other embodiments obtained by those skilled in the art based on the embodiments of the present invention without creative effort are within the scope of protection of the present invention.
[0021] In the description of the embodiments of the present invention, unless otherwise stated, "multiple" means two or more. "And / or" describes the relationship between related objects, indicating that there can be three relationships. For example, A and / or B can represent three situations: A exists alone, A and B exist simultaneously, and B exists alone.
[0022] The terms "first," "second," etc., used in the embodiments of this invention are for descriptive purposes only and should not be construed as indicating or implying their relative importance or implicitly specifying the number of technical features indicated. Therefore, a technical feature defined with "first" or "second" may explicitly or implicitly include at least one of that feature.
[0023] In this document, the term "embodiment" means that a particular feature, structure, or characteristic described in connection with an embodiment may be included in at least one embodiment of the invention. The appearance of this phrase in various places throughout the specification does not necessarily refer to the same embodiment, nor is it a separate or alternative embodiment mutually exclusive with other embodiments. It will be explicitly and implicitly understood by those skilled in the art that the embodiments described herein can be combined with other embodiments.
[0024] In urban transportation systems, shared electric bikes have become an important choice for short-distance travel due to their convenience and environmental friendliness. However, with their widespread adoption, the problem of illegal use of shared electric bikes has become increasingly prominent, especially the illegal carrying of passengers, posing serious threats to traffic safety and presenting numerous challenges to operating platforms. Existing regulatory methods have the following shortcomings.
[0025] First, current regulations on shared electric bikes carrying passengers illegally mainly rely on manual inspections and user reports. However, this method has extremely low coverage and is difficult to achieve comprehensive supervision. Moreover, there is a significant delay between the occurrence of a violation and the implementation of intervention, making dynamic monitoring impossible and unable to cope with rapidly changing on-site environments, thus greatly reducing the effectiveness of supervision.
[0026] Secondly, some technologies employ a single sensing method to detect unauthorized passenger carrying, such as pressure sensing or ultrasonic ranging. However, these methods suffer from significant false positives and false negatives. Because they cannot accurately distinguish between passengers and luggage, a single person carrying heavy loads is easily misjudged as riding as a tandem; and they are difficult to effectively detect atypical postures such as side-sitting or legs-hugging. Furthermore, ultrasonic sensors experience increased ranging errors under adverse weather conditions, further contributing to the higher false positive rate.
[0027] Third, some technologies use monocular camera solutions for monitoring, but this approach has a high false negative rate in nighttime scenarios, especially for small targets such as children standing on platform platforms, where the accuracy is low and there is a significant performance bottleneck. Furthermore, relying solely on visual means cannot determine helmet wearing status in real time, and it is difficult to combine this with information from other sensors, limiting the comprehensiveness and accuracy of the monitoring.
[0028] Fourth, existing technologies lack effective interlocking mechanisms for helmet detection and speed limiting. When users are not wearing helmets, the speed limit enforcement rate is low, and it is impossible to dynamically reduce speed and provide reminders in a timely manner when violations occur, making it difficult to ensure riding safety.
[0029] Fifth, the existing authentication system has design flaws, and it is common for minors to use adult accounts to ride bicycles. Because the system does not verify the age of the rider, it poses legal risks and creates pressure for public oversight.
[0030] Sixth, existing technical solutions are slow in judgment speed and cannot achieve millisecond-level rapid intervention. Furthermore, they have poor environmental adaptability; the misjudgment rate increases significantly under adverse weather conditions such as strong light, rain, and fog. Moreover, the lack of a closed-loop management mechanism makes it difficult to effectively control and correct violations.
[0031] Due to the aforementioned problems, platforms bear certain joint liability in traffic accidents caused by illegal passenger transport, facing numerous compensation and fine cases, resulting in significant economic and reputational losses. Therefore, developing an efficient, accurate, and reliable technology for monitoring and preventing violations by shared electric vehicles is of significant practical importance.
[0032] Based on the above-mentioned technical problems, this invention introduces a fusion strategy of multimodal pressure, distance, vision and near-field communication, combined with edge-cloud collaborative computing, to achieve millisecond-level real-time detection and active intervention of illegal passenger carrying behavior and helmet wearing status.
[0033] Please see Figure 1The first embodiment of the shared electric vehicle illegal passenger-carrying detection and intervention system provided by the first solution of the present invention is shown in the following module diagram: a pressure sensor array module 300, a distance sensor module 400, a vision recognition module 100, a multimodal fusion control module 200, and an execution module 500; the multimodal fusion control module connects the pressure sensor array module, the distance sensor module, the vision recognition module, and the execution module; the pressure sensor array module is distributed and arranged under the saddle to collect seat pressure distribution; the distance sensor module is arranged at the rear of the vehicle to detect the riding distance between the passenger and the vehicle body; the vision recognition module is used to collect riding images in real time; the multimodal fusion control module is used to determine whether the passenger is violating the rules based on the seat pressure distribution, riding distance, and riding images; the execution module is used to disconnect the connection between the motor and the power supply when a passenger violation is detected.
[0034] It is understandable that, such as Figure 2 As shown, the pressure sensor array module 300 uses 64 piezoresistive sensing units (such as the Interlink FSR series or Tekscan FlexiForce), arranged in an 8×8 matrix (64 temperature measurement points) under the saddle, with a sampling frequency of 50Hz. The piezoresistive sensing units are based on piezoresistive materials made of silicon or conductive polymers. Their resistance changes non-linearly with the magnitude of the applied force. Under load, the internal lattice of the material deforms, causing a change in carrier mobility, thereby altering the resistance. Each piezoresistive sensing unit is connected to the 12-bit ADC channel of the STM32F407 (multimodal fusion control module) via a bridge circuit (such as a Wheatstone bridge). The sampled voltage value is mapped to the pressure distribution on the seat surface, and the raw pressure data is transmitted to the STM32F407 MCU via the SPI bus to generate a pressure-thermal map of the seat surface.
[0035] It should be noted that, as Figure 3 As shown, the distance sensor module 400 can be an infrared ranging sensor and an ultrasonic module, which are installed at the rear of the saddle of the shared electric vehicle. The infrared ranging sensor can be a VL53L1X infrared time-of-flight (ToF) sensor, which has the characteristics of high precision, high frequency and low power consumption. The ultrasonic module can be an HC-SR04. The two sensors simultaneously (at a frequency of 20 Hz) sample and detect the distance between the passenger and the vehicle body.
[0036] Specifically, such as Figure 4As shown, the visual recognition module 100 can use a Sony IMX219 miniature camera (640×480 resolution @ 30fps) and a MobileNet-SSD-Depthwise lightweight pose recognition model (INT8 quantized model size < 5MB) equipped with a multimodal fusion control module. It is actually deployed on an NVIDIA Jetson Nano (4 GB), with an average inference latency of about 48ms. The multimodal fusion control module identifies six types of illegal poses in real time, including "forward step," "side step," "leg hug," and "standing on a step stool." MobileNet-SSD-Depthwise uses depthwise separable convolution instead of traditional convolution, which significantly reduces the amount of computation and parameters while maintaining high detection accuracy. The NVIDIA Jetson Nano AI hardware, about the size of a credit card, features a 128-core NVIDIA Maxwell architecture with a frequency of 921MHz, providing 472GFLOPS of computing power. It supports parallel neural network inference for handling multitasking and control logic, and supports 4K@30p encoding, 4K@60p decoding, or 8-channel 1080p@30p parallel processing, making it suitable for video analysis scenarios.
[0037] Optionally, the multimodal fusion control module 200 is based on the STM32F407 and works in collaboration with the NVIDIA Jetson Nano. The STM32F407 performs data preprocessing, filtering, and feature extraction on the pressure heatmap and travel distance, while the NVIDIA Jetson Nano performs inference on the travel images. After aligning the multimodal data through Dynamic Time Warping (DTW), the data is input into a Bi-LSTM + Self-Attention fusion network, which outputs a violation probability. When the violation probability exceeds a preset threshold, a violation by the passenger is determined. Dynamic Time Warping (DTW) addresses the inconsistency of the time axis in the multimodal data, using the pressure heatmap as a reference to align the travel distance and travel images. The Bi-LSTM (Bidirectional Long Short-Term Memory) network processes the forward and reverse sequences of the sequence through two independent LSTM layers (forward + backward), ultimately fusing bidirectional information to improve the ability to capture contextual dependencies. Self-Attention calculates the correlation between each element in the input sequence and all other elements, generating attention weights, and then performs a weighted summation of the value vectors.
[0038] Furthermore, if a violation is detected, the STM32F407 outputs a TTL (digital level) signal, which, after optocoupler isolation, drives the IR2104+FQP30N06L positive voltage side switch to cut off the motor power supply (response time less than 5ms), and controls the headlights to flash at 2 Hz and drives the buzzer to sound an alarm at 1 kHz via the PWM channel.
[0039] In this embodiment, the present invention collects the pressure distribution of the seat surface through a pressure sensor array module installed under the saddle to generate a pressure heat map, detects the riding distance between the occupant and the vehicle body through a distance sensor module, collects riding images through a visual recognition module, and collects the outputs of the three modules through a multimodal fusion control module to determine whether the occupant has violated regulations. If the occupant violates regulations, the execution module cuts off the power.
[0040] Based on the first embodiment of this application, in the second embodiment of this application, the same or similar content as the first embodiment described above can be referred to the above description, and will not be repeated hereafter.
[0041] In some embodiments of the present invention, the pressure sensor array module includes: multiple piezoresistive sensing units; the piezoresistive sensing units are arranged in a matrix under the seat and connected to a multimodal fusion control module; the multiple piezoresistive sensing units are used to collect sampling voltages when the occupant is riding, and map the sampling voltages to the seat pressure distribution; the multimodal fusion control module is used to receive the seat pressure distribution at preset intervals, and calibrate and compensate the seat pressure distribution to obtain a pressure heat map.
[0042] It should be noted that the piezoresistive sensing unit is based on piezoresistive materials made of silicon or conductive polymers, and its resistance changes non-linearly with the magnitude of the applied force. Under load, the internal lattice of the material deforms, causing a change in carrier mobility, thereby altering the resistance. Sixty-four independent piezoresistive sensing units are arranged in an 8×8 array under the seat cushion. Each piezoresistive sensing unit is typically about 10mm×10mm in size, covering the entire seat surface to ensure high-resolution pressure distribution detection. Each piezoresistive sensing unit is connected to the 12-bit ADC channel of the STM32F407 multimodal fusion control module via a bridge circuit (such as a Wheatstone bridge). The sampled voltage is mapped to the pressure value to generate the seat surface pressure distribution. Due to the limited number of ADC channels in the STM32F407 multimodal fusion control module, an external multiplexer or ADC module (such as the ADS1118) is used for channel multiplexing, and then the 64 channels of sampled data are transmitted to the STM32F407 multimodal fusion control module at high speed (up to tens of MHz) via the SPI bus. The multimodal fusion control module performs a full array scan every 20ms with a sampling frequency of 50Hz, ensuring that force changes during cycling are captured in a timely manner and the dynamic pressure distribution is smoothly reproduced. After power-on or periodically under no-load conditions, the multimodal fusion control module acquires baseline values, corrects zero-point offsets in different units, and outputs a pressure thermogram. Since the piezoresistive material is temperature-sensitive, an on-chip temperature sensor (such as an NTC thermistor) is used to compensate for each measurement, ensuring measurement accuracy under varying ambient temperatures. Multi-point calibration (e.g., 0kg, 30kg, 60kg) obtains voltage-pressure mapping curves within the pressure range (using polynomial fitting or table lookup), improving measurement linearity.
[0043] In some embodiments of the present invention, the multimodal fusion control module is further configured to perform weighted clustering of the pressure heat map using a dual clustering algorithm, output the cluster center spacing and the intra-cluster variance ratio, and determine the occupant violation when the cluster center spacing is greater than a preset spacing or the intra-cluster variance ratio is greater than a preset ratio.
[0044] It should be noted that the piezoresistive sensing units (model HSC-50kg, range 0–50 kg, resolution 0.1 kg) are arranged in an 8×8 dot matrix below the saddle base plate. Each row is connected to the multimodal fusion control module STM32F407 (PA5-SCK, PA6-MISO pins) via a custom ribbon cable using an SPI bus. The clock is configured to be 10 MHz, and the sampling frequency is 50 Hz. The array size of the piezoresistive sensing units is 4 cm × 4 cm with a unit pitch, and they are packaged within the saddle base plate using a custom PCB.
[0045] Understandably, the multimodal fusion control module acquires the original seat surface pressure distribution through ADC sampling, then performs denoising (IIR filtering, cutoff frequency 20 Hz) and normalization to generate a 32×32 pressure heat map matrix P. To reduce noise interference caused by deformation of the saddle edge sensors, each piezoresistive sensing unit in the 8×8 pressure array is assigned a position weight. The calculation formula is as follows: ; in, The attenuation coefficient was determined through calibration tests. (Optimized signal-to-noise ratio 12dB), numerator constant 1 to avoid denominator zero, denominator calculated as Euclidean distance from sensor unit to array center (3.5, 3.5), center unit weight. Edge unit weight .
[0046] Specifically, based on the weighted stress value Generate a new heatmap input, using weighted K-means clustering (K=1 or 2). Objective function: ; in, For the k-th cluster, The cluster center. If the cluster center spacing... cm or intra-cluster variance ratio If the result is positive, it is determined to be a two-person test; otherwise, it is a single-person test. In the laboratory, in 1000 tests of single-person mode and 1000 tests of two-person mode, the single-person detection accuracy rate was 98%, the two-person detection accuracy rate was 95%, and the false positive rate was less than 0.8%.
[0047] Specifically, the collected 8×8 seat pressure distribution is denoised to generate a 32×32 pressure heatmap. Each piezoresistive sensor unit is weighted using a distance-to-center weight to obtain a weighted heatmap. This weighted heatmap is then input into a weighted K-means+DBSCAN biclustering algorithm to obtain the cluster center spacing and intra-cluster variance ratio. If the cluster center spacing is >15cm (preset spacing) and the intra-cluster variance ratio is >3 (preset ratio), it is preliminarily determined that the occupant is in violation of regulations and is a two-person load; otherwise, it is considered normal single-person driving.
[0048] In some embodiments of the present invention, the distance sensor module includes: an infrared ranging sensor and an ultrasonic module; the infrared ranging sensor and the ultrasonic module are connected to a multimodal fusion control module; the infrared ranging sensor is installed at the rear end of the vehicle and is used to measure the distance between the vehicle and the occupant as a first distance; the ultrasonic module is installed at the rear end of the vehicle and is used to measure the distance between the vehicle and the occupant as a second distance; the multimodal fusion control module is also used to fuse the first distance and the second distance through Kalman filtering to output a fused distance, and when the fused distance is less than a preset distance and the cluster center spacing is less than a preset spacing but greater than a minimum spacing, the occupant is judged to have violated the rules.
[0049] Understandably, the saddle tail integrates an infrared ranging sensor (VL53L1X) with an infrared time-of-flight (ToF) sensor (range accuracy ±1 mm, FOV 20°) and an ultrasonic module (HC-SR04) (range range 2 cm–4 m, accuracy ±3 mm). The two sensors sample synchronously (20 Hz). Infrared sensor data is acquired via I2C, while the ultrasonic module is triggered by a trigger pin and its echo pin captures the range pulse width. The pulses are then fused using a Kalman filter by the multi-modal fusion control module (STM32F407) to output a fused distance value. This fused distance value assists in determining the distance between two people when the "cluster center distance" in the pressure sensor array module is not significant.
[0050] Specifically, the VL53L1X ToF infrared ranging sensor is based on infrared laser ranging, featuring high precision and strong anti-interference capabilities, with the output ranging value being the first ranging value. The unit is millimeters. The HC-SR04 ultrasonic module emits ultrasonic pulses and measures the echo time to calculate the distance to the second ranging point. The units are either centimeters or millimeters. For effective fusion, a unit conversion is required, converting both to millimeters. A timer TIMx is used with a 20 Hz timer interrupt to drive the two sensors to sample, controlling VL53L1X and HC-SR04 to acquire ranging values within the same interrupt cycle. This ensures consistent data timing and avoids fusion errors caused by delays.
[0051] First ranging Second ranging Input a Kalman filter model, and let the state variable be the true distance. Equation of state: ; in, Let k be the actual state of the system at time k, i.e., the actual distance value (mm). This is an estimate of the true state of the system at the previous time k−1. This is process noise, reflecting the uncertainty of the state during the prediction process. The larger the value, the more "open" the filter is to sudden changes in state, and the more sensitive it is to new observations.
[0052] The observation equation is: ; in, Let H be the observation vector, with size 2×1, and let H be the observation matrix. Assume both sensors directly measure "distance" and are related to the state. Since there is a linear correspondence, H is a 2×1 vector of all 1s. To measure the noise vector, assume it to be a zero-mean Gaussian: R is the measurement noise covariance matrix, with a size of 2×2, in typical form: ; in and These are the first ranging methods. Second ranging The variance.
[0053] Prediction in STM32F407: ; in, To use the previously estimated time k The predicted state, due to the model Since the state transition matrix is 1, we can directly maintain the estimate from the previous step. The prediction error covariance represents the uncertainty of the predicted state. This is the covariance after the last update. The process noise covariance increases prediction uncertainty.
[0054] Measurement update gain calculation: ; in, This is the Kalman gain matrix, with a size of 1×2. This is the transpose of the observation matrix, with a size of 1×2. The observation prediction covariance (size 2×2) is represented by its inverse, which represents the weighted sum of measurement uncertainties.
[0055] Status Update: ; in To measure the residual, representing the difference between the actual observation and the predicted observation, the size is 1×2. To perform weighted correction of the residuals based on the Kalman gain, This is to combine the posterior state estimate after observation, i.e. the distance output after fusion in this period.
[0056] Covariance update: ; in, It is the identity matrix (1×1, i.e., scalar 1). The update factor is used to reduce covariance, reflecting the reduced uncertainty after "fusion" of observations. The posterior error covariance after fusion.
[0057] Output fusion distance: ; Among them, the process noise covariance Q and measurement noise covariance R of the Kalman filtering process were set as follows after calibration experiments in rainy and foggy weather: Output distance after fusion .
[0058] When the multimodal fusion control module determines that there are "two people to be determined" (i.e., the cluster center distance is between the minimum spacing of 10cm and the preset spacing of 15cm), it collects the riding distance output by the distance sensor module to calculate the fusion distance. If the fusion distance is less than the preset distance... When m, the occupants are deemed to have violated the rules, and the violation involves two people; otherwise, the judgment is based on the results of the seat pressure distribution.
[0059] Optionally, a comparative experiment was conducted on 500 sets of test data in a rainy environment. The single ultrasonic ranging error reached ±30 cm, the single infrared error was ±15 cm, while the ranging error after fusion was stable within ±1 cm, and the false detection rate of the assisted two-person system was reduced by 20%.
[0060] In some embodiments of the present invention, the multimodal fusion control module is used to identify the passenger's riding posture based on the riding image, and to determine the passenger's violation when a high-risk posture is detected and the confidence level is greater than the minimum confidence level threshold.
[0061] Understandably, the visual recognition module uses a miniature Sony IMX219 camera (640×480 resolution @ 30 fps, 120° wide-angle), connected to the NVIDIA Jetson Nano (4 GB RAM, Ubuntu 18.04, TensorRT 7.1 environment) of the multimodal fusion control module via a CSI interface. The Jetson Nano performs INT8 quantization inference on the model. A self-built dataset, ShareRider-1.5K, was used, containing 1500 riding videos (720p resolution, 5-10 seconds long) of vehicles operating in Changsha and Hangzhou. Violation postures were manually selected using a bounding box tool, resulting in 9000 training images and 2250 test images labeled, including four categories: "normal riding," "forward / backward straddle," "side-sitting," and "leg-hugging." A public dataset, CityStreet-Rider, was also created, containing 300 samples of nighttime and rain / fog scenes, primarily used for model generalization testing.
[0062] It should be noted that the base model for the multimodal fusion control module is MobileNet-SSD-Depthwise, with an input size of 300×300, and pre-trained weights from the COCO dataset. During training, the Adam optimizer was used with a learning rate of 1e-3, a batch size of 16, and 300 epochs. The loss function is cross-entropy loss plus bounding box regression loss. Data augmentation included spatial augmentation (random rotation ±15°, scaling 0.8-1.2), illumination augmentation (Gamma correction, Gaussian noise), and CycleGAN generative adversarial training to convert sunny images into rainy / foggy scenes. Training results show that on the test set, the mAP@0.5 (mean accuracy) is 87.3%, with mAP@0.5 reaching 87.3% in nighttime and rainy / foggy scenes; the mAP@0.5 for the un-augmented model at night is only 62.1%. The PyTorch model was exported as ONNX (a cross-platform deep learning model exchange format), and then INT8 quantization was performed using TensorRT. The model was compressed to less than 5 MB, and the average inference latency on Jetson Nano was 48 ms, which meets the requirement of less than 50 ms.
[0063] Specifically, the COCO dataset (Common Objects in Context) is an authoritative large-scale benchmark dataset in the field of computer vision, used for tasks such as object detection, instance segmentation, and keypoint detection. The Adam optimizer (Adaptive Moment Estimation) is an adaptive learning rate optimization algorithm that combines the advantages of momentum and adaptive learning rates. It is suitable for most deep learning tasks. It dynamically adjusts the learning rate of each parameter by calculating the first moment estimate (mean) and second moment estimate (uncentered variance) of the gradient, achieving adaptive optimization. In object detection tasks, cross-entropy loss and bounding box regression loss are often used together to handle classification and localization subtasks respectively. Cross-entropy loss measures the difference between the predicted class probability distribution and the true label, suitable for multi-class tasks. Bounding box regression loss measures the geometric difference between the predicted bounding box and the true box, guiding the model to accurately locate the target. CycleGAN (Cyclic Consistency Generative Adversarial Network) is an unsupervised image-to-image translation model that enables style transfer between images from two different domains without requiring paired training data. Its core lies in the cycle consistency loss mechanism, which ensures translation quality through bidirectional generation and constraints. Building and training models in PyTorch is a core part of deep learning development, offering flexibility and dynamic computation graph characteristics. TensorRT is a high-performance deep learning inference optimizer and runtime library designed to improve the inference speed and efficiency of deep learning models on GPUs, achieving low-latency, high-throughput inference capabilities. INT8 (8-bit integer) quantization is a key technique in deep learning model optimization. By converting model weights and activation values from high-precision (e.g., FP32 / FP16) to low-precision (INT8) representations, it significantly reduces computation, memory usage, and power consumption while maintaining model accuracy.
[0064] Category of violation posture:
[0065] Optionally, the visual recognition module continuously acquires images of passengers and feeds them into the MobileNet-SSD-Depthwise model of the multimodal fusion control module, outputting the category and confidence score of the violation posture. The minimum confidence threshold is set to 0.90. If high-risk postures such as "leg hugging" or "standing on the step" are detected and the confidence score is ≥0.90, they are immediately marked as violations. In nighttime and rain / fog environments, the visual recognition module uses CycleGAN data augmentation to generate synthetic samples for offline training, improving the accuracy in low-light scenes from 62.1% to 87.3%.
[0066] In some embodiments of the present invention, the multimodal fusion control module uses the pressure heatmap as a reference, aligns the riding distance and riding image through linear interpolation mapping, and outputs the violation probability using a Bi-LSTM and self-attention fusion network based on the pressure heatmap, fusion distance and riding image. When the violation probability is greater than a preset threshold, the occupant is judged to have violated the rules.
[0067] It should be noted that the pressure heatmap frequency is 50 Hz, the vehicle distance acquisition frequency is 20 Hz, and the vehicle image acquisition frequency is 30 fps. Based on the pressure heatmap data, the vehicle distance and vehicle image features are first mapped to 50 Hz using linear interpolation, and then Dynamic Time Warping (DTW) is applied with a maximum time difference of 20 ms. The aligned feature vector is input into two layers of Bi-LSTM (64 hidden units), a self-attention layer, a multi-head attention layer (8 heads), a 128-unit fully connected layer, and a Softmax layer. The output determines an occupant violation when the violation probability p_violation ≥ 0.85.
[0068] Specifically, based on the pressure heatmap data, the travel distance and travel image data are linearly interpolated in 100 ms windows (50 ms steps) and mapped to a 50 Hz sampling frequency. The aligned multimodal feature vector is as follows: (Pressure mean, pressure variance, fusion distance, number of detection boxes, average confidence, and time characteristics) Construct a Euclidean distance D matrix, and set a slope constraint of less than or equal to 2 in the DTW path search (allowing a maximum delay of 20 ms). Solve for the minimum cost path using dynamic programming. For unmatched time periods, nearest neighbor padding is performed to ensure that all modal features have consistent lengths. The alignment error is pressure-distance. The maximum frame offset of the image data was 2 frames (66 ms), which was reduced to 1 frame (33 ms) after DTW optimization.
[0069] It is understandable that the input vector of the fusion network structure The output dimension is 6; the Bi-LSTM layer consists of two bidirectional LSTM layers with 64 hidden units per layer and an output dimension of 128; the self-attention layer is a multi-head attention layer with 8 heads, used for weighted fusion of temporal features; the fully connected layer has 128 neurons activated by ReLU; the output of the Softmax layer is... The probability of violation is The loss function is cross-entropy loss, which weights positive and negative samples to alleviate the imbalance problem. ReLU (Rectified Linear Unit) is one of the most commonly used activation functions in deep learning, with a concise mathematical form and efficient computation, and is applied in convolutional neural networks (CNNs) and fully connected networks (FCNs). The Softmax layer is a core component in deep learning for multi-class classification tasks, used to transform the model's raw output (logits) into a probability distribution. Optional, when At this time, the multimodal fusion control module (STM32F407) outputs a TTL signal, which, after optocoupler isolation, drives the high-side switching MOSFET (IR2104+FQP30N06L) to cut off the power supply to the motor driver, with a response time of less than 5 ms. Simultaneously, the STM32F407 multimodal fusion control module controls the headlights to flash at 2 Hz and drives the buzzer to sound at 1 kHz via the PWM channel to alert the user. If the pressure sensor array module fails, the pressure heatmap data is bypassed, and the judgment is made directly by the distance sensor module and the vision recognition module. If the vision recognition module fails at night, the judgment is based on the combined judgment of the pressure sensor array module and the distance sensor module. If the network is temporarily unavailable, the execution module retains the speed limit and voice prompt functions, and reports data synchronously after the network is restored.
[0070] In this embodiment, in a laboratory environment, 2000 "single-person + double-person" load tests (1000 in sunny weather and 1000 in rainy weather) were conducted for comparison. The false detection rate of the single pressure clustering scheme reached 8.5%, while this system, through a three-level judgment using a pressure sensor array module + distance sensor module + visual recognition module, reduced the overall false detection rate to 0.8%. The multimodal fusion control module 200 has a fusion network inference latency of less than 50ms, pressure heatmap and vehicle distance data preprocessing and DTW alignment time of less than 20ms, and vehicle graph inference time of only about 48ms. The overall time from data acquisition to violation judgment is only 100ms, which is more than 50% faster than the traditional scheme with a judgment time of more than 200ms. The visual recognition module enhances environmental adaptability, improving detection accuracy by more than 40% in rain, fog, low light, and nighttime environments. If the pressure sensor array module fails, the pressure heat map data is bypassed, and the judgment is made directly by the distance sensor module and the vision recognition module. If the vision recognition module fails at night, the judgment is made based on the combination of the pressure sensor array module and the distance sensor module, which reduces the false alarm rate and false negative rate in multiple environments and improves the detection rate of violations when passengers ride electric vehicles.
[0071] Based on the first and / or second embodiments of this application, in the third embodiment of this application, the contents that are the same as or similar to those in the first and second embodiments described above can be referred to the above description and will not be repeated hereafter.
[0072] In some embodiments of the present invention, the system further includes: a helmet detection module; the helmet detection module is connected to the execution module; the helmet detection module is used to send a helmet violation signal to the execution module when it detects that the helmet has fallen off for more than a set time; the execution module is also used to perform an audible and visual alarm and limit the speed of the vehicle when the helmet violation signal is detected.
[0073] It is understandable that, such as Figure 5 As shown, if the helmet detection module detects that the helmet has been removed or is not being worn (HM-11 Bluetooth beacon / UHF RFID feedback), it sends a helmet violation signal to the execution module. The execution module immediately limits the vehicle speed to 5 km / h and maintains this limit for more than 200ms before returning to normal. When the execution module detects a helmet violation signal, it triggers an audible and visual alarm and limits the vehicle speed. The audible and visual alarm uses a PWM channel to control the vehicle lights to flash at 2 Hz and drives the buzzer to sound an alarm at 1 kHz.
[0074] Specifically, in the experiment, near-field communication (HM-11 Bluetooth beacon / UHF RFID) was used to detect the helmet wearing status in real time, with a detection accuracy of 99%. When the system detects that the user has removed the helmet or is not wearing it, it automatically limits the vehicle speed to 5 km / h within 200ms and activates the buzzer and flashing lights to warn the user, significantly improving safety.
[0075] In some embodiments of the present invention, the system further includes: a communication module and a cloud data analysis module; the communication module connects the multimodal fusion control module and the cloud data analysis module; the communication module is used to send violation event data to the cloud server and output voice prompts when occupant violations are detected; the cloud data analysis module is used to store violation event data, generate heat map reports of high-incidence areas and high-incidence periods through spatiotemporal clustering algorithms, and push safety education videos to occupants.
[0076] It is understandable that, such as Figure 6As shown, when the multimodal fusion control module determines that an occupant has violated regulations, the communication module uploads JSON-formatted data (including GPS coordinates, timestamps, violation type, and helmet status) to the cloud data analysis module via SIM7000G LTE Cat-1 (MQTT, QoS = 1). The cloud server receives the data, stores it in a MySQL database, and combines it with the user's app to push safety education short videos (H.264 encoded, ≤2MB), completing the "detection-execution-education" closed loop. The cloud data analysis module also periodically generates heat maps of high-incidence violation areas based on spatiotemporal clustering algorithms in the background, providing operational support for deployment and early warning decisions. SIM7000G supports LTE Cat-M1 (eMTC) and NB-IoT, and is designed specifically for low-power, wide-coverage, and low-data-rate IoT scenarios. LTE Cat-1 is a terminal device classification standard in 4G LTE networks, specifically designed for IoT scenarios.
[0077] Specifically, in the experiment, the cloud data analysis module pushed safety education short videos, with a user feedback rate of over 85%. The heat map of high-incidence areas of violations based on spatiotemporal clustering provided the operator with a basis for accurate targeting and labeling optimization, which reduced the platform's accident rate by more than 25% compared to before deployment.
[0078] In some embodiments of the present invention, the system further includes: a model update module; the model update module is connected to a cloud data analysis module, a communication module, a pressure sensor array module, a distance sensor module, a visual recognition module, a multimodal fusion control module, and an execution module; the model update module is used to receive false alarm data fed back by the occupants and update and adjust each module through transfer learning.
[0079] Understandably, the model update module collects false alarm data (such as accidental alarms and incorrect identifications) from occupants, driving parameter and model iteration across modules to form a closed-loop system of "perception-decision-feedback-optimization." Update priorities for each module are automatically assigned based on the type of false alarm. Occupants mark false alarm events (such as "alarm triggered without being touched") via the in-vehicle terminal or mobile app, and the system automatically records the timestamp, raw sensor data, and environmental parameters. Common features are extracted from multimodal data to help pinpoint the root cause of problems.
[0080] In this embodiment, when a helmet is detected as removed or not worn, the vehicle speed is immediately limited to ≤5 km / h and remains at this speed for at least 200 ms before returning to normal. Simultaneously, if a passenger is deemed to have violated regulations, the communication module publishes JSON data containing fields such as GPS coordinates, timestamp, violation type, and helmet status to the cloud data analysis module. The cloud data analysis module receives this data, stores it in a MySQL database, and combines it with a short safety education video (H.264 encoded, ≤2MB) pushed to the user's app, completing the "detection-execution-education" closed loop. Furthermore, the cloud data analysis module periodically generates heat maps of high-violation areas based on spatiotemporal clustering algorithms in the background, providing operational support for deployment and early warning decisions.
[0081] The second solution provided by this invention, a method for detecting and intervening in illegal passenger carrying on shared electric vehicles, is applied to a system for detecting and intervening in illegal passenger carrying on shared electric vehicles, as described above. Figure 7 , Figure 7 This is a flowchart illustrating an embodiment of the shared electric vehicle illegal passenger-carrying detection and intervention method proposed in this invention.
[0082] S701: Collect seat pressure distribution, riding distance, and riding images; S702. Preprocess and extract features from seat pressure distribution, riding distance, and riding images, and determine whether the passenger violated regulations through multimodal fusion; S703. In case of occupant violation, control the vehicle to cut off power and issue an alarm. S704. When the occupants resume safe driving, the violation shall be reported and safety education shall be sent to the occupants.
[0083] Understandably, after denoising the collected 8×8 seat pressure distribution, a 32×32 pressure heatmap is generated. Each piezoresistive sensing unit is then weighted using a distance-to-center weight to obtain a weighted heatmap. This weighted heatmap is then input into a weighted K-means+DBSCAN biclustering algorithm to obtain the cluster center spacing and intra-cluster variance ratio. If the cluster center spacing is >15cm (preset spacing) and the intra-cluster variance ratio is >3 (preset ratio), it is preliminarily determined to be a two-person load; otherwise, it is a single-person load.
[0084] It should be noted that when the output result based on the seat pressure distribution is uncertain (the cluster center distance is between the minimum spacing of 10cm and the preset spacing of 15cm), the passenger distance data is used. First, the first and second distance measurements are fused using Kalman filtering, and the fused distance is output. If the fused distance is less than the preset distance (0.5cm), then it is finally determined to be a double passenger; otherwise, the result of the seat pressure distribution is used.
[0085] Specifically, the visual recognition module continuously acquires images of passengers and feeds them into the MobileNet-SSD-Depthwise model, outputting the category and confidence score of the violation posture. The minimum confidence threshold is set to 0.90. If high-risk postures such as "leg hugging" or "standing on the step" are detected and the confidence score is ≥0.90, it is immediately marked as a violation. In nighttime and rain / fog environments, the visual recognition module uses CycleGAN data augmentation to generate synthetic samples for offline training, improving the accuracy in low-light scenes from 62.1% to 87.3%.
[0086] Optionally, the pressure heatmap frequency is 50 Hz, the ride distance acquisition frequency is 20 Hz, and the ride image acquisition frequency is 30 fps. Based on the pressure heatmap data, the ride distance and ride image features are first mapped to 50 Hz using linear interpolation, and then Dynamic Time Warping (DTW) is applied with a maximum time difference of 20 ms. The aligned feature vector is... (Pressure mean, pressure variance, fusion distance, number of detection boxes, average confidence, and temporal features) Input: Two layers of Bi-LSTM (64 hidden units), self-attention layer, multi-head attention (8 heads), 128-unit fully connected layer, and Softmax; Output: When the probability of violation p_violation is ≥ 0.85, the occupant is judged to have violated the rules.
[0087] If the incident is determined to be a violation by the occupants, the execution module cuts off the motor power and controls the headlights to flash and the buzzer to sound an alarm.
[0088] Understandably, when the execution module finishes determining the occupant's violation, the communication module publishes JSON data containing fields such as GPS coordinates, timestamp, violation type, and helmet status to the cloud data analysis module. The cloud data analysis module receives and stores this data, and then pushes short safety education videos to the user's app.
[0089] In this embodiment, the overall false detection rate is reduced to 0.8% through triple detection using a pressure sensor array module, a distance sensor module, and a visual recognition module. The multimodal fusion control module improves real-time performance to the millisecond level, significantly increasing response speed. The visual recognition module enhances environmental adaptability, improving detection accuracy by more than 40% in rain, fog, low light, and nighttime environments. The system detects helmet wearing status in real time with a detection accuracy of 99%. When the system detects that the user has removed the helmet or is not wearing it, it automatically limits the vehicle speed within 200ms and triggers a buzzer and flashing lights as a warning, significantly improving safety.
Claims
1. A system for detecting and intervening in illegal passenger carrying on shared electric vehicles, characterized in that, include: Pressure sensor array module, distance sensor module, vision recognition module, multimodal fusion control module, and execution module; The multimodal fusion control module is connected to the pressure sensor array module, the distance sensor module, the visual recognition module, and the execution module; The pressure sensor array modules are distributed below the saddle and are used to collect the pressure distribution on the seat surface; The distance sensor module is located at the rear of the vehicle and is used to detect the distance between the occupants and the vehicle body. The visual recognition module is used to acquire images of passengers in real time. The multimodal fusion control module is used to determine whether the passenger has violated regulations based on the seat pressure distribution, the riding distance, and the riding image. The execution module is used to disconnect the motor from the power supply when the occupant's violation is detected.
2. The shared electric vehicle illegal passenger-carrying detection and intervention system as described in claim 1, characterized in that, The pressure sensor array module includes: multiple piezoresistive sensing units; The piezoresistive sensing unit is arranged in a dot matrix under the seat cushion, and the piezoresistive sensing unit is connected to the multimodal fusion control module. The plurality of piezoresistive sensing units are used to collect sampling voltages when the occupant is riding, and to map the sampling voltages as seat pressure distributions. The multimodal fusion control module is used to receive the seat pressure distribution at preset intervals and to calibrate and compensate the seat pressure distribution to obtain a pressure heat map.
3. The shared electric vehicle illegal passenger-carrying detection and intervention system as described in claim 2, characterized in that, The multimodal fusion control module is also used to perform weighted clustering on the pressure heat map using a dual clustering algorithm, output the cluster center spacing and the intra-cluster variance ratio, and determine the occupant violation when the cluster center spacing is greater than a preset spacing or the intra-cluster variance ratio is greater than a preset ratio.
4. The shared electric vehicle illegal passenger-carrying detection and intervention system as described in claim 3, characterized in that, The distance sensor module includes: an infrared ranging sensor and an ultrasonic module; The infrared ranging sensor and the ultrasonic module are connected to the multimodal fusion control module; The infrared ranging sensor is installed at the rear end of the vehicle and is used to measure the distance between the sensor and the occupants as the first distance measurement. The ultrasonic module is installed at the rear end of the vehicle and is used to measure the distance between the vehicle and the occupants as a second distance measurement. The multimodal fusion control module is further configured to fuse the first ranging and the second ranging through Kalman filtering to output a fused distance, and to determine that the occupant violates the rules when the fused distance is less than a preset distance and the cluster center spacing is less than a preset spacing but greater than a minimum spacing.
5. The shared electric vehicle illegal passenger-carrying detection and intervention system as described in claim 1, characterized in that, The multimodal fusion control module is used to identify the passenger's sitting posture based on the riding image, and to determine that the passenger has violated the rules when a high-risk posture is detected and the confidence level is greater than the minimum confidence level threshold.
6. The shared electric vehicle illegal passenger-carrying detection and intervention system as described in claim 4, characterized in that, The multimodal fusion control module uses the pressure heatmap as a reference to align the riding distance and the riding image through linear interpolation mapping. Based on the pressure heatmap, the fusion distance, and the riding image, it uses a Bi-LSTM and self-attention fusion network to output the violation probability. When the violation probability is greater than a preset threshold, it determines that the occupant has violated the rules.
7. The shared electric vehicle illegal passenger-carrying detection and intervention system as described in claim 1, characterized in that, The system also includes: a helmet detection module; The helmet detection module is connected to the execution module; The helmet detection module is used to send a helmet violation signal to the execution module when it detects that the helmet has fallen off for more than a set time. The execution module is also used to issue an audible and visual alarm and limit the vehicle's speed when the helmet violation signal is detected.
8. The shared electric vehicle illegal passenger-carrying detection and intervention system as described in claim 1, characterized in that, The system also includes: a communication module and a cloud data analysis module; The communication module connects the multimodal fusion control module and the cloud data analysis module; The communication module is used to send violation event data to the cloud server and output voice prompts when the occupant's violation is detected. The cloud-based data analysis module is used to store the violation data, generate heat map reports of high-incidence areas and high-incidence periods through spatiotemporal clustering algorithms, and push safety education videos to passengers.
9. The shared electric vehicle illegal passenger-carrying detection and intervention system as described in claim 8, characterized in that, The system also includes: a model update module; The model update module is connected to the cloud data analysis module, the communication module, the pressure sensor array module, the distance sensor module, the visual recognition module, the multimodal fusion control module, and the execution module; The model update module is used to receive false alarm data from occupants and update and adjust each module through transfer learning.
10. A method for detecting and intervening in the illegal carrying of passengers on shared electric vehicles, characterized in that, The method, applied to the shared electric vehicle illegal passenger-carrying detection and intervention system as described in any one of claims 1 to 9, comprises: Collect seat pressure distribution, passenger distance, and passenger images; The seat pressure distribution, the riding distance, and the riding image are preprocessed and feature extracted, and multimodal fusion is used to determine whether the passenger violated the rules. When the occupant violates the rules, the vehicle's power will be cut off and an alarm will be triggered. When the occupant resumes safe driving, the violation is reported and safety education is sent to the occupant.