Smart helmet based safety cycling assistance method and system
By converting multi-source sensor data from smart helmets into a unified spatiotemporal framework for feature extraction and fusion, the problems of multi-sensor data misalignment and target detection in complex environments of smart cycling helmets are solved, achieving high-precision risk assessment and graded early warning, and improving cycling safety.
Patent Information
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- HUAZHONG NORMAL UNIV
- Filing Date
- 2026-05-21
- Publication Date
- 2026-07-24
AI Technical Summary
Existing smart cycling helmets suffer from problems such as time-series misalignment of multi-sensor data, spatial coordinate mismatch, insufficient prediction of multi-target trajectories, poor robustness of fall detection, and loose coupling of system modules in complex traffic environments. These issues lead to safety hazards such as low target detection accuracy, redundant warning information, false alarms and missed alarms, and misjudgment of fall events.
By converting multi-source sensor data into an inertial measurement unit coordinate system spatiotemporal framework, feature extraction and state estimation are performed to achieve multimodal data fusion, target motion prediction and risk quantification are carried out, and graded early warning commands are generated. Early warning is then provided in conjunction with voice and vibration motors.
It improves the accuracy and robustness of target detection in complex traffic environments, reduces false positives and false negatives, enables accurate assessment and graded early warning of cycling safety risks, and reduces the probability of accidents.
Smart Images

Figure CN122454785A_ABST
Abstract
Description
Technical Field
[0001] This invention belongs to the field of smart wearable devices and traffic safety technology, and in particular relates to a safe riding assistance method and system based on a smart helmet. Background Technology
[0002] With the rapid development of smart wearable technology, multi-sensor fusion perception technology and urban micro-mobility, two-wheeled riding, as one of the core modes of short-distance travel, has seen its safety protection technology rapidly evolve from traditional passive physical protection to proactive intelligent safety warning. This has led to the emergence of intelligent riding helmet technology that integrates environmental perception, risk warning, and status monitoring functions.
[0003] Smart cycling helmet technology can monitor the surrounding traffic environment and the rider's own condition in real time by incorporating sensing and processing units into the helmet itself. This breaks through the limitations of traditional helmets, which can only provide passive protection after a collision, and has become an important development direction in the field of cycling safety protection. Furthermore, it has led to the development of various cycling safety assistance methods based on sensor detection.
[0004] However, current cycling safety assistance methods and smart cycling helmet technologies still have many technical shortcomings in practical cycling scenarios, failing to meet the high-reliability safety protection requirements in complex traffic environments. Specifically, these shortcomings manifest in the following ways: First, multi-source sensor data lacks spatiotemporal joint calibration and deep fusion. Radar and visual information are processed independently, failing to achieve sensor performance complementarity. This leads to a significant decrease in the accuracy and stability of target detection in adverse environments such as target occlusion, drastic lighting changes, and nighttime rain or fog, easily resulting in missed or false detections. Simultaneously, the lack of a unified time reference and spatial coordinate system alignment for multi-source sensor data causes temporal misalignment, affecting the real-time performance and accuracy of subsequent perception decisions. Second, there is a lack of multi-target dynamic trajectory prediction and comprehensive risk quantification assessment mechanisms. Existing solutions only trigger warnings based on a single distance threshold, failing to jointly model and predict the motion state, relative speed, and collision time of multiple targets. This makes it impossible to quantify and prioritize the collision risks of different targets, easily leading to redundant warning information, false alarms, and missed alarms. Furthermore, the superposition of multiple warning messages can interfere with cyclists, causing secondary safety risks. Third, the robustness of fall detection algorithms is insufficient. Existing solutions mostly rely on a single acceleration threshold to determine fall behavior, which cannot effectively distinguish between normal cycling movements and actual falls, resulting in a high false trigger rate. Simultaneously, the lack of a multi-dimensional continuous fall behavior determination mechanism makes it impossible to perform temporal correlation verification of impacts, abnormal postures, and stationary states, easily leading to missed or false detections of fall events and compromising the reliable triggering of emergency assistance functions. Fourth, the entire system lacks a tightly coupled collaborative design. The modules of perception, decision-making, early warning, communication, and user interaction fail to achieve temporal alignment and state synchronization, easily causing problems such as delayed risk decisions, conflicts between early warning commands and user navigation scenarios, and functional failures in offline states. Furthermore, the lack of a collaborative architecture design at the edge, device, and cloud levels makes it impossible to simultaneously address the low power consumption and real-time requirements of embedded devices with the big data analysis and emergency management capabilities of the cloud, resulting in insufficient overall system reliability and adaptability. Summary of the Invention
[0005] Based on this, it is necessary to provide a safe riding assistance method and system based on a smart helmet that can achieve multi-sensor spatiotemporal joint calibration, multi-modal fusion perception, multi-target trajectory prediction and risk classification quantification, and multi-stage temporal confirmation of fall behavior, in order to address the above-mentioned technical problems.
[0006] Firstly, this application provides a safe riding assistance method based on a smart helmet, including:
[0007] S1. Acquire multi-source sensor data from the smart helmet, convert the multi-source sensor data into a unified spatiotemporal framework based on the inertial measurement unit coordinate system, and generate a perception unit queue; wherein, the perception unit queue is used to represent the multi-source sensor data sequence that is time-synchronized and spatially aligned with the inertial measurement unit coordinate system;
[0008] S2. Perform feature extraction and state estimation on the sensing unit queue to obtain a multi-source sensing feature set; wherein, the multi-source sensing feature set includes panoramic images, dynamic radar point clouds and helmet attitude pre-integration sequence;
[0009] S3. Input the panoramic image into the pre-built lightweight pure visual target detection model to obtain candidate region proposals and the visual feature vectors corresponding to the candidate region proposals. Project the dynamic radar point cloud onto the image plane of the panoramic image to obtain the target feature vectors of each radar point. Then fuse the target feature vectors and visual feature vectors of each radar point to obtain the set of perceived targets.
[0010] S4. Based on the set of perceived targets and the pre-integrated sequence of helmet posture, target motion prediction information is obtained; wherein, the target motion prediction information is used to characterize the predicted motion state and trajectory of each target relative to the smart helmet in a future set time period;
[0011] S5. Based on the helmet posture pre-integration sequence, the helmet's own pose is calculated. With the helmet's own pose as the origin, a dynamic risk field is constructed. Based on the target motion prediction information, the risk potential energy of each target is obtained. Among them, the risk potential energy is used to characterize the degree of collision threat posed by each target to the smart helmet based on the predicted motion state and trajectory.
[0012] S6. Sort the risk potential energy from high to low. Based on the target with the highest risk potential energy, obtain the azimuth angle of the target with the highest risk potential energy. Based on the azimuth angle and the target with the highest risk potential energy, generate a graded warning command. The graded warning command is used to instruct the voice broadcaster of the smart helmet to output the warning voice content corresponding to the azimuth angle, and to instruct the vibration motor located on the side corresponding to the azimuth angle to vibrate according to the vibration frequency and duty cycle that match the warning level.
[0013] Secondly, this application also provides a safe riding assistance system based on a smart helmet, comprising:
[0014] The spatiotemporal alignment module is used to acquire multi-source sensor data from the smart helmet, transform the multi-source sensor data into a unified spatiotemporal framework based on the inertial measurement unit coordinate system, and generate a perception unit queue. The perception unit queue is used to represent a sequence of multi-source sensor data that is time-synchronized and spatially aligned with the inertial measurement unit coordinate system.
[0015] The feature extraction module is used to extract features and estimate the state of the sensing unit queue to obtain a multi-source sensing feature set; the multi-source sensing feature set includes panoramic images, dynamic radar point clouds and helmet attitude pre-integration sequences;
[0016] The fusion perception module is used to input panoramic images into a pre-built lightweight pure vision target detection model to obtain candidate region proposals and corresponding visual feature vectors. It projects dynamic radar point clouds onto the image plane of the panoramic image to obtain target feature vectors of each radar point, and fuses the target feature vectors and visual feature vectors of each radar point to obtain a set of perceived targets.
[0017] The tracking and prediction module is used to obtain target motion prediction information based on the set of perceived targets and the pre-integrated sequence of helmet posture; wherein, the target motion prediction information is used to characterize the predicted motion state and trajectory of each target relative to the smart helmet in a future set time period;
[0018] The risk assessment module is used to calculate the helmet's own pose based on the helmet's pre-integrated sequence of postures. With the helmet's own pose as the origin, a dynamic risk field is constructed. Based on the target motion prediction information, the risk potential energy of each target is obtained. The risk potential energy is used to characterize the degree of collision threat posed by each target to the smart helmet based on the predicted motion state and trajectory.
[0019] The graded early warning module is used to sort risk potential energy from high to low, obtain the azimuth angle of the target with the highest risk potential energy, and generate graded early warning instructions based on the azimuth angle and the target with the highest risk potential energy. The graded early warning instructions are used to instruct the voice broadcaster of the smart helmet to output early warning voice content corresponding to the azimuth angle, and instruct the vibration motor located on the side corresponding to the azimuth angle to vibrate according to the vibration frequency and duty cycle matched with the early warning level.
[0020] Thirdly, this application also provides a computer device, including a memory and a processor, wherein the memory stores a computer program, and the processor executes the computer program to implement the steps of any of the methods described above.
[0021] Fourthly, this application also provides a computer-readable storage medium having a computer program stored thereon, which, when executed by a processor, implements the steps of any of the methods described above.
[0022] The aforementioned safe riding assistance method and system based on smart helmets acquires multi-source sensor data from the smart helmet and transforms this data into a unified spatiotemporal framework based on the inertial measurement unit coordinate system to generate a perception unit queue. This achieves precise time synchronization and spatial alignment of the multi-source sensor data, laying a unified spatiotemporal reference for subsequent multimodal data fusion processing and effectively solving the problem of decreased perception accuracy caused by temporal misalignment and spatial coordinate mismatch in multi-sensor data. By performing feature extraction and state estimation on the perception unit queue, a multi-source perception feature set is obtained, including panoramic images, dynamic radar point clouds, and helmet attitude pre-integration sequences. This enables multi-dimensional feature decoupling and... Accurate extraction and simultaneous screening and standardization of effective perception data provide high-quality, reliable feature input for subsequent environmental perception and risk assessment. By inputting panoramic images into a pre-built lightweight pure visual target detection model, candidate region proposals and corresponding visual feature vectors are obtained. Dynamic radar point clouds are projected onto the image plane of the panoramic image to obtain target feature vectors for each radar point. The two types of feature vectors are then fused to obtain a set of perceived targets. This leverages the advantages of accurate target category identification in visual detection and the range, velocity measurement, and environmental interference resistance of millimeter-wave radar, achieving deep fusion and complementary enhancement of multimodal perception data. This effectively improves target detection in complex cycling scenarios such as target occlusion, drastic lighting changes, and nighttime rain and fog. Accuracy and robustness are improved, reducing the problems of missed and false detections of targets. Based on the perceived target set and the helmet posture pre-integration sequence, target motion prediction information representing the future predicted motion state and trajectory of targets is obtained. This enables the prediction of the future motion trends of multiple targets in the surrounding traffic environment. By fully combining the helmet's own motion state and the historical motion patterns of targets, the accuracy of target trajectory prediction in dynamic traffic scenarios is improved. By solving the helmet's own pose based on the helmet posture pre-integration sequence, a dynamic risk field is constructed with the helmet's own pose as the origin. Combined with target motion prediction information, the risk potential energy representing the degree of target collision threat is obtained. This enables standardized and quantitative assessment of collision risks of different targets, completing multi-target collision assessments with a unified risk potential energy value. The system characterizes the degree of collision threat, breaking through the technical limitations of traditional single-distance threshold warnings and providing a reliable decision-making basis for tiered warnings. By sorting risk potential energy from high to low, it generates tiered warning commands based on the azimuth angle and risk level of the target with the highest risk potential energy. This drives the smart helmet's voice broadcaster to output warning voice content in the corresponding direction, and the vibration motor on the corresponding side vibrates according to the parameters matching the warning level. This enables targeted and tiered warnings for the highest priority risks, effectively avoiding information interference caused by the superposition of multiple target warning information. Through the dual warning mode of voice and vibration matching the direction, it improves the rider's perception efficiency of risk direction and risk level, enabling early intervention and proactive protection against cycling safety risks.This method can comprehensively improve the real-time performance and accuracy of cycling safety protection in complex traffic environments, and effectively reduce the probability of cycling accidents. Attached Figure Description
[0023] To more clearly illustrate the technical solutions in the embodiments or related technologies of this application, the accompanying drawings used in the description of the embodiments or related technologies will be briefly introduced below. Obviously, the accompanying drawings described below are only some embodiments of this application. For those skilled in the art, other drawings can be obtained based on these drawings without creative effort.
[0024] Figure 1 A flowchart illustrating a safe riding assistance method based on a smart helmet, provided as an exemplary embodiment of this application;
[0025] Figure 2 A schematic diagram of a process for generating target motion prediction information is provided as an exemplary embodiment of this application;
[0026] Figure 3 This is a schematic diagram of a safety riding assistance system based on a smart helmet, provided as an exemplary embodiment of this application. Detailed Implementation
[0027] To make the objectives, technical solutions, and advantages of this application clearer, the following detailed description is provided in conjunction with the accompanying drawings and embodiments. It should be understood that the specific embodiments described herein are merely illustrative and not intended to limit the scope of this application.
[0028] In one embodiment, such as Figure 1 As shown, a safe riding assistance method based on a smart helmet is provided. This embodiment illustrates the application of this method to a smart assistance terminal. It is understood that this method can also be applied to a smart assistance server, and further to a system including both a smart assistance terminal and a smart assistance server, and is implemented through the interaction between the smart assistance terminal and the smart assistance server. In this embodiment, the method includes the following steps:
[0029] S1. Acquire multi-source sensor data from the smart helmet, convert the multi-source sensor data into a unified spatiotemporal framework based on the inertial measurement unit coordinate system, and generate a perception unit queue.
[0030] Optionally, the multi-source sensor data of the smart helmet may include, but is not limited to, attitude motion data collected by the inertial measurement unit (IMU), image data collected by the camera module, and radar point cloud data collected by the millimeter-wave radar.
[0031] For example, the intelligent auxiliary terminal can acquire multi-source sensor data collected by the multi-source sensing devices mounted on the intelligent helmet in real time through a standardized communication interface. The intelligent auxiliary terminal can use the inertial measurement unit coordinate system as the reference coordinate system and employ the inertial-aided Kalibr (iKalibr) framework to complete the joint calibration of multiple sensors, thus achieving joint calibration of the multi-source sensing devices. Furthermore, the intelligent auxiliary terminal can acquire the spatial coordinate transformation parameters of the timestamp delay parameters of each sensor, perform timestamp alignment and spatial coordinate transformation on all multi-source sensor data based on the calibration parameters, uniformly map all data to the IMU coordinate system, and arrange them in an ordered manner according to the time sequence to generate a sensing unit queue.
[0032] Optionally, the sensing unit queue can be used to characterize a sequence of multi-source sensor data that is time-synchronized and spatially aligned to the inertial measurement unit coordinate system.
[0033] S2. Perform feature extraction and state estimation on the sensing unit queue to obtain a multi-source sensing feature set.
[0034] For example, the intelligent auxiliary terminal can perform parallel feature extraction and state estimation on different types of multi-source sensor data in the sensing unit queue. For inertial measurement unit data in the sensing unit queue, the intelligent auxiliary terminal can perform pre-integration calculations within a preset sliding time window to obtain a helmet attitude pre-integration sequence, which can be used to characterize the helmet's position, attitude, and velocity changes over continuous time. For image data in the queue, the intelligent auxiliary terminal can perform distortion correction based on pre-calibrated camera intrinsic parameters and distortion coefficients, and stitch together multi-view corrected images using a calibrated fixed transformation matrix to extract a panoramic image. For radar point cloud data in the queue, the intelligent auxiliary terminal can perform static clutter filtering in conjunction with the helmet's motion state to extract a dynamic radar point cloud. Furthermore, the intelligent auxiliary terminal can integrate the helmet attitude pre-integration sequence, panoramic image, and dynamic radar point cloud to generate a multi-source sensing feature set.
[0035] Optionally, the multi-source sensing feature set may include, but is not limited to, panoramic images, dynamic radar point clouds, and helmet pose pre-integration sequences.
[0036] S3. Input the panoramic image into the pre-built lightweight pure vision target detection model to obtain candidate region proposals and corresponding visual feature vectors. Project the dynamic radar point cloud onto the image plane of the panoramic image to obtain the target feature vectors of each radar point. Then, fuse the target feature vectors and visual feature vectors of each radar point to obtain the set of perceived targets.
[0037] For example, the intelligent auxiliary terminal can input a panoramic image from a multi-source sensing feature set into a pre-constructed lightweight pure visual object detection model. This model can perform multi-scale feature extraction and object detection on the panoramic image, outputting candidate region proposals. Further, the lightweight pure visual object detection model can perform region-of-interest alignment processing on the feature maps corresponding to the candidate regions, extracting the visual feature vectors of the corresponding candidate region proposals. Further still, the intelligent auxiliary terminal can project a dynamic radar point cloud onto the image plane of the panoramic image using a pre-calibrated extrinsic matrix, complete spatial matching, and then extract the target feature vectors of each radar point.
[0038] Optionally, the pre-built lightweight pure visual object detection model can be constructed based on the YOLO (You Only LookOnce) series of lightweight network architectures. The lightweight pure visual object detection model may include, but is not limited to, a CSPNet feature extraction layer (Cross Stage Partial Network), a multi-scale feature aggregation layer, and a candidate region generation layer.
[0039] Optionally, candidate region proposals can be used to characterize the bounding box information of the region where a potential target is located in a panoramic image.
[0040] Optionally, the visual feature vector corresponding to the candidate region proposal can be used to characterize the visual semantics and appearance features of the target within the region corresponding to the candidate region proposal.
[0041] Optionally, the target feature vector of each radar point can be used to characterize the spatial position, speed of motion, and radar scattering characteristics of the corresponding target.
[0042] Optionally, the set of perceived targets can be used to characterize the standardized perception information of various traffic targets in the surrounding environment of a cycling scenario after deep fusion of visual features and radar features. The set of perceived targets may include, but is not limited to, the target's category attributes, three-dimensional spatial coordinates, speed of movement relative to the smart helmet, detection confidence, and spatial azimuth.
[0043] S4. Based on the set of perceived targets and the pre-integrated sequence of helmet posture, target motion prediction information is obtained.
[0044] For example, the intelligent auxiliary terminal can use the set of perceived targets and the helmet posture pre-integration sequence as input, perform identity association matching on the target data of consecutive frames in the set of perceived targets, and use a multi-target tracking method that combines Kalman filtering (linear quadratic estimation) and target appearance feature joint matching to continuously update the motion state and fit the trajectory of each target in the set of perceived targets, generating a set of multi-target tracking trajectories. Furthermore, the intelligent auxiliary terminal can perform trajectory stability screening on the set of multi-target tracking trajectories, eliminating invalid trajectories with insufficient tracking time or substandard trajectory continuity, and selecting target trajectories that meet preset stability conditions. The intelligent auxiliary terminal can then combine the helmet's own motion state represented by the helmet posture pre-integration sequence to extrapolate the motion trend of the selected target trajectories over future time periods, obtaining target motion prediction information.
[0045] Optionally, the target motion prediction information can be used to characterize the predicted motion state and trajectory of each target relative to the smart helmet within a future set time period.
[0046] S5. Based on the pre-integrated sequence of helmet posture, the helmet's own pose is calculated. With the helmet's own pose as the origin, a dynamic risk field is constructed. Based on the target motion prediction information, the risk potential energy of each target is obtained.
[0047] For example, the intelligent assistance terminal can calculate the helmet's six-DOF pose in the world coordinate system based on the helmet's pre-integrated posture sequence using visual inertial odometry technology. The six-DOF pose can include the helmet's three-dimensional spatial position and three-dimensional rotational attitude. Furthermore, the intelligent assistance terminal can use the calculated helmet pose as the origin of the coordinate system and combine it with the traffic environment constraints of the cycling scenario to construct a dynamic risk field centered on the helmet. This dynamic risk field can be used to characterize the basic distribution of collision risk at different locations within the surrounding space centered on the helmet.
[0048] Furthermore, the intelligent auxiliary terminal can map the predicted motion trajectory and motion state of each target in the target motion prediction information to the dynamic risk field, and calculate the risk potential energy corresponding to each target.
[0049] Optionally, the dynamic risk field can be used to characterize a three-dimensional spatial collision risk distribution model adapted to two-wheeled riding dynamic scenarios, constructed with the real-time pose of the smart helmet as the coordinate origin. The risk field strength of the dynamic risk field can be dynamically adjusted in real time according to the distance between the spatial point and the helmet, the relative orientation, and the helmet's own movement state. The dynamic risk field can provide a unified spatial benchmark and calculation framework for the quantitative assessment of the collision threat level of different traffic targets.
[0050] Optionally, risk potential energy can be used to characterize the collision threat level quantified by each target based on the predicted motion state and trajectory to the smart helmet. The value of the collision threat level quantified can be positively correlated with the collision threat level of the target to the helmet.
[0051] S6. Sort the risk potential energy from high to low, obtain the azimuth of the target with the highest risk potential energy, and generate a graded early warning command based on the azimuth and the target with the highest risk potential energy.
[0052] For example, the intelligent auxiliary terminal can sort all calculated target risk potential energy values from highest to lowest, select the target with the highest risk potential energy at the top of the list, and designate it as the current highest priority warning target. Further, the intelligent auxiliary terminal can calculate the azimuth angle of the current highest priority warning target relative to the intelligent helmet based on its spatial location information. Further still, the intelligent auxiliary terminal can match the corresponding warning level based on the risk potential energy value of the current highest priority warning target, and combine this with the target's azimuth angle to generate a tiered warning instruction.
[0053] Optionally, the azimuth angle can be used to characterize the angular position of the target in a polar coordinate system centered on the helmet, corresponding to different spatial orientations of the helmet.
[0054] Optionally, the graded warning command can be used to instruct the smart helmet's voice broadcaster to output warning voice content corresponding to the azimuth angle, and to instruct the vibration motor located on the side corresponding to the azimuth angle to vibrate according to the vibration frequency and duty cycle matched with the warning level.
[0055] Optionally, the warning level can be a multi-level safety control level pre-divided according to the numerical range of risk potential energy. The warning level can include, but is not limited to, a safety level, a warning level, and a danger level. Different warning levels can correspond to non-overlapping risk potential energy threshold ranges, differentiated warning execution strategies, and hardware execution parameters. Among them, the safety level can correspond to the normal riding state without warning output, the warning level can correspond to low-intensity prompt warning output, and the danger level can correspond to high-intensity emergency alarm output.
[0056] The aforementioned safety riding assistance method and system based on smart helmets involves an intelligent assistance terminal acquiring multi-source sensor data from the smart helmet and converting this data into a unified spatiotemporal framework based on the inertial measurement unit coordinate system to generate a perception unit queue. This enables precise time synchronization and spatial alignment of the multi-source sensor data, laying a unified spatiotemporal reference for subsequent multimodal data fusion processing and effectively solving the problem of decreased perception accuracy caused by temporal misalignment and spatial coordinate mismatch in multi-sensor data. By performing feature extraction and state estimation on the perception unit queue, a multi-source perception feature set is obtained, including panoramic images, dynamic radar point clouds, and helmet attitude pre-integration sequences. This enables multi-dimensional perception of the surrounding riding environment and the helmet's own motion state. Feature decoupling and precise extraction are performed simultaneously to screen and standardize effective perception data, providing high-quality and reliable feature input for subsequent environmental perception and risk assessment. By inputting panoramic images into a pre-built lightweight pure visual target detection model, candidate region proposals and corresponding visual feature vectors are obtained. Dynamic radar point clouds are projected onto the image plane of the panoramic image to obtain target feature vectors for each radar point. The two types of feature vectors are then fused to obtain a set of perceived targets. This leverages the advantages of visual detection in accurate target category identification and the range, velocity measurement, and environmental interference resistance advantages of millimeter-wave radar, thereby achieving deep fusion and complementary enhancement of multimodal perception data. This effectively improves target detection performance in complex cycling scenarios such as target occlusion, drastic lighting changes, and nighttime rain and fog. This technology improves detection accuracy and robustness, reducing false and missed detections. Based on the perceived target set and helmet posture pre-integration sequence, it obtains target motion prediction information representing the future predicted motion state and trajectory of the target. This enables the prediction of the future motion trends of multiple targets in the surrounding traffic environment. By fully combining the helmet's own motion state and the historical motion patterns of the target, it improves the accuracy of target trajectory prediction in dynamic traffic scenarios. Furthermore, by calculating the helmet's own pose based on the helmet posture pre-integration sequence, a dynamic risk field is constructed with the helmet's own pose as the origin. Combined with target motion prediction information, it obtains the risk potential energy representing the degree of target collision threat. This enables standardized and quantitative assessment of collision risks for different targets, achieving multi-target collision risk assessment with a unified risk potential energy value. This system characterizes the degree of collision threat, overcoming the limitations of traditional single-distance threshold warnings and providing a reliable basis for tiered warnings. By sorting risk potential energy from high to low, it generates tiered warning commands based on the azimuth and risk level of the target with the highest risk potential energy. This drives the smart helmet's voice broadcaster to output warning voice content in the corresponding direction, and the vibration motor on the corresponding side vibrates according to the parameters matching the warning level. This enables targeted and tiered warnings for the highest priority risks, effectively avoiding information interference caused by the superposition of multiple target warning information. Furthermore, through the dual warning mode of voice and vibration matching the direction, it improves the rider's perception efficiency of risk direction and risk level, enabling early intervention and proactive protection against cycling safety risks.This method can comprehensively improve the real-time performance and accuracy of cycling safety protection in complex traffic environments, and effectively reduce the probability of cycling accidents.
[0057] In one embodiment, S2 may include:
[0058] S21. Perform pre-integration calculation on the inertial measurement unit data within the sliding window to obtain the helmet attitude pre-integration sequence.
[0059] For example, the intelligent auxiliary terminal can acquire inertial measurement unit (IMU) data from the sensing unit queue after spatiotemporal alignment. Within a preset sliding time window, it performs integration calculations on the raw measurement data of the accelerometer and gyroscope in consecutive frames based on the IMU pre-integration theory to calculate the relative motion increment between adjacent data frames. It also compensates for and corrects zero-bias errors and random noise during the IMU measurement process, suppressing error accumulation during integration. Furthermore, the intelligent auxiliary terminal can arrange the relative motion increments between adjacent data frames in each sliding window in a time sequence to generate a helmet attitude pre-integration sequence.
[0060] Optionally, the helmet attitude pre-integral sequence can be used to characterize the continuous changes in the three-dimensional spatial position, three-dimensional rotational attitude, linear velocity, and angular velocity of the smart helmet relative to the initial reference coordinate system in a continuous time dimension.
[0061] S22. Based on the pre-calibrated intrinsic parameter matrix and distortion coefficients, the image data is distorted to obtain a multi-view calibrated image. The multi-view calibrated image is then stitched together according to the pre-calibrated fixed transformation matrix to obtain a panoramic image.
[0062] For example, the intelligent auxiliary terminal can acquire multi-channel image data from the sensing unit queue after spatiotemporal alignment. Based on the pre-calibrated camera intrinsic parameter matrix and lens distortion coefficients, it performs distortion correction processing on each channel of image data to eliminate radial and tangential distortion caused by wide-angle lenses, thereby obtaining multi-view calibrated images. Furthermore, the intelligent auxiliary terminal can perform feature matching and projection transformation on the multi-view calibrated images based on the pre-calibrated fixed transformation matrix between the multi-channel cameras, and uniformly map the multi-view calibrated images to the same plane coordinate system to complete seamless image stitching and generate a panoramic image covering a set field of view around the helmet.
[0063] Optionally, the multi-view corrected image can be used to characterize the distortion-free, pixel coordinate-normalized single-view image acquired by each camera after distortion correction.
[0064] Optionally, the pre-calibrated fixed transformation matrix can be an extrinsic parameter matrix obtained through multi-camera joint calibration, used to characterize the spatial pose transformation relationship between different camera coordinate systems, and may include rotation matrix and translation vector.
[0065] S23. Perform kinematic calculations on the inertial measurement unit data to obtain a helmet velocity estimate, and perform static clutter filtering on the radar point cloud data based on the velocity estimate to obtain a dynamic radar point cloud.
[0066] For example, the intelligent auxiliary terminal can perform rigid body kinematics calculations on the inertial measurement unit data in the sensing unit queue after spatiotemporal alignment. Combining the pre-integration calculation results, it can calculate the helmet's velocity in three-dimensional space, i.e., the helmet velocity estimate. Furthermore, based on the helmet velocity estimate, the intelligent auxiliary terminal can calculate the theoretical radial velocity corresponding to each radar point in the radar point cloud data, and obtain the velocity residual by subtracting the theoretical radial velocity from the radar's measured Doppler radial velocity.
[0067] Furthermore, the intelligent auxiliary terminal can combine the radar cross section (RCS) threshold to filter out radar points whose velocity residuals and radar cross sections meet the preset static conditions as clutter, while retaining the effective radar points that characterize moving targets, thus obtaining a dynamic radar point cloud.
[0068] Optionally, the helmet velocity estimation can be used to characterize the three-dimensional linear velocity vector of the smart helmet relative to the ground in the world coordinate system, reflecting the real-time motion state and travel trend of the smart helmet itself.
[0069] Preferably, the expression for helmet speed estimation can be:
[0070]
[0071] In the formula, This indicates the helmet's speed estimate. This represents the initial linear velocity of the helmet at the start of the integration. This represents the rotation matrix from the helmet-mounted coordinate system to the world coordinate system. This represents the specific force measurement value in the body coordinate system acquired by the inertial measurement unit. Represents the local gravitational acceleration vector. Indicates the starting time of the integration operation. Indicates the time at which the integration operation terminates.
[0072] S24. Combine the helmet attitude pre-integration sequence, panoramic image and dynamic radar point cloud to obtain a multi-source perception feature set.
[0073] For example, the intelligent auxiliary terminal can perform frame synchronization matching of the helmet attitude pre-integration sequence, panoramic image, and dynamic radar point cloud based on the timestamp under a unified spatiotemporal framework. The intelligent auxiliary terminal performs structured encapsulation of the matched multi-source data and adds a unified index identifier and time tag to each set of synchronized data, completing the orderly integration and processing of multi-dimensional perception data and generating a standardized multi-source perception feature set.
[0074] In this embodiment, the intelligent auxiliary terminal constructs a spatiotemporally synchronized and highly robust multimodal perception feature set by performing parallel preprocessing and deep fusion of multi-source sensor data, providing standardized and high-quality feature inputs for subsequent target detection, trajectory prediction and dynamic risk quantification.
[0075] In one embodiment, S23 may include:
[0076] S231. Obtain the three-dimensional position coordinates of each radar point in the radar point cloud data in the coordinate system of the inertial measurement unit corresponding to the unified spatiotemporal frame.
[0077] For example, the intelligent auxiliary terminal can acquire spatiotemporally aligned radar point cloud data from the sensing unit queue. Based on the spatial coordinate transformation parameters obtained from the multi-sensor joint calibration in step S1, the intelligent auxiliary terminal can transform each radar point in the radar point cloud data from the radar device's own coordinate system to the inertial measurement unit coordinate system corresponding to the unified spatiotemporal framework, and calculate the three-dimensional position coordinates of each radar point in this coordinate system. Furthermore, the intelligent auxiliary terminal can perform validity verification on the three-dimensional position coordinates of each radar point, discarding invalid radar points whose coordinate values exceed the preset effective detection range, thus completing the extraction and storage of the valid radar point three-dimensional position coordinates.
[0078] S232. Based on the helmet velocity estimation and three-dimensional position coordinates, calculate the theoretical radial velocity of each radar point, and subtract the theoretical radial velocity from the corresponding Doppler velocity observation value of the radar point to obtain the velocity residual.
[0079] For example, the intelligent auxiliary terminal can calculate the theoretical radial velocity of each radar point relative to the radar device based on the helmet velocity estimate obtained from the solution, combined with the three-dimensional position coordinates of each radar point in the inertial measurement unit coordinate system, and through rigid body kinematic projection relationships. Furthermore, the intelligent auxiliary terminal can extract the Doppler velocity observation value corresponding to each radar point from the raw measurement information of the radar point cloud data, and perform a difference operation between the theoretical radial velocity and the Doppler velocity observation value corresponding to the same radar point to obtain the velocity residual corresponding to each radar point.
[0080] Optionally, the theoretical radial velocity of each radar point can be used to characterize the theoretical radial velocity of a static object at a corresponding spatial position in the radar coordinate system relative to the radar device when the helmet is in motion.
[0081] Optionally, the Doppler velocity observation value corresponding to the radar point can be obtained by millimeter-wave radar through the Doppler effect, and the radial velocity measurement value of the reflecting object relative to the radar equipment can be used.
[0082] S233. Radar points with velocity residuals less than the first preset threshold and radar cross section values less than the second preset threshold are filtered out as static clutter points to obtain the remaining radar points after filtering, and the remaining radar points are used as dynamic radar point clouds.
[0083] For example, the intelligent auxiliary terminal can acquire the velocity residual and radar cross section measurement value corresponding to each radar point. Further, the intelligent auxiliary terminal can compare the velocity residual with a predefined first preset threshold and the radar cross section measurement value with a predefined second preset threshold. The intelligent auxiliary terminal can mark radar points that simultaneously satisfy the condition of a velocity residual less than the first preset threshold and a radar cross section value less than the second preset threshold as static clutter points, remove all marked static clutter points from the radar point cloud data, retain the remaining valid radar points, and use the set of remaining valid radar points after filtering as the dynamic radar point cloud.
[0084] Optionally, the first preset threshold can be used to characterize the maximum allowable velocity residual limit when the radar point and the smart helmet remain relatively stationary.
[0085] Optionally, the second preset threshold can be used to characterize the minimum radar cross section threshold for filtering weak environmental reflection interference.
[0086] In this embodiment, the intelligent auxiliary terminal achieves accurate filtering of static background clutter and effective separation of dynamic targets by using dual criteria based on the velocity residual of the helmet's motion state and the radar cross section. This significantly improves the accuracy of dynamic target identification and the robustness of the perception algorithm in radar point cloud data under complex cycling scenarios.
[0087] In one embodiment, such as Figure 2 As shown, S4 may include:
[0088] S41. Based on the set of perceived targets, a multi-target tracking method using Kalman filtering and joint matching of appearance features is used to perform identity association, thereby obtaining a set of multi-target tracking trajectories.
[0089] For example, the intelligent auxiliary terminal can, based on a set of perceived targets, take the set of perceived targets in consecutive time frames as input, and use Kalman filtering to recursively predict and update the motion state of each target, thus obtaining the motion state vector of each target. Further, the intelligent auxiliary terminal can extract the visual appearance features corresponding to each target and construct a feature matching matrix. The intelligent auxiliary terminal can combine motion information and appearance features to complete target identity association between consecutive frames, assign a unique identity identifier to each target, integrate continuous observation data of targets with the same identity according to the time series, and fit to generate continuous tracking trajectories. Finally, it can summarize all target trajectories to obtain a multi-target tracking trajectory set.
[0090] Optionally, the multi-target tracking trajectory set can be used to characterize the continuous trajectory data set representing the spatial location and motion state changes of each moving target in the surrounding environment of the cycling scene over a continuous time dimension. The multi-target tracking trajectory set may include, but is not limited to, the unique identifier of each target, temporal spatial coordinates, motion parameters, and tracking confidence information.
[0091] S42. Select target trajectories that meet the preset stability conditions from the set of multi-target tracking trajectories.
[0092] For example, the intelligent auxiliary terminal can, based on a multi-target tracking trajectory set, extract the continuous tracking duration, time-series trajectory point missing rate, position observation fluctuation amplitude, and average tracking confidence value for each target trajectory within the set. Furthermore, the intelligent auxiliary terminal can compare and verify each characteristic parameter of each target trajectory against preset stability conditions, eliminating invalid and interfering trajectories that do not meet the preset stability conditions, and retaining all target trajectories that meet the preset stability conditions, thus completing the effectiveness screening and purification process for the target trajectories.
[0093] Optionally, the preset stability conditions can be quantitative judgment rules for determining the effectiveness and stability of the target tracking trajectory. The preset stability conditions may include, but are not limited to, the continuous tracking duration of the trajectory being greater than a preset duration threshold, the trajectory point missing rate being less than a preset missing threshold, and the mean tracking confidence being greater than a preset confidence threshold.
[0094] S43. Perform interactive multi-model trajectory prediction on the helmet attitude pre-integration sequence and the target trajectory to obtain target motion prediction information.
[0095] For example, the intelligent auxiliary terminal can calculate the real-time motion state and travel trend of the helmet itself based on the helmet posture pre-integration sequence, and use the real-time motion state and travel trend of the helmet itself as the reference state for target relative motion analysis.
[0096] Furthermore, the intelligent auxiliary terminal can construct a set of motion models adapted to different traffic target motion patterns for each target, and update and interact with the matching probabilities of each motion model set in real time. The intelligent auxiliary terminal can combine historical trajectory data of the target with the helmet's own motion state to infer the target's motion state and spatial position changes within a set future time period, generating target motion prediction information.
[0097] Optionally, the target motion prediction information can be used to characterize the predicted results of the relative motion state, temporal spatial position, and motion trajectory of each target relative to the smart helmet within a set time period in the future. The target motion prediction information may include, but is not limited to, the predicted coordinates, relative velocity, direction of motion, and prediction confidence information of the target at each future moment.
[0098] In this embodiment, the intelligent auxiliary terminal achieves accurate prediction of the future movement trend of dynamic targets by integrating appearance features and motion filtering for stable trajectory screening and interactive multi-model prediction, which can provide a reliable time-dimensional decision basis for high-risk collision warning.
[0099] In one embodiment, step S3, which involves inputting the panoramic image into a pre-built lightweight pure visual object detection model to obtain candidate region proposals and corresponding visual feature vectors, may include:
[0100] S31. The panoramic image is downsampled stepwise and connected across stages through the CSPNet feature extraction layer to obtain initial feature maps at different scales.
[0101] For example, the intelligent auxiliary terminal can input panoramic images into a pre-constructed lightweight pure vision object detection model. Further, the CSPNet feature extraction layer of the lightweight pure vision object detection model can receive panoramic images. The CSPNet feature extraction layer can perform progressive downsampling of the panoramic image through multiple convolutional operations, gradually reducing the feature map space size while expanding the feature receptive field. It also splits the input feature channels into two branches for feature extraction and gradient propagation respectively, reducing computational cost and feature redundancy while mitigating the gradient vanishing problem in deep networks. This preserves the basic visual information of the panoramic image, such as texture and contour, and outputs multiple sets of initial feature maps at different resolutions.
[0102] Optionally, initial feature maps of different scales can be used to represent visual feature information at different levels in the panoramic image. Large-size feature maps carry detailed features of the target, adapting to the needs of small-size target detection, while small-size feature maps carry deep semantic features of the target, adapting to the needs of large-size target detection. Detailed features may include, but are not limited to, the edges and textures of the target; deep semantic features may include, but are not limited to, the category and global structure of the target.
[0103] S32. The initial feature map is subjected to top-down path aggregation and bottom-up feature enhancement through a multi-scale feature aggregation layer to obtain enhanced feature maps corresponding to different scales.
[0104] For example, the multi-scale feature aggregation layer can receive initial feature maps from the CSPNet feature extraction layer. Further, the multi-scale feature aggregation layer can construct a feature pyramid structure via a top-down path, progressively passing the deep semantic information of high-level feature maps downwards and fusing it with low-level feature maps of the same scale. It can also construct an enhancement pyramid structure via a bottom-up path, progressively passing the precise location information of low-level feature maps upwards, enhancing the spatial localization capability of high-level features. Through bidirectional feature fusion and complementarity, it completes the enhancement processing of initial feature maps at different scales, generating enhanced feature maps corresponding to different scales.
[0105] Optionally, the enhanced feature maps corresponding to different scales can be used to characterize multi-dimensional visual features enhanced by both semantic and positional information.
[0106] S33. Perform bounding box regression and object property prediction on each enhanced feature map through the candidate region generation layer to obtain the prediction score. Select a preset number of candidate region proposals based on the prediction score, and align the enhanced feature map with the region of interest based on each candidate region proposal to obtain the visual feature vector.
[0107] For example, the candidate region generation layer can receive enhanced feature maps from the multi-scale feature aggregation layer. Further, the candidate region generation layer can perform region-by-region processing on the enhanced feature maps using sliding window convolution operations, and perform regression fitting of the target bounding box coordinates and object property prediction of the probability of object presence within the region, outputting a prediction score for each candidate region.
[0108] Preferably, the expression for the predicted score corresponding to each candidate region can be:
[0109]
[0110] In the formula, Indicates the first The predicted score corresponding to each candidate region Indicates the first The object prediction confidence of the nth candidate region is used to characterize the nth candidate region. The probability that a valid traffic target related to cycling safety exists within each candidate region. Indicates the first The maximum class prediction confidence of the candidate region is the first The candidate region is the maximum value of the predicted probability of the candidate region under all preset traffic target categories. Indicates the first The bounding box regression quality score of each candidate region is used to characterize the matching accuracy between the bounding box of the candidate region obtained by regression fitting and the real spatial contour of the target.
[0111] Furthermore, the candidate region generation layer can sort the candidate regions according to their prediction scores from high to low, and select a preset number of high-confidence candidate region proposals. Further, the candidate region generation layer can use the coordinates of each candidate region proposal as a reference to align the enhanced feature map with the Region of Interest (ROI), eliminate quantization errors, and normalize the features of regions of different sizes into fixed-dimensional feature representations, outputting a standardized visual feature vector.
[0112] Optionally, the prediction score can be used to characterize the confidence level that there are valid traffic targets related to cycling safety in the corresponding candidate area, and the score value is positively correlated with the probability that the candidate area is a valid target area.
[0113] Optionally, the visual feature vector can be used to characterize the high-dimensional visual semantic features of the target within the candidate region. The visual feature vector may include, but is not limited to, the target's category attributes, appearance contours, and spatial structure.
[0114] In this embodiment, the intelligent auxiliary terminal inputs panoramic images into a pre-constructed lightweight pure visual target detection model. The lightweight pure visual target detection model performs efficient feature extraction and bidirectional multi-scale feature aggregation, which significantly reduces computational redundancy while enhancing target semantics and spatial positioning capabilities, thereby providing accurate and robust visual feature representation for subsequent multimodal fusion.
[0115] In one embodiment, the expression for the risk potential energy can be:
[0116]
[0117]
[0118] In the formula, Indicates risk potential energy. This represents the Euclidean distance between the target and the smart helmet. This represents the radial velocity of the target relative to the smart helmet. This indicates that the radial velocity is taken as a positive value. Represented by natural constant An exponentially decaying function with base 0. Indicates the distance attenuation coefficient. This indicates the collision time between the target and the smart helmet. This represents the weighting coefficient of the distance term. This represents the weighting coefficient for the approximation velocity term. This represents the weighting coefficient for the collision time term. Represents the absolute value of radial velocity. This indicates the preset minimum speed threshold. This indicates taking the maximum value within the parentheses.
[0119] For example, the intelligent auxiliary terminal can acquire the Euclidean distance between each target and the intelligent helmet, and the radial velocity of the target relative to the intelligent helmet, obtained through trajectory prediction processing. Further, based on the Euclidean distance, the absolute value of the radial velocity, and a preset minimum velocity threshold, the intelligent auxiliary terminal can calculate the collision time corresponding to each target using the Time To Collision (TTC) calculation formula. Further, the intelligent auxiliary terminal can calculate the distance term, approach velocity term, and collision time term in the risk potential energy expression separately, and combine pre-configured weighting coefficients for the distance term, approach velocity term, collision time term, and distance attenuation coefficient to complete the quantitative calculation of the risk potential energy corresponding to each target.
[0120] In this embodiment, the intelligent auxiliary terminal constructs a multi-factor risk potential quantification expression that integrates the target distance dimension, the relative approach speed dimension, and the collision time dimension. This enables standardized and refined quantitative assessment of the collision threat level of various traffic targets in cycling scenarios. By using configurable weight coefficients to adapt to the risk assessment needs of different cycling scenarios, and by using an exponential decay function to suppress the risk interference of distant targets, it overcomes the technical limitations of traditional single distance threshold warnings. It can distinguish the collision risk levels of different targets, providing accurate and quantifiable core decision-making basis for subsequent high-priority risk screening and generation of graded warning instructions, effectively improving the accuracy and robustness of risk assessment in complex traffic scenarios.
[0121] The aforementioned safety riding assistance method and system based on smart helmets utilizes a smart auxiliary terminal to construct a unified spatiotemporal framework for sensing unit queues through spatiotemporal alignment of multi-source sensor data. Parallel feature extraction and state estimation generate a multi-source sensing feature set. Multimodal target perception is achieved through lightweight visual detection and radar projection fusion. Target trajectory prediction is completed using Kalman filtering and interactive multi-models. Collision threats are quantified through a multi-factor risk potential energy model. Finally, a graded directional warning command is generated based on the azimuth angle of the highest-risk target. This technical solution effectively addresses the technical problems of spatiotemporal misalignment of multi-source data, low detection accuracy in complex scenes due to independent processing of radar and vision, and the susceptibility to false alarms and missed alarms and information redundancy in single-threshold warnings in existing technologies. It achieves a tightly coupled design across the entire link of perception, decision-making, and warning, significantly improving the real-time performance and robustness of riding safety protection in complex traffic scenarios and reducing the probability of riding accidents.
[0122] It should be understood that although the steps in the flowcharts of the embodiments described above are shown sequentially according to the arrows, these steps are not necessarily executed in the order indicated by the arrows. Unless explicitly stated herein, there is no strict order restriction on the execution of these steps, and they can be executed in other orders. Moreover, at least some steps in the flowcharts of the embodiments described above may include multiple steps or multiple stages. These steps or stages are not necessarily completed at the same time, but can be executed at different times. The execution order of these steps or stages is not necessarily sequential, but can be performed alternately or in turn with other steps or at least some of the steps or stages of other steps.
[0123] Based on the same inventive concept, this application also provides a system for implementing the aforementioned smart helmet-based safe riding assistance method. The solution provided by this system is similar to the implementation described in the above method; therefore, the specific limitations of one or more smart helmet-based safe riding assistance system embodiments provided below can be found in the above-described limitations of a smart helmet-based safe riding assistance method, and will not be repeated here.
[0124] In one exemplary embodiment, such as Figure 3 As shown, a smart helmet-based safe riding assistance system 60 is provided, comprising:
[0125] The spatiotemporal alignment module 61 can be used to acquire multi-source sensor data from the smart helmet, convert the multi-source sensor data into a unified spatiotemporal framework based on the inertial measurement unit coordinate system, and generate a perception unit queue; wherein, the perception unit queue is used to represent a sequence of multi-source sensor data that is time-synchronized and spatially aligned with the inertial measurement unit coordinate system.
[0126] The feature extraction module 62 can be used to extract features and estimate the state of the sensing unit queue to obtain a multi-source sensing feature set; wherein, the multi-source sensing feature set includes panoramic images, dynamic radar point clouds and helmet attitude pre-integration sequences;
[0127] The fusion perception module 63 can be used to input panoramic images into a pre-built lightweight pure visual target detection model to obtain candidate region proposals and corresponding visual feature vectors, project dynamic radar point clouds onto the image plane of the panoramic image to obtain target feature vectors of each radar point, and fuse the target feature vectors and visual feature vectors of each radar point to obtain a set of perceived targets.
[0128] The tracking and prediction module 64 can be used to obtain target motion prediction information based on the set of perceived targets and the pre-integrated sequence of helmet posture; wherein, the target motion prediction information is used to characterize the predicted motion state and trajectory of each target relative to the smart helmet in a future set time period.
[0129] The risk assessment module 65 can be used to calculate the helmet's own pose based on the helmet's pre-integrated sequence of postures, construct a dynamic risk field with the helmet's own pose as the origin, and obtain the risk potential energy of each target based on the target motion prediction information. Among them, the risk potential energy is used to characterize the degree of collision threat posed by each target to the smart helmet based on the predicted motion state and trajectory.
[0130] The graded early warning module 66 can be used to sort risk potential energy from high to low, obtain the azimuth angle of the target with the highest risk potential energy, and generate graded early warning instructions based on the azimuth angle and the target with the highest risk potential energy. The graded early warning instructions are used to instruct the voice broadcaster of the smart helmet to output early warning voice content corresponding to the azimuth angle, and instruct the vibration motor located on the side corresponding to the azimuth angle to vibrate according to the vibration frequency and duty cycle matched with the early warning level.
[0131] In one embodiment, the feature extraction module includes:
[0132] The attitude pre-integration unit can be used to pre-integrate the inertial measurement unit data within a sliding window to obtain the helmet attitude pre-integration sequence.
[0133] The panoramic stitching unit can be used to perform distortion correction on image data based on a pre-calibrated intrinsic parameter matrix and distortion coefficients to obtain multi-view calibrated images, and then stitch the multi-view calibrated images according to a pre-calibrated fixed transformation matrix to obtain a panoramic image.
[0134] The dynamic point cloud extraction unit can be used to perform kinematic calculations on inertial measurement unit data to obtain helmet velocity estimates, and to perform static clutter filtering on radar point cloud data based on the velocity estimates to obtain dynamic radar point clouds.
[0135] The feature combination unit can be used to combine helmet attitude pre-integration sequences, panoramic images, and dynamic radar point clouds to obtain a multi-source perception feature set.
[0136] In one embodiment, the dynamic point cloud extraction unit includes:
[0137] The coordinate transformation subunit can be used to obtain the three-dimensional position coordinates of each radar point in the radar point cloud data in the coordinate system of the inertial measurement unit corresponding to the unified spatiotemporal frame;
[0138] The velocity residual calculation subunit can be used to calculate the theoretical radial velocity of each radar point based on helmet velocity estimation and three-dimensional position coordinates, and then subtract the theoretical radial velocity from the corresponding Doppler velocity observation value of the radar point to obtain the velocity residual.
[0139] The static-dynamic separation subunit can be used to filter out radar points with velocity residuals less than a first preset threshold and radar cross-section values less than a second preset threshold as static clutter points, obtaining the remaining radar points after filtering, and using the remaining radar points as a dynamic radar point cloud; wherein, the first preset threshold is used to characterize the maximum allowable velocity residual limit when the radar point and the smart helmet remain relatively stationary; the second preset threshold is used to characterize the minimum radar cross-section threshold for filtering weak reflection interference in the environment.
[0140] In one embodiment, the tracking prediction module includes:
[0141] The multi-target tracking unit can be used to perform identity association based on a set of perceived targets, using a multi-target tracking method that employs Kalman filtering and joint matching of appearance features, to obtain a set of multi-target tracking trajectories;
[0142] The trajectory filtering unit can be used to filter out target trajectories that meet preset stability conditions from a set of multi-target tracking trajectories;
[0143] The trajectory prediction unit can be used to perform interactive multi-model trajectory prediction on the helmet attitude pre-integration sequence and the target trajectory to obtain target motion prediction information.
[0144] In one embodiment, the fusion sensing module can also be used for:
[0145] By using the CSPNet feature extraction layer to perform stepwise downsampling and cross-stage local connections on the panoramic image, initial feature maps at different scales are obtained.
[0146] The initial feature map is subjected to top-down path aggregation and bottom-up feature enhancement through a multi-scale feature aggregation layer to obtain enhanced feature maps corresponding to different scales.
[0147] The candidate region generation layer performs bounding box regression and object property prediction on each enhanced feature map to obtain a prediction score. Based on the prediction score, a preset number of candidate region proposals are selected. Based on each candidate region proposal, the enhanced feature map is aligned with the region of interest to obtain a visual feature vector.
[0148] In one embodiment, the risk assessment module can also be used for:
[0149] The expression for risk potential energy is:
[0150]
[0151]
[0152] In the formula, Indicates risk potential energy. This represents the Euclidean distance between the target and the smart helmet. This represents the radial velocity of the target relative to the smart helmet. This indicates that the radial velocity is taken as a positive value. Represented by natural constant An exponentially decaying function with base 0. Indicates the distance attenuation coefficient. This indicates the collision time between the target and the smart helmet. This represents the weighting coefficient of the distance term. This represents the weighting coefficient for the approximation velocity term. This represents the weighting coefficient for the collision time term. Represents the absolute value of radial velocity. This indicates the preset minimum speed threshold. This indicates taking the maximum value within the parentheses.
[0153] In one embodiment, a computer device is provided, including a memory and a processor, the memory storing a computer program, the processor executing the computer program to implement the steps of the aforementioned smart helmet-based safe riding assistance method.
[0154] In one embodiment, a computer-readable storage medium is provided having a computer program stored thereon, which, when executed by a processor, implements the steps in the above method embodiments.
[0155] For the device embodiments, since they basically correspond to the method embodiments, the relevant parts can be referred to in the description of the method embodiments. The device embodiments described above are merely illustrative. The components described as separate parts may or may not be physically separate, and the components shown as units may or may not be physical units, that is, they may be located in one place or distributed across multiple network units. Some or all of the modules can be selected to achieve the purpose of this disclosure according to actual needs. Those skilled in the art can understand and implement this without creative effort.
[0156] The above-described embodiments are merely illustrative of several implementation methods of the embodiments of this application, and their descriptions are relatively specific and detailed. However, they should not be construed as limiting the scope of the patent application. It should be noted that those skilled in the art can make various modifications and improvements without departing from the concept of the embodiments of this application, and these modifications and improvements all fall within the protection scope of the embodiments of this application.
Claims
1. A safe riding assistance method based on a smart helmet, characterized in that, The method includes: S1. Acquire multi-source sensor data from the smart helmet, and convert the multi-source sensor data into a unified spatiotemporal framework based on the inertial measurement unit coordinate system to generate a perception unit queue; wherein, the perception unit queue is used to represent a multi-source sensor data sequence that is time-synchronized and spatially aligned with the inertial measurement unit coordinate system. S2. Perform feature extraction and state estimation on the sensing unit queue to obtain a multi-source sensing feature set; wherein, the multi-source sensing feature set includes panoramic images, dynamic radar point clouds, and helmet attitude pre-integration sequences; S3. Input the panoramic image into a pre-constructed lightweight pure visual target detection model to obtain candidate region proposals and visual feature vectors corresponding to the candidate region proposals. Project the dynamic radar point cloud onto the image plane of the panoramic image to obtain the target feature vectors of each radar point. Then fuse the target feature vectors and visual feature vectors of each radar point to obtain a set of perceived targets. S4. Based on the set of perceived targets and the pre-integrated sequence of helmet posture, target motion prediction information is obtained; wherein, the target motion prediction information is used to characterize the predicted motion state and trajectory of each target relative to the smart helmet in a future set time period; S5. Based on the helmet posture pre-integration sequence, the helmet's own pose is calculated. A dynamic risk field is constructed with the helmet's own pose as the origin. Based on the target motion prediction information, the risk potential energy of each target is obtained. The risk potential energy is used to characterize the degree of collision threat posed by each target to the smart helmet based on the predicted motion state and the trajectory. S6. Sort the risk potential energy from high to low, obtain the azimuth angle of the target with the highest risk potential energy, and generate a graded warning instruction based on the azimuth angle and the target with the highest risk potential energy; wherein, the graded warning instruction is used to instruct the voice broadcaster of the smart helmet to output warning voice content corresponding to the azimuth angle, and instruct the vibration motor located on the side corresponding to the azimuth angle to vibrate according to the vibration frequency and duty cycle matching the warning level.
2. The method according to claim 1, characterized in that, The sensing unit queue includes inertial measurement unit data, image data, and radar point cloud data; The S2 includes: S21. Perform pre-integration calculation on the inertial measurement unit data within the sliding window to obtain the helmet attitude pre-integration sequence; S22. Based on the pre-calibrated intrinsic parameter matrix and distortion coefficients, the image data is distorted to obtain a multi-view calibrated image, and the multi-view calibrated image is stitched together according to the pre-calibrated fixed transformation matrix to obtain the panoramic image. S23. Perform kinematic calculations on the inertial measurement unit data to obtain a helmet velocity estimate, and perform static clutter filtering on the radar point cloud data based on the velocity estimate to obtain the dynamic radar point cloud; S24. Combine the helmet attitude pre-integration sequence, the panoramic image, and the dynamic radar point cloud to obtain the multi-source perception feature set.
3. The method according to claim 2, characterized in that, S23 includes: S231. Obtain the three-dimensional position coordinates of each radar point in the radar point cloud data under the coordinate system of the inertial measurement unit corresponding to the unified spatiotemporal frame; S232. Based on the helmet velocity estimation and the three-dimensional position coordinates, calculate the theoretical radial velocity of each radar point, and subtract the theoretical radial velocity from the Doppler velocity observation value corresponding to the radar point to obtain the velocity residual. S233. Radar points with velocity residuals less than a first preset threshold and radar cross-section values less than a second preset threshold are filtered out as static clutter points to obtain the remaining radar points after filtering, and the remaining radar points are used as the dynamic radar point cloud; wherein, the first preset threshold is used to characterize the maximum allowable velocity residual limit when the radar point and the smart helmet are relatively stationary; the second preset threshold is used to characterize the minimum radar cross-section threshold for filtering weak reflection interference in the environment.
4. The method according to claim 1, characterized in that, The S4 includes: S41. Based on the set of perceived targets, a multi-target tracking method using Kalman filtering and joint matching of appearance features is used to perform identity association, thereby obtaining a set of multi-target tracking trajectories. S42. Select target trajectories that meet preset stability conditions from the multi-target tracking trajectory set; S43. Perform interactive multi-model trajectory prediction on the helmet posture pre-integration sequence and the target trajectory to obtain the target motion prediction information.
5. The method according to claim 1, characterized in that, The lightweight pure visual object detection model includes a CSPNet feature extraction layer, a multi-scale feature aggregation layer, and a candidate region generation layer. In step S3, the panoramic image is input into a pre-constructed lightweight pure visual object detection model to obtain candidate region proposals and corresponding visual feature vectors, including: S31. The panoramic image is downsampled stepwise and connected across stages through the CSPNet feature extraction layer to obtain initial feature maps at different scales; S32. The initial feature map is subjected to top-down path aggregation and bottom-up feature enhancement through the multi-scale feature aggregation layer to obtain enhanced feature maps corresponding to different scales. S33. The candidate region generation layer performs bounding box regression and object property prediction on each of the enhanced feature maps to obtain a prediction score. A preset number of candidate region proposals are selected based on the prediction scores. Based on each candidate region proposal, the enhanced feature maps are aligned with regions of interest to obtain the visual feature vector.
6. The method according to claim 1, characterized in that, The expression for the risk potential energy is: In the formula, This represents the risk potential energy. This represents the Euclidean distance between the target and the smart helmet. This represents the radial velocity of the target relative to the smart helmet. This indicates that the radial velocity is taken as a positive value. Represented by natural constant An exponentially decaying function with base 0. Indicates the distance attenuation coefficient. This indicates the collision time between the target and the smart helmet. This represents the weighting coefficient of the distance term. This represents the weighting coefficient for the approximation velocity term. This represents the weighting coefficient for the collision time term. This represents the absolute value of the radial velocity. This indicates the preset minimum speed threshold. This indicates taking the maximum value within the parentheses.
7. A safety riding assistance system based on a smart helmet, characterized in that, The system includes: The spatiotemporal alignment module is used to acquire multi-source sensor data from the smart helmet, convert the multi-source sensor data into a unified spatiotemporal framework based on the inertial measurement unit coordinate system, and generate a perception unit queue; wherein, the perception unit queue is used to represent a multi-source sensor data sequence that is time-synchronized and spatially aligned with the inertial measurement unit coordinate system. The feature extraction module is used to extract features and estimate the state of the sensing unit queue to obtain a multi-source sensing feature set; wherein, the multi-source sensing feature set includes panoramic images, dynamic radar point clouds and helmet attitude pre-integration sequences; The fusion perception module is used to input the panoramic image into a pre-constructed lightweight pure visual target detection model to obtain candidate region proposals and visual feature vectors corresponding to the candidate region proposals, project the dynamic radar point cloud onto the image plane of the panoramic image to obtain the target feature vectors of each radar point, and fuse the target feature vectors and visual feature vectors of each radar point to obtain a set of perceived targets. The tracking and prediction module is used to obtain target motion prediction information based on the set of perceived targets and the helmet posture pre-integration sequence; wherein, the target motion prediction information is used to characterize the predicted motion state and trajectory of each target relative to the smart helmet in a future set time period; The risk assessment module is used to calculate the helmet's own pose based on the helmet's pre-integrated sequence of postures, construct a dynamic risk field with the helmet's own pose as the origin, and obtain the risk potential energy of each target based on the target motion prediction information; wherein, the risk potential energy is used to characterize the degree of collision threat posed by each target to the smart helmet based on the predicted motion state and the trajectory. The graded early warning module is used to sort the risk potential energy from high to low, obtain the azimuth angle of the target with the highest risk potential energy, and generate a graded early warning command based on the azimuth angle and the target with the highest risk potential energy. The graded early warning command is used to instruct the voice broadcaster of the smart helmet to output early warning voice content corresponding to the azimuth angle, and instruct the vibration motor located on the side corresponding to the azimuth angle to vibrate according to the vibration frequency and duty cycle matched with the early warning level.
8. A computer device comprising a memory and a processor, wherein the memory stores a computer program, characterized in that, When the processor executes the computer program, it implements the method of any one of claims 1 to 6.
9. A computer-readable storage medium having a computer program stored thereon, characterized in that, When the computer program is executed by a processor, it implements the method of any one of claims 1 to 6.