An engineered vehicle multi-modal safety monitoring system and method
By using dynamic confidence-weighted fusion of multi-source heterogeneous sensor modules and intelligent vehicle-mounted processing terminals, the problem of perception system failure in engineering vehicles under harsh environments has been solved. This has enabled dynamic fusion of driver attention and vehicle status, improving the accuracy and reliability of early warnings.
Patent Information
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- SHANTUI CONSTR MASCH CO LTD
- Filing Date
- 2026-04-13
- Publication Date
- 2026-06-09
Smart Images

Figure CN122166122A_ABST
Abstract
Description
Technical Field
[0001] This application belongs to the field of engineering machinery safety monitoring technology, specifically relating to a multimodal safety monitoring system and method for engineering vehicles. Background Technology
[0002] Construction machinery (such as bulldozers, excavators, and mining trucks) often operate in complex environments with mixed human and machine activity, high dust levels, and high noise levels, such as mines, tunnels, and large construction sites. These environments commonly present problems such as large blind spots for drivers, variable vehicle postures (such as steep inclines and tilts), and harsh environments that can cause sensor failures, leading to frequent safety accidents such as collisions and rollovers.
[0003] Currently, existing technologies mostly employ independent or simply integrated safety systems. Examples include: 1) Reversing camera / radar systems, which only address blind spots at the rear when reversing, failing to detect side and frontal hazards or the driver's status. 2) Independent Driver Monitoring Systems (DMS), which only monitor driver fatigue and distraction, but are unaware of external dynamic hazards, and provide isolated warnings. 3) Simple 360° surround view systems, which only provide panoramic images, lacking accurate obstacle distance measurement and active warnings, and suffer from severe image quality degradation in adverse weather conditions. Existing technologies primarily suffer from the following drawbacks: First, the warning logic is passive and prone to false alarms. Warnings based on fixed distance thresholds fail to consider the driver's current attention distribution and the vehicle's future movement trends, leading to frequent interruptions from alarms. Second, it suffers from poor environmental adaptability and insufficient reliability. In typical engineering scenarios such as dust, rain, snow, and mud, pure vision systems fail, while single radar systems are susceptible to clutter interference, resulting in low confidence levels in perception results. Finally, the depth of information fusion is insufficient. It merely overlays and displays data from people, vehicles, and the environment or triggers rules, lacking in-depth modeling and prediction of the dynamic coupling relationship among these three elements, thus failing to achieve truly intelligent risk prediction.
[0004] Therefore, there is an urgent need for a safety monitoring system for engineering vehicles that can dynamically adapt to harsh environments and provide predictive early warnings by integrating driver attention and vehicle status. Summary of the Invention
[0005] In a first aspect, embodiments of this application provide a multimodal safety monitoring system for engineering vehicles, including a multi-source heterogeneous sensing module and an intelligent vehicle-mounted processing terminal; Multi-source heterogeneous sensor modules are installed on engineering vehicles to simultaneously collect environmental multimodal data, driver status data, and vehicle motion status data. The intelligent vehicle-mounted processing terminal is connected to a multi-source heterogeneous sensor module. The intelligent vehicle-mounted processing terminal includes an automotive-grade embedded processor, which is connected to an in-vehicle touch screen and an audible and visual alarm. The automotive-grade embedded processor is configured as follows: The system receives and preprocesses data collected by multi-source heterogeneous sensing modules, and calculates the dynamic confidence level of each sensor output in real time. Based on the dynamic confidence level, it performs weighted fusion of the perception results from different sensors to generate fused environmental obstacle information and driver attention information. Based on the vehicle's own motion state data, it predicts the dynamic dangerous motion trajectory sector of the vehicle within a preset time window in real time. It maps the driver attention information to the vehicle's surrounding environment to generate an attention heatmap, and couples the attention heatmap, the fused environmental obstacle information, and the dynamic dangerous motion trajectory sector to calculate and output the comprehensive risk coefficient of each obstacle. The in-vehicle touchscreen display is used to display high-risk targets and 360° panoramic images with differentiated visual identifiers based on a comprehensive risk factor. An audible and visual alarm is used to perform graded audible and visual alarms based on the comprehensive risk coefficient.
[0006] Furthermore, the multi-source heterogeneous sensor module includes a driver status monitoring camera, an inertial measurement unit, a vehicle attitude sensor, at least four surround-view cameras, and at least four millimeter-wave radars. The driver status monitoring camera is installed in the cab of the engineering vehicle and faces the driver's face to collect facial images, gaze direction and head posture data of the driver. The inertial measurement unit is installed on the vehicle body to acquire real-time data on the vehicle's three-axis acceleration, roll angle, and pitch angle. Vehicle attitude sensor, used to acquire vehicle steering angle and speed data; Each surround-view camera is installed at the front, rear, left, and right sides of the engineering vehicle to collect visual images of the environment around the vehicle. Each millimeter-wave radar is installed on the front bumper, rear bumper, and both sides of the engineering vehicle to detect the distance, relative speed, and azimuth of obstacles. The vehicle-mounted environmental perception sensor array is installed on the exterior of the vehicle or on the top of the cab to collect data on rain and snow intensity, light intensity, and PM2.5 concentration. The vehicle-mounted environmental perception sensor module includes a rain sensor, a light sensor, and an onboard air quality sensor.
[0007] Furthermore, the automotive-grade embedded processor is also connected to volatile memory, non-volatile memory, and an in-vehicle communication interface; Volatile memory is used to temporarily store data acquired by the sensor and intermediate calculation data; Non-volatile memory used to permanently store executable program code; The vehicle communication interface is used to receive data from multi-source heterogeneous sensor modules and vehicle status signals via CAN bus or Ethernet.
[0008] Secondly, embodiments of this application also provide a multimodal safety monitoring method for engineering vehicles, applied to the system described in the first aspect, comprising the following steps: S1. Simultaneously collect multimodal environmental data, driver status data, and vehicle motion status data around the engineering vehicle through a multi-source heterogeneous sensing module, and transmit the collected data as raw data to the intelligent vehicle processing terminal. S2. The intelligent vehicle-mounted processing terminal preprocesses the raw data from each sensor and calculates the dynamic confidence level of each sensor output in real time. S3. Based on dynamic confidence, the perception results of different sensors are weighted and fused to generate fused environmental obstacle information and driver attention information; S4. Based on the vehicle's own motion status data, predict the dynamic dangerous motion trajectory sector of the vehicle within a preset time window in real time; S5. Map the driver's attention information to the vehicle's surrounding environment to generate an attention heatmap. Couple the attention heatmap, the fused environmental obstacle information, and the dynamic dangerous movement trajectory sectors to calculate and output the comprehensive risk coefficient of each obstacle. S6. When the overall risk coefficient exceeds the preset threshold, a graded warning is issued, and high-risk targets are highlighted on the vehicle touch screen.
[0009] Furthermore, the specific steps of step S1 are as follows: S11. By installing surround-view cameras at the front, rear, left and right sides of the engineering vehicle, the system simultaneously collects visual images of the environment around the vehicle, generates multiple video streams, and transmits them to the intelligent vehicle processing terminal via the vehicle Ethernet. S12. By using a driver status monitoring camera installed in the driver's cab and facing the driver's face, the driver's facial image is simultaneously captured, driver video stream data is generated, and transmitted to the intelligent vehicle processing terminal via vehicle Ethernet. S13. By installing millimeter-wave radars on the front bumper, rear bumper and both sides of the vehicle, the original echo signals of obstacles are collected synchronously, radar point cloud data is generated, and transmitted to the intelligent vehicle processing terminal via CAN bus or vehicle Ethernet. S14. The inertial measurement unit installed on the vehicle body synchronously collects the vehicle's three-axis acceleration, roll angle and pitch angle data, and transmits them to the intelligent vehicle processing terminal via the CAN bus. S15. The vehicle's steering angle and speed data are collected synchronously through the vehicle attitude sensor and transmitted to the intelligent vehicle processing terminal via the CAN bus. S16. Through the vehicle-mounted environmental perception sensor group, the current environmental parameters such as rain and snow intensity, light intensity, and PM2.5 concentration are collected synchronously and transmitted to the intelligent vehicle processing terminal via CAN bus or vehicle Ethernet.
[0010] Furthermore, the specific steps of step S2 are as follows: S21. The intelligent vehicle processing terminal receives raw data from various sensors, decodes, corrects distortion and preprocesses noise in the video stream data of the surround view camera and the driver status monitoring camera, filters and clusters the point cloud data of the millimeter-wave radar, performs timestamp alignment and coordinate system normalization preprocessing on the data of the inertial measurement unit and the vehicle attitude sensor, and filters and verifies the validity of the data of the environmental perception sensor group. S22. Based on the current environmental parameters of light intensity, rain / snow intensity, and PM2.5 concentration, calculate the environmental confidence score of the i-th sensor at time t. ; S23. Obtain the signal strength, data loss rate, and self-test fault code for each sensor, and calculate the self-confidence score of the i-th sensor at time t. ; S24. Based on environmental confidence scores And the sensor's self-confidence score The dynamic confidence level of the i-th sensor at time t is calculated using the following formula. :
[0011] in, and For preset weighting coefficients and .
[0012] Furthermore, the specific steps of step S3 are as follows: S31. Calculate the obstacle detection results of each sensor, and perform spatiotemporal synchronization correlation of the obstacle detection results of different sensors. Based on the azimuth, distance and timestamp of the obstacle, determine whether it points to the same obstacle, and establish an obstacle cross-sensor correlation matrix. S32. When multiple sensors detect the same obstacle, obtain the current dynamic confidence level of each successfully associated sensor. and the distance measurement value output by the corresponding sensor. The distance to the merged obstacle is calculated using the following formula:
[0013] Where n is the total number of sensors that detected this obstacle; S33. Based on the image recognition results from the visual sensor with the highest confidence, determine the category of obstacles using a pre-trained deep learning model for object detection, and output a fused list of environmental obstacle information, with each obstacle containing a corresponding fused distance. Category and azimuth; S34. Using the driver's facial image captured by the driver status monitoring camera, extract the driver's gaze direction vector and head posture Euler angles using computer vision algorithms, and output the driver's attention information, including the azimuth and pitch angles of the gaze point in the vehicle coordinate system.
[0014] Furthermore, the specific steps of step S4 are as follows: S41. Obtain the current vehicle speed from the vehicle's CAN bus. and steering angle And obtain the vehicle's current roll angle from the inertial measurement unit. and pitch angle ; S42. Based on the Ackermann steering kinematics model, calculate the vehicle's position within a preset time window. Longitudinal driving distance within and maximum lateral offset ; in, This refers to the vehicle's wheelbase. S43. Determine a point on the ground projection with the vehicle's current position as the starting point and the longitudinal travel distance as the starting point. Radial length, with maximum lateral offset A sector-shaped region with a horizontal boundary; The sector is enclosed by the minimum turning radius boundary and the straight-line driving boundary, thus forming a dynamic dangerous movement trajectory sector. S44. Determine the roll angle Does it exceed the preset roll threshold? ; If so, extend the sector range in the lateral tilt direction by a preset rollover risk zone width. This area is designated as an extended zone and marked as a high-risk rollover zone. If not, proceed to step S5; Among them, the preset rollover risk zone width Calculated dynamically based on vehicle speed.
[0015] Furthermore, the specific steps of step S5 are as follows: S51. Based on the driver's attention information, fuse the driver's gaze direction vector with the head posture Euler angles to calculate the driver's actual gaze direction in the vehicle coordinate system: Let the head yaw angle be Pitch angle is The yaw angle of the line of sight relative to the head is Pitch angle The horizontal angle of the driver's actual gaze direction in the vehicle coordinate system. and vertical angle They are respectively:
[0016] ; S52. Determine the driver's actual gaze direction By combining the spatial geometric relationships in the vehicle coordinate system with the driver's head as the origin, the projected coordinates of the gaze point on the ground are calculated. : Let the driver's head be at the height of the ground. Then the distance between the gaze point and the vehicle and gaze coordinates They are respectively:
[0017]
[0018] ; S53. Using the coordinates of the gaze point Generate a two-dimensional Gaussian-distributed attention heatmap centered on the target: The ground area around the vehicle in a 360° radius is divided into... Discrete grid, each grid cell Attention coefficient The calculation formula is:
[0019] in, For grid cells The coordinates of the center point The radius of attention decay; S54. For each fused obstacle j generated in step S3, obtain the position coordinates of this obstacle in the vehicle coordinate system. Locate the corresponding grid cell and extract the attention coefficient for that grid cell. ; in, The value ranges from 0 to 1, with a larger value indicating a more focused driver attention on the area where the obstacle is located. S55. Based on the position coordinates of the j-th obstacle. Whether the vehicle falls within the dynamic dangerous motion trajectory sector generated in step S4, the vehicle's situational danger coefficient is obtained. : If it falls in, then ; If it does not fall in, then ; S56. Obtain the fusion distance of the obstacle generated in step S3. and read the preset safe distance threshold. ; Preset safe distance threshold According to vehicle speed Dynamic adjustment:
[0020] In the formula, To preset the braking response time, To pre-set a safety margin; S57. The comprehensive risk coefficient shall be calculated using the following formula. :
[0021] in, Let the overall risk coefficient be the j-th obstacle; For the distance term, an S-shaped function, when When the function value approaches 1, The function value approaches 0 as time progresses; , , These are the preset weight coefficients for the distance term, the attention deficit term, and the situation term, respectively, satisfying... ; is the steepness coefficient of the sigmoid function of the distance term.
[0022] Furthermore, step S6 is detailed as follows: S61. A first-level emergency warning threshold is preset in the intelligent vehicle-mounted processing terminal. and Level II emergency warning threshold ,and ; S62. Identify the overall risk factor for each obstacle. ; like This will trigger a Level 1 emergency alert: On the 360° panoramic image of the vehicle's touch screen, the obstacle target is highlighted with a first visual identifier, and the fused distance value and risk warning text are dynamically displayed. At the same time, the sound and light alarm emits a continuous sound and light alarm at the first frequency. like This will trigger a level 2 warning: On the 360° panoramic image on the in-vehicle touch screen, the obstacle target is marked with a second visual identifier and the fusion distance value is displayed. At the same time, the sound and light alarm emits an intermittent sound and light alarm at a second frequency. like If the obstacle is not triggered, the sound and light alarm will not be activated. Instead, the obstacle and its fusion distance will be displayed in the 360° panoramic image on the vehicle's touch screen in a normal color. A semi-transparent color block will be used to indicate the driver's blind spot in the low attention area of the attention heatmap.
[0023] As can be seen from the above technical solutions, this application has the following advantages: The multimodal safety monitoring system and method for engineering vehicles provided in this application achieves adaptability and accuracy of the perception system for engineering machinery vehicles in harsh environments such as rain, snow, and dust by using multi-source heterogeneous sensing and dynamic confidence weighted fusion, thus solving the problem of single sensor failure under extreme conditions. Through the coupled calculation of driver attention heatmaps and vehicle dynamic hazard trajectory sectors, accurate quantitative assessment of operational risks is achieved, solving the problems of false alarms and missed alarms in traditional systems. Through graded early warning and differentiated visual labels for panoramic images, precise situational perception assistance for drivers is achieved, avoiding invalid alarm interference and ensuring immediate response in high-risk scenarios, thereby improving the active safety performance of engineering vehicles. Attached Figure Description
[0024] To more clearly illustrate the technical solution of this application, the accompanying drawings used in the description will be briefly introduced below. Obviously, the accompanying drawings described below are only some embodiments of this application. For those skilled in the art, other drawings can be obtained based on these drawings without creative effort.
[0025] Figure 1 This is a schematic diagram of the multimodal safety monitoring system for engineering vehicles according to the present invention.
[0026] Figure 2 This is a flowchart illustrating the multimodal safety monitoring method for engineering vehicles according to the present invention. Detailed Implementation
[0027] The various embodiments of this disclosure will be described more fully in the following detailed description of the multimodal safety monitoring system for engineering vehicles. This disclosure may have various embodiments, and adjustments and changes may be made therein. However, it should be understood that there is no intention to limit the various embodiments of this disclosure to the specific embodiments disclosed herein, but rather this disclosure should be understood to cover all adjustments, equivalents, and / or alternatives falling within the spirit and scope of the various embodiments of this disclosure.
[0028] This embodiment provides a multimodal safety monitoring system for engineering vehicles, which achieves high robustness and accurate early warning in harsh environments through multimodal fusion and dynamic confidence weighting.
[0029] The technical solutions of the embodiments of the present invention will be clearly and completely described below with reference to the accompanying drawings. Obviously, the described embodiments are only some embodiments of the present invention, and not all embodiments. Based on the embodiments of the present invention, all other embodiments obtained by those skilled in the art without creative effort are within the scope of protection of the present invention.
[0030] Please see Figure 1 The diagram shown is a schematic of a multimodal safety monitoring system for engineering vehicles in a specific embodiment. The system includes a multi-source heterogeneous sensing module and an intelligent vehicle-mounted processing terminal. Multi-source heterogeneous sensor modules are installed on engineering vehicles to simultaneously collect environmental multimodal data, driver status data, and vehicle motion status data. It should be noted that the multi-source heterogeneous sensor module breaks through the limitations of a single mode, simultaneously acquiring environmental images, radar point clouds, vehicle motion parameters, and environmental meteorological parameters, providing a comprehensive data foundation for subsequent fusion analysis and ensuring no information blind spots. The intelligent vehicle-mounted processing terminal is connected to a multi-source heterogeneous sensor module. The intelligent vehicle-mounted processing terminal includes an automotive-grade embedded processor, which is connected to an in-vehicle touch screen and an audible and visual alarm. The automotive-grade embedded processor is configured as follows: The system receives and preprocesses data collected by multi-source heterogeneous sensing modules, and calculates the dynamic confidence level of each sensor output in real time. Based on the dynamic confidence level, it performs weighted fusion of the perception results from different sensors to generate fused environmental obstacle information and driver attention information. Based on the vehicle's own motion state data, it predicts the dynamic dangerous motion trajectory sector of the vehicle within a preset time window in real time. It maps the driver attention information to the vehicle's surrounding environment to generate an attention heatmap, and couples the attention heatmap, the fused environmental obstacle information, and the dynamic dangerous motion trajectory sector to calculate and output the comprehensive risk coefficient of each obstacle. It should be noted that the automotive-grade embedded processor performs complex confidence calculations, data fusion, and risk assessments locally on the vehicle, without relying on a remote server, reducing data transmission latency and meeting the safety requirements of millisecond-level response during vehicle operation. The in-vehicle touchscreen display is used to display high-risk targets and 360° panoramic images with differentiated visual identifiers based on a comprehensive risk factor. An audible and visual alarm device is used to perform graded audible and visual alarms based on the comprehensive risk coefficient. It should be noted that the in-vehicle touch screen and the sound and light alarm output different sound and light stimuli according to the risk level (such as low, medium and high), avoiding the continuous beeping that would cause the driver to become irritated and shut down the system. Strong interference is only provided in critical moments, thus improving the availability of the system.
[0031] This embodiment achieves robustness in perception of harsh environments through multi-source heterogeneous sensing and dynamic confidence weighting; furthermore, by coupling attention heatmaps with dynamic trajectory sectors, it enables precise graded early warnings only when the driver is not paying attention and the vehicle is in danger.
[0032] Furthermore, as a refinement and extension of the specific implementation of the above embodiments, in order to fully illustrate the specific implementation process in this embodiment, another engineering vehicle multimodal safety monitoring system is provided, which includes a multi-source heterogeneous sensing module and an intelligent vehicle-mounted processing terminal. Multi-source heterogeneous sensor modules are installed on engineering vehicles to simultaneously collect environmental multimodal data, driver status data, and vehicle motion status data. The intelligent vehicle-mounted processing terminal is connected to a multi-source heterogeneous sensor module. The intelligent vehicle-mounted processing terminal includes an automotive-grade embedded processor, which is connected to an in-vehicle touch screen and an audible and visual alarm. The automotive-grade embedded processor is configured as follows: The system receives and preprocesses data collected by multi-source heterogeneous sensing modules, and calculates the dynamic confidence level of each sensor output in real time. Based on the dynamic confidence level, it performs weighted fusion of the perception results from different sensors to generate fused environmental obstacle information and driver attention information. Based on the vehicle's own motion state data, it predicts the dynamic dangerous motion trajectory sector of the vehicle within a preset time window in real time. It maps the driver attention information to the vehicle's surrounding environment to generate an attention heatmap, and couples the attention heatmap, the fused environmental obstacle information, and the dynamic dangerous motion trajectory sector to calculate and output the comprehensive risk coefficient of each obstacle. The in-vehicle touchscreen display is used to display high-risk targets and 360° panoramic images with differentiated visual identifiers based on a comprehensive risk factor. An audible and visual alarm device is used to perform graded audible and visual alarms based on the comprehensive risk coefficient. The multi-source heterogeneous sensor module includes a driver status monitoring camera, an inertial measurement unit, a vehicle attitude sensor, at least four surround-view cameras, and at least four millimeter-wave radars; The driver status monitoring camera is installed in the cab of the engineering vehicle and faces the driver's face to collect facial images, gaze direction and head posture data of the driver. The inertial measurement unit is installed on the vehicle body to acquire real-time data on the vehicle's three-axis acceleration, roll angle, and pitch angle. Vehicle attitude sensor, used to acquire vehicle steering angle and speed data; Each surround-view camera is installed at the front, rear, left, and right sides of the engineering vehicle to collect visual images of the environment around the vehicle. Each millimeter-wave radar is installed on the front bumper, rear bumper, and both sides of the engineering vehicle to detect the distance, relative speed, and azimuth of obstacles. The vehicle-mounted environmental perception sensor array is installed on the exterior of the vehicle or on the top of the cab to collect data on rain and snow intensity, light intensity, and PM2.5 concentration. The vehicle-mounted environmental perception sensor module includes a rain sensor, a light sensor, and an on-board air quality sensor. The automotive-grade embedded processor is also connected to volatile memory, non-volatile memory, and an in-vehicle communication interface. Volatile memory is used to temporarily store data acquired by the sensor and intermediate calculation data; Non-volatile memory used to permanently store executable program code; The vehicle communication interface is used to receive data from multi-source heterogeneous sensor modules and vehicle status signals via CAN bus or Ethernet.
[0033] like Figure 2 As shown, the following are embodiments of the multimodal safety monitoring method for engineering vehicles provided in this disclosure. This method and the multimodal safety monitoring system for engineering vehicles in the above embodiments belong to the same inventive concept. For details not described in detail in the embodiments of the multimodal safety monitoring method for engineering vehicles, please refer to the embodiments of the multimodal safety monitoring system for engineering vehicles described above.
[0034] The method includes the following steps: S1. Simultaneously collect multimodal environmental data, driver status data, and vehicle motion status data around the engineering vehicle through a multi-source heterogeneous sensing module, and transmit the collected data as raw data to the intelligent vehicle processing terminal. It should be noted that this step realizes a full mapping from the physical world to the digital world, especially the synchronous collection of environmental meteorological parameters, which provides prior knowledge for subsequent judgment on whether the sensor is affected by environmental interference, and achieves highly reliable sensing. S2. The intelligent vehicle-mounted processing terminal preprocesses the raw data from each sensor and calculates the dynamic confidence level of each sensor output in real time. It should be noted that this step, by quantifying the impact of the environment on the sensor, can eliminate interference from invalid or low-quality data, ensuring that the data input to the fusion layer is high-quality information that has been filtered, thus achieving data cleaning and weighting. S3. Based on dynamic confidence, the perception results of different sensors are weighted and fused to generate fused environmental obstacle information and driver attention information; It should be noted that this step involves two aspects: firstly, using high-confidence data to correct low-confidence data and outputting accurate obstacle locations and categories; secondly, extracting the driver's line of sight through algorithms to digitize human factors, enabling the system not only to identify obstacles but also to determine whether the driver has identified them, thus achieving a combination of virtual and real-world data.
[0035] S4. Based on the vehicle's own motion status data, predict the dynamic dangerous motion trajectory sector of the vehicle within a preset time window in real time; It should be noted that, unlike traditional panoramic images which simply display images passively, this step uses a kinematic model to calculate the vehicle's predicted trajectory. Combined with slope tilt extension logic, it can identify rollover risk areas that cannot be displayed in static panoramic images in advance, thus achieving spatiotemporal prediction. S5. Map the driver's attention information to the vehicle's surrounding environment to generate an attention heatmap. Couple the attention heatmap, the fused environmental obstacle information, and the dynamic dangerous movement trajectory sectors to calculate and output the comprehensive risk coefficient of each obstacle. It should be noted that this step is based on the coupled calculation of driver attention distribution and vehicle dynamic dangerous trajectory, and comprehensively evaluates obstacle position, driver attention deficit and vehicle motion state, outputting quantitative risk coefficient, thus realizing proactive prediction and accurate identification of potential collision risks. S6. When the overall risk coefficient exceeds the preset threshold, a graded warning is issued, and high-risk targets are highlighted on the vehicle touch screen. It should be noted that this step transforms complex calculation results into visual symbols and sound frequencies that are easy for drivers to understand. Through a differentiated feedback mechanism, it ensures that drivers receive warnings at the right time and in the right way, without overlooking risks or being disturbed by noise, thus achieving precise intervention.
[0036] This embodiment achieves highly reliable perception under rain, snow, and dust by calculating the dynamic confidence of sensors in real time and weighted fusion; by mapping the driver's attention to a heat map and coupling it with the vehicle's dynamic dangerous trajectory sector for calculation, it achieves graded audio-visual warnings only when attention is lost and the trajectory conflicts occur, solving the problems of high false alarm rate and poor environmental adaptability of traditional systems.
[0037] Furthermore, as a refinement and extension of the specific implementation methods of the above embodiments, in order to fully illustrate the specific implementation process of this embodiment, another multimodal safety monitoring method for engineering vehicles is provided, taking a bulldozer operating in a mine as an example. A 2-megapixel ultra-wide-angle fisheye camera is installed at the front, rear, and left and right sides behind the cab of the bulldozer, forming a surround-view camera group to collect 360° environmental images around the vehicle. An infrared driver status monitoring camera is installed directly in front of the driver in the cab to ensure clear capture of the driver's face. Two millimeter-wave radars are installed on the front bumper, two on the rear bumper, and one on each side, for a total of six millimeter-wave radars, achieving coverage of approximately 5 meters around the vehicle. A rain sensor, a light sensor, and an onboard air quality sensor are installed on the roof to form an environmental perception sensor group. An inertial measurement unit is installed on the body and connected to the vehicle's CAN bus to obtain data on vehicle speed, steering angle, roll angle, and pitch angle. The method includes the following steps: S1. Simultaneously collect multimodal environmental data, driver status data, and vehicle motion status data around the engineering vehicle through a multi-source heterogeneous sensing module, and transmit the collected data as raw data to the intelligent vehicle processing terminal. The specific steps of step S1 are as follows: S11. By installing surround-view cameras at the front, rear, left and right sides of the engineering vehicle, the system simultaneously collects visual images of the environment around the vehicle, generates multiple video streams, and transmits them to the intelligent vehicle processing terminal via the vehicle Ethernet. S12. By using a driver status monitoring camera installed in the driver's cab and facing the driver's face, the driver's facial image is simultaneously captured, driver video stream data is generated, and transmitted to the intelligent vehicle processing terminal via vehicle Ethernet. S13. By installing millimeter-wave radars on the front bumper, rear bumper and both sides of the vehicle, the original echo signals of obstacles are collected synchronously, radar point cloud data is generated, and transmitted to the intelligent vehicle processing terminal via CAN bus or vehicle Ethernet. S14. The inertial measurement unit installed on the vehicle body synchronously collects the vehicle's three-axis acceleration, roll angle and pitch angle data, and transmits them to the intelligent vehicle processing terminal via the CAN bus. S15. The vehicle's steering angle and speed data are collected synchronously through the vehicle attitude sensor and transmitted to the intelligent vehicle processing terminal via the CAN bus. S16. Through the vehicle-mounted environmental perception sensor group, the current environmental parameters such as rain and snow intensity, light intensity, and PM2.5 concentration are collected synchronously and transmitted to the intelligent vehicle processing terminal via CAN bus or vehicle Ethernet. Among them, illuminance is collected by a light sensor or estimated based on the automatic exposure parameters of the surround view camera, rain and snow intensity is collected by a rain sensor, and PM2.5 concentration is collected by an on-board air quality sensor. For example, suppose the bulldozer is reversing in rainy conditions at 20:00 at night, with an ambient light intensity of 50 lux, a rain / snow intensity of 8 mm / h, and a PM2.5 concentration of 120 μg / m³. Four surround-view cameras simultaneously capture visual images of the environment around the vehicle, generating four video streams at a rate of 30 frames per second, which are then transmitted to the intelligent vehicle processing terminal via the vehicle Ethernet. The driver status monitoring camera simultaneously captures the driver's facial images and generates driver video stream data at a rate of 30 frames per second, which is then transmitted to the intelligent vehicle processing terminal via the vehicle Ethernet. Six millimeter-wave radars simultaneously acquire the raw echo signals of obstacles, generate radar point cloud data at a rate of 20Hz, and transmit it to the intelligent vehicle processing terminal via CAN bus. The inertial measurement unit synchronously collects the vehicle's three-axis acceleration, roll angle, and pitch angle data, and transmits them to the intelligent vehicle processing terminal via the CAN bus at a rate of 100Hz. At this time, the vehicle is reversing at a speed of 5 km / h, with a steering angle of 0° (i.e., reversing in a straight line) and a roll angle of 3° (i.e., no obvious body roll). The vehicle attitude sensor synchronously collects the vehicle's steering angle and speed data, and transmits them to the intelligent vehicle processing terminal via the CAN bus. The rain sensor collects the rain and snow intensity at 8 mm / h, the light sensor collects the illuminance at 50 lux, and the vehicle air quality sensor collects the PM2.5 concentration at 120 μg / m³. These data are transmitted to the intelligent vehicle processing terminal via the CAN bus. S2. The intelligent vehicle-mounted processing terminal preprocesses the raw data from each sensor and calculates the dynamic confidence level of each sensor output in real time. The specific steps of step S2 are as follows: S21. The intelligent vehicle processing terminal receives raw data from various sensors, decodes, corrects distortion and preprocesses noise in the video stream data of the surround view camera and the driver status monitoring camera, filters and clusters the point cloud data of the millimeter-wave radar, performs timestamp alignment and coordinate system normalization preprocessing on the data of the inertial measurement unit and the vehicle attitude sensor, and filters and verifies the validity of the data of the environmental perception sensor group. For example, the intelligent vehicle processing terminal receives all sensor data and performs preprocessing: The four surround-view video streams are decoded, distortion correction is performed based on the intrinsic and extrinsic parameters of the fisheye camera, and Gaussian filtering is used to remove image noise. Decode and denoise the driver's video stream; Point cloud data from six millimeter-wave radars were filtered to remove outliers, and the DBSCAN clustering algorithm was used to aggregate point clouds that are spatially close into candidate targets. The data from the inertial measurement unit and the vehicle attitude sensor are aligned according to the timestamp, and the data from each sensor are uniformly converted to the vehicle coordinate system with the rear axle center as the origin. Median filtering is applied to the data from the environmental perception sensor array to remove outliers; S22. Based on the current environmental parameters of light intensity, rain / snow intensity, and PM2.5 concentration, calculate the environmental confidence score of the i-th sensor at time t. ; For vision-based sensors (surround-view cameras, driver status monitoring cameras), the following formula is used for calculation:
[0038] in, , To preset a reference illuminance (e.g., a value of 100 lux), when the illuminance is lower than... The time fraction decreases linearly; , A preset rain / snow intensity threshold (e.g., 10 mm / h) is set when the rain / snow intensity exceeds... The hour and minute fractions are set to zero; , A preset PM2.5 concentration threshold (e.g., 150 μg / m³) is set. When the PM2.5 concentration exceeds... The hour and minute fractions are set to zero; , , For the preset weighting coefficients, satisfy ; For radar sensors (millimeter-wave radar), the following formula is used for calculation:
[0039] in, , A preset rain / snow intensity threshold (e.g., 10 mm / h) is set when the rain / snow intensity exceeds... The hour and minute fractions are set to zero; , A preset PM2.5 concentration threshold (e.g., 150 μg / m³) is set. When the PM2.5 concentration exceeds... The hour and minute fractions are set to zero; Furthermore, the environmental confidence score of radar sensors is not affected by illumination. For example, the environmental confidence score for each sensor is calculated: For surround-view cameras and driver status monitoring cameras (visual sensors), based on illuminance... =50 lux, rain and snow intensity =8 mm / h, PM2.5 concentration =120 μg / m³, and preset parameters =100 lux、 =10 mm / h =150 μg / m³, weighting coefficient =0.4、 =0.3、 =0.3, calculate:
[0040]
[0041]
[0042]
[0043] For millimeter-wave radar (radar-type sensors), the environmental confidence score is only affected by rain / snow intensity and PM2.5 concentration, with weighting coefficients... =0.5、 =0.5, calculate:
[0044] S23. Obtain the signal strength, data loss rate, and self-test fault code for each sensor, and calculate the self-confidence score of the i-th sensor at time t. ;
[0045] in, It is the signal strength of the i-th sensor; It is the data packet loss rate of the i-th sensor, with a value ranging from 0 to 1; These are self-test fault codes; 0 indicates normal operation, and non-zero indicates a fault. , A preset reference signal strength (e.g., -60 dBm) is used when the signal strength is lower than... The time fraction decreases linearly; This represents the data packet reception success rate; the higher the packet loss rate, the lower this value. This value is set to zero when the self-test fault code is in an abnormal state. , , For the preset weighting coefficients, satisfy ; For example, calculate the self-confidence score for each sensor: Assuming a surround-view camera has good signal strength, the data packet loss rate is... =0.05, self-test fault code =0; Preset fractional function value corresponding to the reference signal strength =0.9, weighting coefficient =0.4、 =0.3、 =0.3, calculate: =0.9
[0046] =1
[0047] S24. Based on environmental confidence scores And the sensor's self-confidence score The dynamic confidence level of the i-th sensor at time t is calculated using the following formula. :
[0048] in, and For preset weighting coefficients and ; For example, calculate the dynamic confidence level: set up =0.6, =0.4, then the dynamic confidence level of the camera is: =0.6×0.32+0.4×0.945=0.57 Similarly, to calculate the dynamic confidence level of a millimeter-wave radar, assuming its own confidence score is 0.9, then: =0.6×0.2+0.4×0.9=0.48 S3. Based on dynamic confidence, the perception results of different sensors are weighted and fused to generate fused environmental obstacle information and driver attention information; The specific steps of step S3 are as follows: S31. Calculate the obstacle detection results of each sensor, and perform spatiotemporal synchronization correlation of the obstacle detection results of different sensors. Based on the azimuth, distance and timestamp of the obstacle, determine whether it points to the same obstacle, and establish an obstacle cross-sensor correlation matrix. The specific steps for calculating the obstacle detection results of each sensor are as follows: S311. Perform visual target detection on the preprocessed surround-view camera video stream data. Input the multi-channel surround-view camera images into a pre-trained target detection deep learning model (such as YOLOv8 or SSD), detect obstacles in each channel image, output the category of each obstacle, the pixel coordinates in the image and the detection confidence, and map the pixel coordinates to the azimuth and distance in the vehicle coordinate system through inverse perspective transformation to generate a list of surround-view camera obstacle detection results. S312. Perform target detection on the preprocessed millimeter-wave radar point cloud data, perform clustering processing on the radar point cloud (such as DBSCAN clustering algorithm), cluster point clouds belonging to the same obstacle into one target, calculate the distance, relative velocity and azimuth of the target, and generate a list of millimeter-wave radar obstacle detection results. S313. Calculate the dynamic confidence scores of each sensor obtained in step S2. The obstacle detection results from the surround-view camera and the millimeter-wave radar are assigned to the corresponding obstacle detection results, respectively, to generate independent obstacle detection results for each sensor with confidence weights; For example, each sensor detects obstacles independently: The surround-view camera detected a pedestrian about 2.5 meters behind the vehicle using the YOLOv8 object detection model, with a detection confidence score of 0.85. After inverse perspective transformation, the pedestrian's position coordinates in the vehicle coordinate system were obtained as (0, -2.5) (X is horizontal, Y is vertical, and the positive Y direction is directly in front). The millimeter-wave radar detected a moving target approximately 2.6 meters behind the vehicle, with a relative speed of -0.5 m / s (approaching the vehicle) and an azimuth angle of 0° (directly behind). S32. When multiple sensors detect the same obstacle, obtain the current dynamic confidence level of each successfully associated sensor. and the distance measurement value output by the corresponding sensor. The distance to the merged obstacle is calculated using the following formula:
[0049] Where n is the total number of sensors that detected this obstacle; For example, the two detection results above are spatiotemporally synchronized and correlated: both have an azimuth angle of 0°, a distance of 2.5 meters and 2.6 meters respectively, and a timestamp difference of less than 50ms, and are judged to be the same obstacle (i.e., pedestrian). S33. Based on the image recognition results from the visual sensor with the highest confidence, determine the category of obstacles (such as pedestrians, vehicles, equipment, cones, etc.) using a pre-trained deep learning model for object detection (such as YOLO or SSD), and output a fused list of environmental obstacles, with each obstacle containing a corresponding fused distance. Category and azimuth; For example, weighted fusion distance calculation: n=2, =0.57, =2.5, =0.48, =2.6
[0050] S34. Using computer vision algorithms, extract the driver's gaze direction vector and head posture Euler angles (including pitch, yaw, and roll angles) from the driver's facial image captured by the driver status monitoring camera, and output the driver's attention information, including the azimuth and pitch angles of the gaze point in the vehicle coordinate system. For example, the obstacle category is determined. The camera confidence score is 0.57, which is higher than that of the radar. The camera identification result is a pedestrian. Therefore, the obstacle category is determined to be a pedestrian. Driver attention information extraction: The driver's facial image is captured by the driver status monitoring camera. The computer vision algorithm extracts that the driver's gaze direction is shifted to the right and the head posture is turned to the right by about 30°. The horizontal angle of the gaze point in the vehicle coordinate system is calculated to be 30° and the vertical angle is -15° (i.e. looking down). The geometric calculation shows that the projection coordinates of the gaze point on the ground are approximately (5,0) (i.e., about 5 meters to the right of the vehicle). S4. Based on the vehicle's own motion status data, predict the dynamic dangerous motion trajectory sector of the vehicle within a preset time window in real time; The specific steps of step S4 are as follows: S41. Obtain the current vehicle speed from the vehicle's CAN bus. and steering angle And obtain the vehicle's current roll angle from the inertial measurement unit. and pitch angle ; For example, the vehicle's current speed is obtained from the CAN bus. =5 km / h≈1.39 m / s, steering angle =0°, the roll angle φ=3° and the pitch angle θ=0° are obtained from the inertial measurement unit; S42. Based on the Ackermann steering kinematics model, calculate the vehicle's position within a preset time window. Longitudinal driving distance within and maximum lateral offset ; in, This refers to the vehicle's wheelbase. The value range is 1-3 seconds; For example, a preset time window =2 seconds, vehicle wheelbase =3.5 meters, calculate: Longitudinal driving distance =1.39 × 2 = 2.78 meters; Maximum lateral offset ; S43. Determine a point on the ground projection with the vehicle's current position as the starting point and the longitudinal travel distance as the starting point. Radial length, with maximum lateral offset A sector-shaped region with a horizontal boundary; The sector is enclosed by the minimum turning radius boundary and the straight-line driving boundary, thus forming a dynamic dangerous movement trajectory sector. For example, since the steering angle is 0°, the dynamic dangerous movement trajectory sector degenerates into a rectangular area centered on the rear of the vehicle, that is, the area covered by the vehicle's straight reversing trajectory, with a width equal to the width of the vehicle body (approximately 3 meters) and a length of 2.78 meters; S44. Determine the roll angle Does it exceed the preset roll threshold? ; If so, extend the sector range in the lateral tilt direction by a preset rollover risk zone width. This area is designated as an extended zone and marked as a high-risk rollover zone. If not, proceed to step S5; in, The value range is 10°-15°, and the preset rollover risk zone width is... Calculated dynamically based on vehicle speed; For example, determine whether the roll angle φ=3° exceeds a preset roll threshold. =10°? The risk of rollover has not been exceeded, therefore the risk zone will not be expanded. S5. Map the driver's attention information to the vehicle's surrounding environment to generate an attention heatmap. Couple the attention heatmap, the fused environmental obstacle information, and the dynamic dangerous movement trajectory sectors to calculate and output the comprehensive risk coefficient of each obstacle. The specific steps of step S5 are as follows: S51. Based on the driver's attention information, fuse the driver's gaze direction vector with the head posture Euler angles to calculate the driver's actual gaze direction in the vehicle coordinate system: Let the head yaw angle be Pitch angle is The yaw angle of the line of sight relative to the head is Pitch angle The horizontal angle of the driver's actual gaze direction in the vehicle coordinate system. and vertical angle They are respectively:
[0051] ; For example, merging line of sight and head posture: head yaw angle =30°, the yaw angle of the line of sight relative to the head. =5°, then the actual horizontal gaze angle =35°; Head tilt angle =-10°, the angle of view relative to the head's tilt. =-5°, then the actual vertical angle of gaze =-15° S52. Determine the driver's actual gaze direction By combining the spatial geometric relationships in the vehicle coordinate system with the driver's head as the origin, the projected coordinates of the gaze point on the ground are calculated. : Let the driver's head be at the height of the ground. Then the distance between the gaze point and the vehicle and gaze coordinates They are respectively:
[0052]
[0053] ; For example, to calculate the ground projection coordinates of the gaze point: if the driver's head is 1.5 meters above the ground (H_head = 1.5 meters), then: =1.5×tan(90°-(-15°))≈-5.6 meters (the negative sign indicates the rear) =5.6×sin(35°)≈3.21 meters =5.6×cos(35°)≈4.59 meters S53. Using the coordinates of the gaze point Generate a two-dimensional Gaussian-distributed attention heatmap centered on the target: The ground area around the vehicle in a 360° radius is divided into... Discrete grid, each grid cell Attention coefficient The calculation formula is:
[0054] in, For grid cells The coordinates of the center point The attention decay radius (e.g., ranging from 0.5 to 2.0 meters, representing the size of the area of significant attention around the fixation point) is a Gaussian distribution function that ensures that the attention coefficient of the grid cell closer to the fixation point is higher, while the attention decays exponentially with increasing distance. For example, an attention heatmap is generated: Centered on (3.21, 4.59), =1.0 meter, divide the ground around the vehicle into a grid of 0.1 meter × 0.1 meter, and calculate the attention coefficient for each grid; S54. For each fused obstacle j generated in step S3, obtain the position coordinates of this obstacle in the vehicle coordinate system. Locate the corresponding grid cell and extract the attention coefficient for that grid cell. ; in, The value ranges from 0 to 1, with a larger value indicating a more focused driver attention on the area where the obstacle is located. For example, the attention coefficient of the obstacle is extracted: Given an obstacle with coordinates (0, -2.5), calculate its distance from the gaze point: =0 - 3.21 = -3.21 meters =-2.5-4.59=-7.09 meters Distance = ≈7.78 meters Attention coefficient ≈0 (very small, indicating that the driver's attention is completely off the obstacle behind them); S55. Based on the position coordinates of the j-th obstacle. Whether the vehicle falls within the dynamic dangerous motion trajectory sector generated in step S4, the vehicle's situational danger coefficient is obtained. : If it falls in, then ; If it does not fall in, then ; For example, to determine whether an obstacle falls into a dangerous trajectory sector, if the obstacle is located 2.5 meters directly behind the vehicle, and the vehicle's reversing trajectory sector covers a range of 0-2.78 meters directly behind the vehicle, then the obstacle falls into the sector. ; S56. Obtain the fusion distance of the obstacle generated in step S3. and read the preset safe distance threshold. ; Preset safe distance threshold According to vehicle speed Dynamic adjustment:
[0055] In the formula, To preset the braking response time, To pre-set a safety margin; For example, obtaining the fusion distance =2.55 meters, calculate the dynamic safety distance threshold: preset braking reaction time =1.0 second, safety margin =1.0 meter, vehicle speed =1.39 m / s, then: =1.39 × 1.0 + 1.0 = 2.39 meters S57. The comprehensive risk coefficient shall be calculated using the following formula. :
[0056] in, Let the overall risk coefficient be the j-th obstacle; For the distance term, an S-shaped function, when When the function value approaches 1, The function value approaches 0 as time progresses; , , These are the preset weight coefficients for the distance term, the attention deficit term, and the situation term, respectively, satisfying... ; is the steepness coefficient of the S-shaped function of the distance term, with a value ranging from 0.5 to 2.0; For example, calculate the comprehensive risk coefficient: Preset weighting coefficients =0.5, =0.3, =0.2, steepness coefficient =1.0, =2.55 meters, =2.39 meters, then - = -0.16 meters, distance term S-shaped function:
[0057] Attention deficit: ≈1-0=1.0 Situation item: =1
[0058] S6. When the overall risk coefficient exceeds the preset threshold, a graded warning is issued, and high-risk targets are highlighted on the vehicle touch screen. The specific steps of step S6 are as follows: S61. A first-level emergency warning threshold is preset in the intelligent vehicle-mounted processing terminal. and Level II emergency warning threshold ,and ; in, The value range is 0.7-0.9. The value range is 0.3-0.5; For example, a preset first-level emergency warning threshold is set. =0.7, Level II emergency warning threshold =0.4; S62. Identify the overall risk factor for each obstacle. ; like This will trigger a Level 1 emergency alert: On the 360° panoramic image of the in-vehicle touch screen, the obstacle target is highlighted with a first visual identifier (such as a red flashing border and an enlarged obstacle icon), and the fused distance value and risk warning text are dynamically displayed. At the same time, the sound and light alarm emits a continuous sound and light alarm at the first frequency (such as a continuous rapid beeping sound and a red LED flashing rapidly at a frequency of 5Hz). like This will trigger a level 2 warning: On the 360° panoramic image on the in-vehicle touch screen, the obstacle target is marked with a second visual identifier (such as a yellow border) and the fusion distance value is displayed. At the same time, the sound and light alarm emits an intermittent sound and light alarm at a second frequency (such as an intermittent beeping sound with a 1-second interval and a yellow LED flashing slowly at a frequency of 1Hz). like If the obstacle is not triggered, the sound and light alarm will not be triggered. Instead, the obstacle target and the fusion distance value will be displayed in the normal color on the 360° panoramic image of the vehicle touch screen, and a semi-transparent color block will be used to indicate the driver's blind spot in the low attention area of the attention heatmap. For example, identify =0.73, because > This triggers a Level 1 emergency alert. On the 360° panoramic image displayed on the in-vehicle touch screen, a flashing red pedestrian icon is shown 2.55 meters behind the vehicle, with a flashing red border, and dynamically displays the distance value "2.55m" and the risk warning text "Pedestrian behind!". At the same time, the audible and visual alarm emits a continuous, rapid beeping sound (frequency approximately 5Hz) and the red LED flashes rapidly. Although the driver was looking to the right at that moment, the system effectively drew the driver's attention through both sound and visual alerts, thus preventing a collision. If the overall risk coefficient of the obstacle is 0.5 (i.e., between 0.4 and 0.7), a level 2 warning will be triggered. Pedestrians are marked with a yellow border on the display screen, and the distance value is displayed. The sound and light alarm emits an intermittent buzzing sound (buzzing interval of 1 second) and the yellow LED flashes slowly. If the overall risk coefficient is below 0.4, the audible and visual alarms will not be triggered. Instead, the obstacle and its distance will be displayed on the screen in a normal color, and a semi-transparent color block will be used to indicate to the driver that the low attention area (i.e. the rear area) on the attention heatmap is a blind spot.
[0059] It should be understood that the sequence number of each step in the above embodiments does not imply the order of execution. The execution order of each process should be determined by its function and internal logic, and should not constitute any limitation on the implementation process of the embodiments of the present invention.
[0060] The above description of the disclosed embodiments enables those skilled in the art to make or use the invention. Various modifications to these embodiments will be readily apparent to those skilled in the art, and the general principles defined herein may be implemented in other embodiments without departing from the spirit or scope of the invention. Therefore, the invention is not to be limited to the embodiments shown herein, but is to be accorded the widest scope consistent with the principles and novel features disclosed herein.
Claims
1. A multimodal safety monitoring system for engineering vehicles, characterized in that, Including multi-source heterogeneous sensing modules and intelligent vehicle-mounted processing terminals; Multi-source heterogeneous sensor modules are installed on engineering vehicles to simultaneously collect environmental multimodal data, driver status data, and vehicle motion status data. The intelligent vehicle-mounted processing terminal is connected to a multi-source heterogeneous sensor module. The intelligent vehicle-mounted processing terminal includes an automotive-grade embedded processor, which is connected to an in-vehicle touch screen and an audible and visual alarm. The automotive-grade embedded processor is configured as follows: Receive and preprocess data collected by multi-source heterogeneous sensing modules, and calculate the dynamic confidence level of each sensor output in real time; Based on dynamic confidence, the perception results from different sensors are weighted and fused to generate fused environmental obstacle information and driver attention information; Based on the vehicle's own motion state data, the dynamic dangerous motion trajectory sector within the vehicle's future preset time window is predicted in real time; the driver's attention information is mapped to the vehicle's surrounding environment to generate an attention heatmap, and the attention heatmap, the fused environmental obstacle information, and the dynamic dangerous motion trajectory sector are coupled and calculated to output the comprehensive risk coefficient of each obstacle. The in-vehicle touchscreen display is used to display high-risk targets and 360° panoramic images with differentiated visual identifiers based on a comprehensive risk factor. An audible and visual alarm is used to perform graded audible and visual alarms based on the comprehensive risk coefficient.
2. The multimodal safety monitoring system for engineering vehicles according to claim 1, characterized in that, The multi-source heterogeneous sensor module includes a driver status monitoring camera, an inertial measurement unit, a vehicle attitude sensor, at least four surround-view cameras, and at least four millimeter-wave radars; The driver status monitoring camera is installed in the cab of the engineering vehicle and faces the driver's face to collect facial images, gaze direction and head posture data of the driver. The inertial measurement unit is installed on the vehicle body to acquire real-time data on the vehicle's three-axis acceleration, roll angle, and pitch angle. Vehicle attitude sensor, used to acquire vehicle steering angle and speed data; Each surround-view camera is installed at the front, rear, left, and right sides of the engineering vehicle to collect visual images of the environment around the vehicle. Each millimeter-wave radar is installed on the front bumper, rear bumper, and both sides of the engineering vehicle to detect the distance, relative speed, and azimuth of obstacles. The vehicle-mounted environmental perception sensor array is installed on the exterior of the vehicle or on the top of the cab to collect data on rain and snow intensity, light intensity, and PM2.5 concentration. The vehicle-mounted environmental perception sensor module includes a rain sensor, a light sensor, and an onboard air quality sensor.
3. The multimodal safety monitoring system for engineering vehicles according to claim 1, characterized in that, The automotive-grade embedded processor is also connected to volatile memory, non-volatile memory, and an in-vehicle communication interface. Volatile memory is used to temporarily store data acquired by the sensor and intermediate calculation data; Non-volatile memory used to permanently store executable program code; The vehicle communication interface is used to receive data from multi-source heterogeneous sensor modules and vehicle status signals via CAN bus or Ethernet.
4. A multimodal safety monitoring method for engineering vehicles, applied to the system described in any one of claims 1 to 3, characterized in that, Includes the following steps: S1. Simultaneously collect multimodal environmental data, driver status data, and vehicle motion status data around the engineering vehicle through a multi-source heterogeneous sensing module, and transmit the collected data as raw data to the intelligent vehicle processing terminal. S2. The intelligent vehicle-mounted processing terminal preprocesses the raw data from each sensor and calculates the dynamic confidence level of each sensor output in real time. S3. Based on dynamic confidence, the perception results of different sensors are weighted and fused to generate fused environmental obstacle information and driver attention information; S4. Based on the vehicle's own motion status data, predict the dynamic dangerous motion trajectory sector of the vehicle within a preset time window in real time; S5. Map the driver's attention information to the vehicle's surrounding environment to generate an attention heatmap. Couple the attention heatmap, the fused environmental obstacle information, and the dynamic dangerous movement trajectory sectors to calculate and output the comprehensive risk coefficient of each obstacle. S6. When the overall risk coefficient exceeds the preset threshold, a graded warning is issued, and high-risk targets are highlighted on the vehicle touch screen.
5. The multimodal safety monitoring method for engineering vehicles according to claim 4, characterized in that, The specific steps of step S1 are as follows: S11. By installing surround-view cameras at the front, rear, left and right sides of the engineering vehicle, the system simultaneously collects visual images of the environment around the vehicle, generates multiple video streams, and transmits them to the intelligent vehicle processing terminal via the vehicle Ethernet. S12. By using a driver status monitoring camera installed in the driver's cab and facing the driver's face, the driver's facial image is simultaneously captured, driver video stream data is generated, and transmitted to the intelligent vehicle processing terminal via vehicle Ethernet. S13. By installing millimeter-wave radars on the front bumper, rear bumper and both sides of the vehicle, the original echo signals of obstacles are collected synchronously, radar point cloud data is generated, and transmitted to the intelligent vehicle processing terminal via CAN bus or vehicle Ethernet. S14. The inertial measurement unit installed on the vehicle body synchronously collects the vehicle's three-axis acceleration, roll angle and pitch angle data, and transmits them to the intelligent vehicle processing terminal via the CAN bus; S15. The vehicle's steering angle and speed data are collected synchronously through the vehicle attitude sensor and transmitted to the intelligent vehicle processing terminal via the CAN bus. S16. Through the vehicle-mounted environmental perception sensor group, the current environmental parameters such as rain and snow intensity, light intensity, and PM2.5 concentration are collected synchronously and transmitted to the intelligent vehicle processing terminal via CAN bus or vehicle Ethernet.
6. The multimodal safety monitoring method for engineering vehicles according to claim 5, characterized in that, The specific steps of step S2 are as follows: S21. The intelligent vehicle processing terminal receives raw data from various sensors, decodes, corrects distortion and preprocesses noise in the video stream data of the surround view camera and the driver status monitoring camera, filters and clusters the point cloud data of the millimeter-wave radar, performs timestamp alignment and coordinate system normalization preprocessing on the data of the inertial measurement unit and the vehicle attitude sensor, and filters and verifies the validity of the data of the environmental perception sensor group. S22. Based on the current environmental parameters of light intensity, rain / snow intensity, and PM2.5 concentration, calculate the environmental confidence score of the i-th sensor at time t. ; S23. Obtain the signal strength, data loss rate, and self-test fault code for each sensor, and calculate the self-confidence score of the i-th sensor at time t. ; S24. Based on environmental confidence scores And the sensor's self-confidence score The dynamic confidence level of the i-th sensor at time t is calculated using the following formula. : in, and For preset weighting coefficients and .
7. The multimodal safety monitoring method for engineering vehicles according to claim 4, characterized in that, The specific steps of step S3 are as follows: S31. Calculate the obstacle detection results of each sensor, and perform spatiotemporal synchronization correlation of the obstacle detection results of different sensors. Based on the azimuth, distance and timestamp of the obstacle, determine whether it points to the same obstacle, and establish an obstacle cross-sensor correlation matrix. S32. When multiple sensors detect the same obstacle, obtain the current dynamic confidence level of each successfully associated sensor. and the distance measurement value output by the corresponding sensor. The distance to the merged obstacle is calculated using the following formula: Where n is the total number of sensors that detected this obstacle; S33. Based on the image recognition results from the visual sensor with the highest confidence, determine the category of obstacles using a pre-trained deep learning model for object detection, and output a fused list of environmental obstacle information, with each obstacle containing a corresponding fused distance. Category and azimuth; S34. Using the driver's facial image captured by the driver status monitoring camera, extract the driver's gaze direction vector and head posture Euler angles using computer vision algorithms, and output the driver's attention information, including the azimuth and pitch angles of the gaze point in the vehicle coordinate system.
8. The multimodal safety monitoring method for engineering vehicles according to claim 5, characterized in that, The specific steps of step S4 are as follows: S41. Obtain the current vehicle speed from the vehicle's CAN bus. and steering angle And obtain the vehicle's current roll angle from the inertial measurement unit. and pitch angle ; S42. Based on the Ackermann steering kinematics model, calculate the vehicle's position within a preset time window. Longitudinal driving distance within and maximum lateral offset ; in, This refers to the vehicle's wheelbase. S43. Determine a point on the ground projection with the vehicle's current position as the starting point and the longitudinal travel distance as the starting point. Radial length, with maximum lateral offset A sector-shaped region with a horizontal boundary; The sector is enclosed by the minimum turning radius boundary and the straight-line driving boundary, thus forming a dynamic dangerous movement trajectory sector. S44. Determine the roll angle Does it exceed the preset roll threshold? ; If so, extend the sector range in the lateral tilt direction by a preset rollover risk zone width. This area is designated as an extended zone and marked as a high-risk rollover zone. If not, proceed to step S5; Among them, the preset rollover risk zone width Calculated dynamically based on vehicle speed.
9. The multimodal safety monitoring method for engineering vehicles according to claim 7, characterized in that, The specific steps of step S5 are as follows: S51. Based on the driver's attention information, fuse the driver's gaze direction vector with the head posture Euler angles to calculate the driver's actual gaze direction in the vehicle coordinate system: Let the head yaw angle be... Pitch angle is The yaw angle of the line of sight relative to the head is Pitch angle The horizontal angle of the driver's actual gaze direction in the vehicle coordinate system. and vertical angle They are respectively: ; S52. Determine the driver's actual gaze direction By combining the spatial geometric relationships in the vehicle coordinate system with the driver's head as the origin, the projected coordinates of the gaze point on the ground are calculated. : Let the height of the driver's head from the ground be... Then the distance between the gaze point and the vehicle and gaze coordinates They are respectively: ; S53. Using the coordinates of the gaze point Generate a two-dimensional Gaussian-distributed attention heatmap centered on the target: The ground area around the vehicle in a 360° radius is divided into... Discrete grid, each grid cell Attention coefficient The calculation formula is: in, For grid cells The coordinates of the center point The radius of attention decay; S54. For each fused obstacle j generated in step S3, obtain the position coordinates of this obstacle in the vehicle coordinate system. Locate the corresponding grid cell and extract the attention coefficient for that grid cell. ; in, The value ranges from 0 to 1, with a larger value indicating a more focused driver attention on the area where the obstacle is located. S55. Based on the position coordinates of the j-th obstacle. Whether the vehicle falls within the dynamic dangerous motion trajectory sector generated in step S4, the vehicle's situational danger coefficient is obtained. : If it falls in, then ; If it does not fall in, then ; S56. Obtain the fusion distance of the obstacle generated in step S3. and read the preset safe distance threshold. ; Preset safe distance threshold According to vehicle speed Dynamic adjustment: In the formula, To preset the braking response time, To pre-set a safety margin; S57. The comprehensive risk coefficient shall be calculated using the following formula. : in, Let the overall risk coefficient be the j-th obstacle; For the distance term, an S-shaped function, when When the function value approaches 1, The function value approaches 0 as time progresses; , , These are the preset weight coefficients for the distance term, the attention deficit term, and the situation term, respectively, satisfying... ; is the steepness coefficient of the sigmoid function of the distance term.
10. The multimodal safety monitoring method for engineering vehicles according to claim 4, characterized in that, The specific steps of step S6 are as follows: S61. A first-level emergency warning threshold is preset in the intelligent vehicle-mounted processing terminal. and Level II emergency warning threshold ,and ; S62. Identify the overall risk factor for each obstacle. ; like This will trigger a Level 1 emergency alert: On the 360° panoramic image of the vehicle's touch screen, the obstacle target is highlighted with a first visual identifier, and the fused distance value and risk warning text are dynamically displayed. At the same time, the sound and light alarm emits a continuous sound and light alarm at the first frequency. like This will trigger a level 2 warning: On the 360° panoramic image on the in-vehicle touch screen, the obstacle target is marked with a second visual identifier and the fusion distance value is displayed. At the same time, the sound and light alarm emits an intermittent sound and light alarm at a second frequency. like If the obstacle is not triggered, the sound and light alarm will not be activated. Instead, the obstacle and its fusion distance will be displayed in the 360° panoramic image on the vehicle's touch screen in a normal color. A semi-transparent color block will be used to indicate the driver's blind spot in the low attention area of the attention heatmap.