Closed-loop alerting method, system, and electronic device based on pedestrian attention
Patent Information
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- SHANDONG XIEHE UNIV
- Filing Date
- 2026-07-03
- Publication Date
- 2026-08-04
AI Technical Summary
[0005]当前车辆行人警示技术领域存在的核心问题可以概括为:警示系统缺乏对行人认知状态的感知与闭环验证能力,导致简单地认为 “警示发出” 就等于 “警示有效”,从而在高危场景下,车辆与行人之间的安全风险无法得到可靠降低
[0019] According to one aspect of the present invention, this approach shifts from conventional forward prediction (determining whether a pedestrian will cross the road) to backward verification (determining whether a pedestrian has heard the warning). By establishing a differentiated secondary reminder mechanism guided by pedestrian attention closed-loop perception and intervention strategies, the effectiveness of the warning is effectively improved, making this approach more applicable to various scenarios, especially low-speed driving scenarios.
Smart Images

Figure CN122501248A_ABST
Abstract
Description
Technical Field
[0001] This invention relates to the field of vehicle engineering, and more particularly to a closed-loop warning method, system, and electronic device based on pedestrian attention. Background Technology
[0002] With the widespread adoption of new energy vehicles, their safety has become a major concern. Low-speed pedestrian warning systems (AVAS), as a crucial component ensuring safe interaction between vehicles and pedestrians, play an indispensable role. However, most current mainstream AVAS systems for new energy vehicles employ a single sensor input, fixed threshold triggering, and unidirectional open-loop control model, which has revealed numerous problems in practical applications.
[0003] In typical operation, the system primarily uses millimeter-wave radar or a monocular camera to detect the presence of pedestrians ahead, and combines this with the vehicle's speed and the distance to the target to issue a warning sound at a preset, fixed volume. While some models incorporate environmental noise compensation, automatically increasing the volume when ambient noise is high, this is merely a simple adjustment of volume and still fails to fundamentally address the crucial issue of "whether the warning is effective."
[0004] After a thorough review of existing technical solutions, the following specific shortcomings were identified: Pedestrian recognition and state perception are not precise enough: Under complex lighting conditions, inclement weather, and partial occlusion of the target, existing pedestrian recognition systems are prone to drift and abrupt changes in results. For example, single vision algorithms show a significantly higher false negative rate under conditions such as backlighting, rain, snow, and nighttime. Even though some systems employ multi-sensor fusion technology, most use fixed weights, which still easily leads to abrupt changes in recognition results in transitional lighting conditions. More importantly, existing systems can only identify "whether there is a pedestrian," but cannot identify "what the pedestrian is doing" or "whether they have seen a vehicle." For example, pedestrians distracted by their phones and those focused on crossing have drastically different probabilities of responding to warning sounds, but existing systems treat them equally, failing to adopt differentiated warning strategies for pedestrians in different states.
[0005] The core problems in the current field of vehicle pedestrian warning technology can be summarized as follows: warning systems lack the ability to perceive and verify the pedestrian's cognitive state, leading to the simplistic assumption that "issuing a warning" equals "the warning is effective." Consequently, in high-risk scenarios, the safety risks between vehicles and pedestrians cannot be reliably reduced. This not only affects pedestrian safety but also hinders the further promotion and application of new energy vehicles. Therefore, there is an urgent practical need for a new type of pedestrian warning system that can accurately perceive the pedestrian's cognitive state and possess closed-loop verification capabilities. Summary of the Invention
[0006] The technical problem to be solved by the present invention is to provide a closed-loop warning method, system and electronic device based on pedestrian attention.
[0007] To achieve the above-mentioned objectives, this invention provides a closed-loop warning method based on pedestrian attention, comprising the following steps: S1. Sensing pedestrians around the vehicle and identifying their attention state; where attention state is one of the following: focused, distracted, hesitant, or preparing to avoid. S2. Based on the attention state and physical collision risk, determine the initial warning strategy and output the initial warning signal; S3. Within a dynamic time window after the first warning signal is output, detect the pedestrian's feedback behavior to determine whether the pedestrian's attention has been successfully captured; S4. If it is determined that the traveler's attention has not been successfully captured, then based on the attention state, target characteristics, and environmental characteristics, the intervention strategy type is determined according to the preset priority rules. The intervention strategy type is one of the following: distraction behavior type, high-frequency insensitive response type, environmental masking type, and default type. S5. Based on the obtained intervention strategy type, adaptively select and execute a secondary reminder from a variety of preset cross-modal reminder strategies, wherein the cross-modal reminder strategies corresponding to different intervention strategy types use at least one of acoustic and auxiliary sensory modalities for reminder.
[0008] According to one aspect of the present invention, step S1, which involves sensing pedestrians around the vehicle and identifying their attentional states, includes: Acquire pedestrians' head posture, gaze direction, and handheld device status; Attention state determination is performed using a weighted voting fusion mode and / or a single feature determination mode, and the determination result is output. If the determination result satisfies both hesitation and avoidance preparation, then avoidance preparation shall prevail.
[0009] According to one aspect of the present invention, step S2, which involves determining the initial warning strategy and outputting the initial warning signal based on the attention state and physical collision risk, includes: The initial physical risk score is calculated based on three indicators: time to collision (TTC), relative velocity, and crossing intention. The initial physical risk score is corrected based on the pedestrian's attention state to obtain the final physical risk score, and a first mapping relationship is established between the final physical risk score and the volume compensation value; wherein, the higher the final physical risk score, the larger the volume compensation value. The lower limit of the warning sound volume is determined based on the final physical risk score, the first mapping relationship, and the environmental noise, while the upper limit of the warning sound volume is determined based on the pedestrian type and the pedestrian hearing protection threshold. The warning sound characteristics are determined based on the final physical risk score and the pedestrian's attention state. An initial warning strategy is constructed using the warning sound characteristics, the lower limit of volume, and the upper limit of volume. The initial warning signal is generated using the initial warning strategy. The warning sound characteristics include: warning audio segment, timbre, rhythm, and duration of a single sound.
[0010] According to one aspect of the present invention, in step S3, within a dynamic time window after the output of the first warning signal, the step of detecting the pedestrian's feedback behavior to determine whether the pedestrian's attention has been successfully captured, wherein the dynamic time window has a monotonically decreasing negative correlation with the collision time TTC; wherein, when the vehicle speed is less than or equal to a first speed threshold, the dynamic time window is calculated using a linear function, and the linear function is expressed as:
[0011] in, Indicates a dynamic time window. Indicates the collision time. , Indicates coefficient; When the vehicle speed exceeds the first speed threshold, the dynamic time window is calculated using an exponential function, which is expressed as:
[0012] in, , Indicates coefficient; The duration of the dynamic time window is limited to between the preset minimum and maximum window duration.
[0013] According to one aspect of the present invention, in step S3, within a dynamic time window after the first warning signal is output, the step of detecting the pedestrian's feedback behavior to determine whether the pedestrian's attention has been successfully captured includes continuously collecting the pedestrian's head deflection angle within the dynamic time window, calculating the change in head deflection angle relative to a preset deflection angle reference, and determining that the attention has been successfully captured if the change in head deflection angle is greater than or equal to a fourth preset angle threshold; or, within the dynamic time window, continuously collecting the pedestrian's gait and walking speed, calculating the change in walking speed relative to a preset walking speed reference, and determining that the attention has been successfully captured if the change in walking speed is greater than or equal to a preset percentage. If the change in head deflection angle is less than the fifth preset angle threshold and the pedestrian's gait remains unchanged, it is determined that the pedestrian's attention has not been successfully captured.
[0014] According to one aspect of the present invention, in step S4, the step of determining the intervention strategy type based on attention state, target characteristics, and environmental characteristics according to a preset priority rule is performed in the following priority order: First priority: if the attention state is distracted and the change in the pedestrian's head deflection angle is less than the fifth preset angle threshold, then it is judged as a distracted behavior type. The second priority is that if the pedestrian's age exceeds the age threshold and the main frequency of the first warning signal is higher than the preset frequency threshold and has been continuously sounding for more than the preset duration threshold, then it is determined to be a high-frequency insensitive response type. The third priority is that if the energy of the environmental noise in the preset frequency band near the main frequency of the first warning signal exceeds the energy ratio threshold of the first warning signal in the preset frequency band, it is determined to be an environmental masking type. Fourth priority: If none of the above conditions are met, it is determined to be the default type.
[0015] According to one aspect of the present invention, step S5, which involves adaptively selecting and executing a secondary reminder from a set of preset cross-modal reminder strategies based on the obtained intervention strategy type, includes: For distraction-related behaviors, play low-frequency pulse sounds within the first frequency range and simultaneously activate the visual alert module to flash lights; For high-frequency insensitive response types, reduce the main frequency of the secondary reminder tone to the second frequency range, shorten the pulse interval, and turn off the additional lights; For environmental masking type, the overall volume is increased by the first volume increment based on frequency domain masking compensation, and a rapid pulse mode is adopted; For the default type, the volume is increased based on the first warning signal and by the first volume increment, maintaining the tone of the first warning signal.
[0016] According to one aspect of the present invention, in step S1, in the step of perceiving pedestrians around the vehicle and identifying the attention state of the pedestrians, visual images, radar point clouds, environmental noise and illumination data around the vehicle are simultaneously collected as perception inputs to realize the perception of pedestrians around the vehicle; wherein, the fusion weight of the visual image and the radar point cloud is dynamically adjusted based on the illumination data, and the partial recognition confidence based on the visual image is corrected according to the attention state of the pedestrians. Step S5 also includes: After the second reminder is completed, step S3 is executed again to detect the pedestrian's attention. If the pedestrian's attention is still not captured, the event is marked as a high-risk unresponsive sample, and the complete record is uploaded to the cloud.
[0017] To achieve the above-mentioned objectives, the present invention provides a closed-loop warning system for the aforementioned pedestrian attention-based closed-loop warning method, comprising: The perception module is used to collect visual images, radar point clouds, environmental noise, and lighting data around the vehicle. The processing module is used to perceive pedestrians around the vehicle and identify their attention state, which is one of the following: focused, distracted, hesitant, or preparing to avoid a collision. Based on the attention state and the risk of physical collision, it determines an initial warning strategy and outputs an initial warning signal. Within a dynamic time window after outputting the initial warning signal, it detects the pedestrian's feedback behavior to determine whether the pedestrian's attention has been successfully captured. If it is determined that the pedestrian's attention has not been successfully captured, it determines an intervention strategy type according to a preset priority rule based on the attention state, target characteristics, and environmental characteristics. The intervention strategy type is one of the following: distracted behavior type, high-frequency insensitive response type, environmental masking type, or default type. Based on the obtained intervention strategy type, it adaptively selects and executes a secondary warning from a variety of preset cross-modal reminder strategies. The cross-modal reminder strategies corresponding to different intervention strategy types use at least one of the following methods: acoustic and auxiliary sensory modalities. The execution module includes a vehicle speaker and a lighting unit that support independent amplitude adjustment, and is used to output corresponding warning signals according to the control instructions of the processing module; The communication and storage module is used to cyclically store runtime data and support data interaction with the cloud.
[0018] To achieve the above-mentioned objectives, the present invention provides an electronic device, including a memory and a processor, wherein the memory stores a computer program, and the processor executes the computer program to implement the steps of the aforementioned closed-loop warning method based on pedestrian attention.
[0019] According to one aspect of the present invention, this approach shifts from conventional forward prediction (determining whether a pedestrian will cross the road) to backward verification (determining whether a pedestrian has heard the warning). By establishing a differentiated secondary reminder mechanism guided by pedestrian attention closed-loop perception and intervention strategies, the effectiveness of the warning is effectively improved, making this approach more applicable to various scenarios, especially low-speed driving scenarios.
[0020] According to one aspect of the present invention, this approach fundamentally changes the open-loop operation mode of existing vehicle and pedestrian warning systems, which "ends as soon as it is issued," by constructing a closed-loop control architecture of perception-decision-execution-feedback-optimization.
[0021] According to one aspect of the present invention, this approach effectively establishes a perceptible and verifiable interactive channel for pedestrian attention. For the first time, pedestrian attention states (focused, distracted, hesitant, and prepared to avoid) are introduced into the control loop as quantifiable feedback signals, enabling vehicles to actively "ask" and "verify" the warning effect, thus achieving a fundamental leap from one-way warning to two-way interaction.
[0022] According to one aspect of the present invention, this solution achieves precise secondary intervention based on cause diagnosis. When the initial warning is ineffective, the cause of failure is automatically diagnosed through priority rules (distraction behavior type > high-frequency insensitive response type > environmental masking type), and cross-modal alert strategies (low-frequency pulse, frequency reduction, light flashing, etc.) are adaptively selected, significantly reducing unnecessary acoustic interference while ensuring safety.
[0023] According to one aspect of the present invention, this solution effectively provides a low-cost, highly robust engineering solution; it does not rely on dedicated acoustic arrays or high-computing platforms, but achieves frontal near-field pointing enhancement through the left and right channel amplitude ratio, is compatible with high and low computing power platforms through a hot-switching mechanism, and ensures basic warning functions under abnormal operating conditions through a graded degradation strategy. Thus, this solution can be widely deployed in various types of vehicles, from economy to high-end, and has strong industrial universality.
[0024] According to one aspect of the present invention, this solution directly serves pedestrian safety protection systems for new energy vehicles, intelligent connected vehicles, and autonomous vehicles, and has broad application potential.
[0025] According to one aspect of the present invention, this approach enables a closed-loop optimization method that combines vehicle-cloud data with reinforcement learning, allowing warning strategies to be continuously iterated via OTA. As road test data accumulates, the system can automatically evolve differentiated warning strategies for different regions, populations, and traffic cultures, ultimately forming a personalized pedestrian interaction capability that is "one for each person" and has cross-scenario (such as construction zone warnings, school bus pick-up and drop-off, blind spot monitoring, etc.) transfer value.
[0026] According to one aspect of the present invention, this solution significantly reduces noise pollution in residential areas and at nighttime environments caused by unnecessary high-volume warnings while improving pedestrian safety, thus helping to alleviate the conflict between the fear of silent trams and the nuisance caused by warning sounds. Through a more precise and pedestrian-friendly interaction method, it promotes harmony and efficiency in the mixed-traffic environment for pedestrians and vehicles. Attached Figure Description
[0027] Figure 1 This is a flowchart illustrating the steps of the closed-loop warning method based on pedestrian attention according to the present invention. Detailed Implementation
[0028] The present invention will now be described in detail with reference to the accompanying drawings and specific embodiments. The embodiments cannot be described in detail here, but the embodiments of the present invention are not limited to the following embodiments.
[0029] like Figure 1 As shown, according to one embodiment of the present invention, a closed-loop warning method based on pedestrian attention includes the following steps: S1. Sensing pedestrians around the vehicle and identifying their attention state; where attention state is one of the following: focused, distracted, hesitant, or preparing to avoid. S2. Based on the attention level and the risk of physical collision, determine the initial warning strategy and output the initial warning signal; S3. Within a dynamic time window after the first warning signal is output, detect the pedestrian's feedback behavior to determine whether the pedestrian's attention has been successfully captured; S4. If it is determined that the traveler's attention has not been successfully captured, then based on the attention state, target characteristics, and environmental characteristics, the intervention strategy type is determined according to the preset priority rules. The intervention strategy type is one of the following: distraction behavior type, high-frequency insensitive response type, environmental masking type, and default type. S5. Based on the obtained intervention strategy type, adaptively select and execute a secondary reminder from a variety of preset cross-modal reminder strategies, wherein the cross-modal reminder strategies corresponding to different intervention strategy types use at least one of acoustic and auxiliary sensory modalities for reminder.
[0030] According to one embodiment of the present invention, in step S1, the step of perceiving pedestrians around the vehicle and identifying their attention status involves simultaneously acquiring visual images, radar point clouds, environmental noise, and illumination data around the vehicle as perception inputs to achieve the perception of pedestrians around the vehicle. Specifically, data from a camera, millimeter-wave radar, ultrasonic radar, a four-channel microphone array, and an illumination sensor are simultaneously acquired at a frequency of 50Hz. Timestamp alignment is performed using an onboard clock to control the time deviation of the multi-source data to within 50 milliseconds. When driving on bumpy roads or at high speeds, coordinate offsets are calculated in real time and dynamically compensated. Furthermore, to ensure the accuracy of data processing, the additional visual images, radar point clouds, environmental noise, and illumination data need to be preprocessed; specifically, this includes: Visual image preprocessing involves processing images captured by the camera. For example, it performs multi-level exposure fusion in low-light environments, local brightness compensation in backlit scenes, and rain removal algorithms in rainy or snowy weather. Simultaneously, it corrects inherent radial distortion of the lens and unifies the resolution to 1080P. These preprocessing steps do not alter the outlines and body features of pedestrians in the images, but significantly improve the robustness of subsequent algorithms.
[0031] Millimeter-wave radar data preprocessing: The raw point cloud from millimeter-wave radar (detection range 0.5-150 meters) contains a large amount of road surface reflections and clutter interference. Steady-state Kalman filtering is used for frame-by-frame processing to remove invalid point clouds and retain the trajectory points of moving targets. For monocular camera mode, the ground position is determined by the bottom contact point of the target detection box, combined with radar ranging to achieve pseudo-3D spatial mapping; in binocular mode, precise alignment between the radar coordinate system and the image pixel coordinate system is directly achieved.
[0032] For environmental noise preprocessing, a four-channel microphone array employs a delay-summing beamforming algorithm to direct the beam towards the center of the vehicle's front. Based on statistics of typical pedestrian crossing distances, it directionally picks up acoustic signals from the core interaction area within 3 meters in front of the vehicle. By running the LMS adaptive filtering algorithm, it eliminates vehicle self-noise such as engine noise, wind noise, and tire noise in real time, outputting the A-weighted environmental noise level with a measurement error controlled within 2 dB(A) and a self-noise filtering efficiency of no less than 80%.
[0033] Ultrasonic radar data preprocessing, based on ultrasonic radar (detection distance 0.1-5 meters), is used for blind spot compensation at close range, and is especially suitable for pedestrians who suddenly appear close to the front of the vehicle during the vehicle's starting phase or in narrow road sections.
[0034] like Figure 1 As shown, according to one embodiment of the present invention, step S1, which involves sensing pedestrians around the vehicle and identifying their attentional states, includes: S11. Obtain the pedestrian's head posture, gaze direction, and handheld device status; In this embodiment, pedestrians are detected from the collected multi-source data based on a target detection method to obtain the pedestrian's head posture, gaze direction, and handheld device status for subsequent attention state recognition. In this embodiment, each pedestrian in the scene is identified based on the target detection method, their type is distinguished (adult, elderly, child, wheelchair user, stroller companion, etc.), and a unique ID is assigned to each target for continuous tracking: Specifically, this includes: S111. Pedestrian target detection and classification, wherein a pre-trained YOLOv8s network is used for pedestrian target detection and classification, which outputs pedestrian bounding boxes and classifies them into five categories: adults (18-60 years old), elderly (over 60 years old), minors (under 12 years old), wheelchair users, and stroller caregivers.
[0035] S112. Special target differentiation (stroller, wheelchair, suitcase), which is implemented using triple verification logic, specifically including: Geometric feature verification is achieved based on the aspect ratio of the detection box falling within different preset ranges, such as strollers (approximately 1:2 to 1:3), wheelchairs (approximately 1:1.8 to 1:2.8), and rigid suitcases (approximately 1:1 to 1:1.5).
[0036] Motion trajectory jitter verification is achieved by calculating the displacement residual fluctuation frequency of the target center between consecutive frames. For example, due to the small wheel diameter and flexible load, the jitter frequency of a stroller is usually between 0.5-2Hz; the jitter characteristics of a wheelchair are different; and a suitcase has basically no jitter.
[0037] For static state verification, targets that have not moved for a continuous preset time (e.g., 300 milliseconds) are identified as stationary obstacles (e.g., temporarily placed suitcases) and no warning is triggered; only targets that are continuously moving are identified as vulnerable moving targets.
[0038] With the above settings, the special target differentiation can be used as a key influencing factor in pedestrian target detection based on the identified special targets. This can more accurately distinguish wheelchair users, stroller caregivers, etc., to achieve accurate judgment of their attention, and avoid misjudging special targets as ordinary pedestrians, which would cause the risk model to completely fail or fail to identify the additional risks brought by special targets, resulting in the risk of underestimating the real danger.
[0039] S113. Head posture, gaze direction, and handheld device status are detected through parallel processing. (1) Head posture estimation: The pitch angle, yaw angle and roll angle of the head are calculated using 68 facial key points.
[0040] (2) Direction of gaze: Based on head posture, the direction of eye gaze is further detected.
[0041] (3) Handheld device detection: The target detection network identifies whether the hand is holding an electronic device such as a mobile phone, and determines whether it is being operated based on the hand posture.
[0042] In this embodiment, when a pedestrian is occluded by more than 50% during visual detection, radar data is used to supplement the position and motion features to avoid target loss. Specifically, the DBSCAN density clustering algorithm is used to filter static obstacles, extract point cloud clusters of dynamic pedestrian targets, and calculate relative distance, relative velocity, and collision time TTC to supplement the occluded parts of the pedestrian.
[0043] In this embodiment, during the continuous tracking of each target, a Kalman filter is used to predict its position in the next frame, and Hungarian matching is performed with the detection results to achieve continuous tracking.
[0044] S12. Attention state determination is performed using a weighted voting fusion mode and / or a single feature determination mode, and the determination result is output. If the determination result simultaneously satisfies hesitation and avoidance preparation, avoidance preparation takes precedence. In this embodiment, in the step of determining attention state using a single feature determination mode and outputting the determination result, in head posture estimation, when the absolute value of the yaw angle exceeds a first preset angle threshold (preferably 30°) or the pitch angle downwards exceeds a second preset angle threshold (preferably 15°), it is directly determined as distraction. In gaze direction detection, if only the gaze deviates from the vehicle's driving direction by more than a third preset angle threshold (preferably 20°), it is also determined as temporary distraction, and the risk correction coefficient is halved. In handheld device status detection, if only a handheld electronic device is detected and the hand posture is in an operating state, it is directly determined as distraction.
[0045] Furthermore, in the step of determining attention state and outputting the determination result using a weighted voting fusion mode, the weighted voting fusion rule adopted is: head posture weight 0.5, gaze direction weight 0.3, and handheld device weight 0.2. When both head posture and handheld device are determined to be distracted, the overall determination is distracted; if only gaze is determined to be distracted while head posture is normal, it is determined to be "brief distraction," and the risk correction coefficient is halved. The above weights and thresholds can be calibrated based on a preset number (e.g., 500 groups) of pedestrian behavior samples collected from real vehicles. In this embodiment, halving the risk correction coefficient means that when the system determines that a pedestrian is in a "brief distraction" state, the negative impact of the distraction state on the risk score is reduced by half. This acknowledges the existence of risk while avoiding overreaction. The risk correction coefficient is a multiplier factor used to adjust the final risk score in the risk weighting classification. Under normal circumstances (complete distraction), such as looking down at a mobile phone, the risk correction coefficient = 1.0 (or other benchmark value). At this time, the distraction state will increase the final risk score. During brief moments of distraction, such as a quick glance at a phone or to the side while the head remains facing the road, the risk adjustment factor is 0.5 (halved).
[0046] In this implementation, the weighted voting fusion mode can be used in parallel with the single feature determination mode. When either rule is triggered, the system outputs the result with the higher confidence level. In situations where computing power is limited or in specific scenarios, the system can enable the single feature determination mode independently.
[0047] Therefore, the judgment results output based on the weighted voting fusion mode and / or the single feature judgment mode can accurately classify the attention state, that is: Focus: Head facing the road, eyes looking at vehicles, no handheld devices.
[0048] Distractions: looking down at a phone, wearing headphones, talking to someone, or looking away from the point of view.
[0049] Hesitation: lingering at the edge of the road, reciprocating forward-backward-forward movement, and signs of accelerating across the road (sudden increase in speed but unstable direction). Specific quantitative criteria: within 1 second, the standard deviation of the pedestrian's center of gravity displacement in the lateral (perpendicular to the road direction) is greater than 0.3m, or the rate of change of speed is greater than 50%.
[0050] Avoidance preparation: The pedestrian's head or body is clearly turned towards the vehicle, indicating that the pedestrian has noticed the vehicle. Specific quantitative judgment: The head yaw angle changes more than 10° towards the vehicle and the angular velocity is greater than 30° / s, lasting for more than 0.3 seconds.
[0051] If both hesitation and preparedness for avoidance are met, the preparedness for avoidance should be taken as the standard to avoid making the same decision repeatedly.
[0052] According to one embodiment of the present invention, in step S1, to ensure efficient and accurate pedestrian perception and attention state recognition, a computational power adaptive scheduling scheme is further set, which is based on real-time monitoring of the detection frame rate and inference latency. Specifically, when the detection frame rate is ≥50Hz and the inference latency is ≤80ms, it is determined to be in high-performance state, and the full-function network (including target detection, pose estimation, and attention state recognition) is run; when the frame rate is between 30 and 50Hz or the latency is between 80 and 100ms, it switches to a lightweight network (pose recognition is turned off, but target detection and attention state recognition are retained); when the frame rate is below 30Hz or the latency exceeds 100ms, it switches to an ultra-lightweight network (only target detection is run, and attention recognition is not run). The algorithm switching adopts a hot switching mechanism. When switching to the lightweight network, the single feature judgment mode is used by default.
[0053] In this implementation, for different computing power conditions, when switching algorithm branches, the features output by the new model are mapped to the feature space of the original model through the cached target feature vector and the pre-trained linear mapping matrix, so as to achieve smooth inheritance of target features with the same ID. The specific network switching mechanism is as follows: Feature caching: A FIFO queue stores the 128-dimensional feature vectors of each target from the most recent 30 frames, including position and velocity (8-dimensional), historical trajectory offset (20-dimensional), appearance features (64-dimensional), and motion features (36-dimensional). To ensure effective matching during switching, the system performs feature space alignment between the lightweight and ultra-lightweight networks during training: using the same batch of labeled data, feature vectors for the same target are extracted from both the full-featured and lightweight networks, and a linear mapping matrix is trained to map the features of the lightweight network to the feature space of the full-featured network. During switching, the features output by the new model are first transformed by this mapping matrix before similarity calculation is performed with the cached features.
[0054] The aforementioned linear mapping matrix was trained offline during the vehicle development phase: using a labeled dataset (no less than 10,000 frames) covering various lighting, angle, and occlusion conditions, feature pairs of the same target were extracted simultaneously through a full-featured network and a lightweight network, and the mapping matrix was solved using the least squares method. ,in, This represents the feature vector of a pedestrian target extracted by a fully functional network (high-performance model). It is a high-dimensional vector (e.g., 128-dimensional) representing the feature representation of the target under the "ideal model". Represents a linear mapping matrix that maps the features of a lightweight network. Mapping to the feature space of a full-featured network makes as close as possible This enables a smooth switching between different computing power branches. This represents the feature vector of the same pedestrian target extracted by a lightweight network (low-computing-power model). Dimension and The same (also 128-dimensional), but due to the simplification of the network structure, its original feature space is the same as... The values are not aligned; argmin represents the values of the independent variables that minimize the objective function; the least squares method for solving the mapping matrix means finding a matrix. This minimizes the sum of squared errors that follow. For example, this minimizes the sum of squared errors. The smallest x should be 5 / 3.
[0055] After training, the mapping matrix is stored as model parameters in the system and does not consume online inference computing power. The mapping matrices corresponding to different computing power branches are stored separately and dynamically loaded when switching.
[0056] ID Mapping Table: Employs a hash table structure. The key is the globally unique ID of the target, and the value is the corresponding 128-dimensional feature vector and the last occurrence timestamp. If a target is not detected for a preset time (2 seconds), it is deleted from the mapping table.
[0057] Matching during switching: The new model weights are loaded in the background. After loading, the inference interface is switched atomically without interrupting the video stream. For each detection box output by the new branch, candidate cache targets (within the range of predicted position ± preset distance) are first filtered according to the Kalman prediction position. Then, the cosine similarity between the current feature (after mapping) and the candidate features is calculated. The one with the maximum value and exceeding the preset similarity threshold (preferably 0.85) inherits the original ID; if all are below the threshold, a new ID is assigned.
[0058] Similarity threshold calibration: This threshold is calibrated based on 500 sets of target switching samples collected from real vehicles. The value that maximizes the accuracy of ID inheritance is selected. The default value is 0.85, and the configurable range is 0.80 to 0.90.
[0059] According to one embodiment of the present invention, in step S1, the step of sensing pedestrians around the vehicle and identifying their attention status involves dynamically adjusting the fusion weights of the visual image and the radar point cloud based on illumination data, and correcting the partial recognition confidence level based on the visual image according to the pedestrian's attention status. Specifically, in the step of dynamically adjusting the fusion weights of the visual image and the radar point cloud based on illumination data, 500 lux is used as the illumination boundary point. Under normal illumination (≥500 lux), the visual image is assigned a higher weight (approaching the upper limit of weight), while under severe illumination (<500 lux), the radar point cloud is assigned a higher weight (approaching the lower limit of weight). Within the transition range of 400-600 lux, a linear interpolation function is used to achieve a continuous and gradual change in weights, avoiding abrupt changes that could lead to jumps in the recognition results. The visual weights... ,in Light intensity (lux) The value range is 0.02-0.05. The value ranges from 0.3 to 0.5, and the final weight is capped between 0.2 and 0.8. Radar weight. The above parameters can be pre-determined based on 1000 hours of real vehicle testing.
[0060] In this embodiment, in the step of correcting the partial recognition confidence level based on visual images according to the pedestrian's attention state, for each pedestrian target, after completing the regular fusion (i.e., dynamically adjusting the fusion weights of the visual sensor and millimeter-wave radar at the decision layer and calculating the comprehensive confidence level of each pedestrian target), their attention state is additionally assessed. If it is determined to be distracted (especially looking down at a mobile phone or having their gaze deviated), the confidence level of the visual sensor in recognizing the target's crossing intention is automatically reduced by a preset weighting factor (50%). This ensures that the recognition process does not mistakenly determine that a distracted pedestrian has no crossing intention due to their unchanged head posture, but instead relies more on the continuous movement trajectory of the radar to determine whether they are about to cross.
[0061] Furthermore, a comprehensive confidence assessment is performed: For each target, the comprehensive confidence score is calculated as follows: Comprehensive Confidence = Target Detection Confidence × Visual Weight + Radar Matching Confidence × Radar Weight. Based on the comprehensive confidence assessment, the perception results for each pedestrian target are scored to filter out unreliable detections, providing credible input for subsequent risk decisions while simultaneously reducing computational resources. For distracted targets, the confidence score of the intent recognition part is multiplied by a weighting factor. If the comprehensive confidence score is lower than a preset threshold (0.7), the system re-infers once; if it is invalid three times consecutively, the target is marked as "risk-free" and enters the basic monitoring mode. In this embodiment, during the process of multiplying the confidence score of the intent recognition part by a weighting factor for distracted targets, if the system determines that a pedestrian is distracted (looking down at a mobile phone, gaze deviating), then the prediction result of their crossing intent will be discounted (e.g., multiplied by 0.5). This is because the head of a distracted pedestrian may always be facing down, and the crossing intent judged by the visual model based on head posture will be seriously unreliable (they may just be looking down at a mobile phone, but may suddenly take a step at any time). By reducing the confidence level of intent recognition, the reliance on visual intent is reduced when calculating risk scores in the future. Instead, the risk of crossing is judged more by the continuous movement trajectory of radar, thus avoiding misjudgment as "no intention to cross" when a distracted pedestrian's head is not moving.
[0062] Furthermore, if the overall confidence score is lower than a preset threshold (0.7), the system re-infers the target. If the result is invalid after three consecutive attempts, the target is marked as "risk-free" and enters the basic monitoring mode. During this process, for each target, the overall confidence score (detection reliability) is calculated first. If it is less than 0.7, it indicates that the current fusion result is unreliable (possibly due to temporary occlusion, sensor noise, etc.). The system does not discard the target immediately but re-executes the target detection and fusion (re-inference). If the overall confidence score is still <0.7 after three consecutive re-inferences, the target is abandoned, and no further attempts are made to enter the risk scoring process. Instead, the target is marked as "risk-free" (without triggering any warnings). Alternatively, it can enter "basic monitoring mode" (retaining only the most basic tracking, without performing complex calculations such as attention recognition and risk scoring). Its purpose is to prevent false triggers (low-confidence detections, such as tree shadows and radar clutter, will not be included in subsequent risk scoring, avoiding false warnings), handle transient disturbances (such as low confidence caused by brief obstruction or noise, which can be recovered through retry), protect computing power (after consecutive failures, there is no more meaningless consumption of computing resources, and it is downgraded to lightweight tracking), and provide a safety net (even if it cannot be reliably detected, it will not be mistakenly judged as "high risk", thus avoiding false braking or false horn blaring).
[0063] According to one embodiment of the present invention, step S2, which involves determining the initial warning strategy and outputting the initial warning signal based on the attention state and physical collision risk, includes: S21. Calculate the initial physical risk score using three indicators: Time to Collision (TTC), relative speed, and crossing intent. In this implementation, a three-part scoring system is used, with each indicator normalized to the 0-1 range and then weighted and summed. The weights are allocated as follows: TTC 0.4, relative speed 0.3, and crossing intent 0.3. The total score is mapped to a 0-10 range, with 0-3 indicating low risk, 4-7 indicating medium risk, and 8-10 indicating high risk. Crossing intent is determined by fitting the curvature of the motion trajectory over five consecutive frames. When the curvature exceeds 0.5 rad / m (approximately 30° / m) and the heading is towards the lane plane, a crossing intent is determined to exist.
[0064] S22. The initial physical risk score is corrected based on the pedestrian's attention state to obtain the final physical risk score, and a first mapping relationship is established between the final physical risk score and the volume compensation value; wherein, the higher the final physical risk score, the larger the volume compensation value; in this embodiment, if the pedestrian is in a distracted state, their equivalent vulnerability is increased by one level: adults are protected according to the hearing threshold of minors (75dB to 70dB), and minors are protected according to the hearing threshold of the elderly (70dB to 65dB). If the pedestrian is in a hesitant state (wandering, signs of accelerating across), the risk score increases by 2 points (out of 10), and the TTC trigger window is shortened by 20%. If the pedestrian is in a "preparing to avoid" state (head or body turning towards the vehicle), even if the current TTC is low, the system temporarily suspends increasing the volume and maintains the current volume output to avoid startling pedestrians who have already noticed the vehicle. Therefore, the first mapping relationship can be determined as follows: determine the volume compensation value based on the final risk score, where high risk is compensated by 15-20dB, medium risk by 5-10dB, and low risk by no sound or only outputting a very low volume (configurable).
[0065] S23. The lower limit of the warning sound volume is determined based on the final physical risk score, the first mapping relationship, and the environmental noise. The upper limit of the warning sound volume is determined based on the pedestrian type and the pedestrian hearing protection threshold. In this embodiment, the step of determining the upper limit of the warning sound volume based on the pedestrian type and the pedestrian hearing protection threshold is determined using a multi-objective priority processing method: when there are multiple pedestrians in the scene, the system executes control according to the highest risk score and the strictest hearing threshold. The protection priority is as follows: infants > elderly > people with mobility impairments (wheelchair users) > distracted adults > normal adults. The preset hearing protection thresholds for various types of pedestrians are: 75dB(A) for normal adults, 70dB(A) for minors (including infants), and 70dB(A) for the elderly. In quiet areas or special working conditions, the system can further reduce the upper limit to 65dB(A), which applies to all pedestrians. The above thresholds are set with reference to psychoacoustic experimental data and can be fine-tuned according to vehicle model calibration.
[0066] S24. Based on the final physical risk score and the pedestrian's attention state, the warning sound characteristics are determined, and an initial warning strategy is constructed using the warning sound characteristics, a lower volume limit, and a higher volume limit. An initial warning signal is then generated using this strategy. The warning sound characteristics are composed of several elements selected from volume, warning audio range, timbre, rhythm, single sound duration, masking compensation, and quiet adaptation. In this embodiment, over-frequency masking compensation technology reduces the peak sound pressure level while maintaining subjective loudness, thus minimizing disturbance to pedestrians. Furthermore, the initial warning strategy specifically includes the following aspects: The base volume is interlocked, and the final volume is: ; The safety lower limit is the larger of 30dB(A) and ambient noise + 10dB(A), and the volume upper limit is no more than 85dB(A). For example, if the ambient noise is 55dB(A), the medium-risk compensation is 10dB, and the adult hearing threshold is 75dB, then min(65,75) = 65dB, and the larger of the two is taken as min(65,75) = 65dB. The output volume is 65dB.
[0067] Frequency domain masking compensation is achieved by pre-storing various typical environmental noise spectral templates and their corresponding masking compensation gain curves. Each template contains multiple typical energy values (in dB) across 1 / 3 octave bands (20Hz-20kHz), such as 31 templates. After real-time acquisition of environmental noise, its spectral similarity (e.g., cosine similarity) with each template is calculated. The best-matching template is selected, and its pre-calculated compensation gain curve is loaded to perform frequency domain compensation for the warning sound.
[0068] The alert frequency band switching feature proactively shifts the main frequency component of the warning sound from the default range to a preset alert frequency band (3.5-4.5kHz) for pedestrians deemed distracted (especially those looking down at their phones). Given the human ear's sensitivity to this frequency band, it is easy to attract attention even at low volumes. This switching only changes the frequency components and does not increase the overall sound pressure level, therefore it will not cause additional disturbance to others.
[0069] Quiet zone adaptation: When the system determines that it is currently in a quiet zone, the final output volume is forced to not exceed 45dB(A), and a soft sound effect is switched (rising edge of a sine wave, without abrupt impact). Quiet zone determination uses a weighted scoring method: ambient noise ≤40dB weight 40%, vehicle speed ≤20km / h weight 30%, time period 22:00-6:00 weight 20%, residential / school zone roads weight 10% (can be calibrated). Quiet mode is entered when the total weighted score is ≥60%. The above weights and thresholds can be adjusted through preset parameters, which will not be elaborated here.
[0070] Based on the various constraints generated by the aforementioned initial warning strategy, warning sounds of varying intensities and rhythms are output according to risk levels. These sounds are then directed forward via vehicle speakers to reduce noise diffusion from the sides and rear. Specifically, the initial warning strategy is tiered, and the warning sound output is based on this tiered approach. Low risk (0-3 points): No acoustic alerts are triggered.
[0071] Medium risk (4-7 points): Outputs a soft warning sound with a low frequency and a smooth rise edge, with a single sound lasting 1 second.
[0072] High risk (8-10 points): Outputs a standard warning tone (similar to the sweep tone of traditional AVAS), with a single sound duration of 1.5 seconds.
[0073] In this embodiment, during the output of the initial warning signal, the vehicle's speakers can be further utilized to enhance the near-field sound level. For vehicles equipped only with left and right channel speakers, the system achieves near-directional sound emission by adjusting the amplitude ratio of the left and right channels. Specifically, the main channel (the side closest to the pedestrian) outputs 100% amplitude, the secondary channel outputs 50% to 70% amplitude, and the amplitude difference between the left and right channels is controlled within the range of 30% to 50%. This amplitude difference causes the sound image to shift towards the main channel. Simultaneously, due to the natural directivity of the speakers at the front bumper, the sound pressure level in the horizontal region 30°-90° in front of the vehicle is 10-15 dB higher than that in the rear. For vehicles equipped with multi-channel arrays, beamforming algorithms can be used to achieve more precise directional control.
[0074] In this embodiment, a strong reflection scene compensation scheme can be further set during the output of the first warning signal. When the vehicle is driving near strong reflection scenes such as road guardrails and tunnel walls, the radar RCS value will increase significantly (the RCS of multiple consecutive frames exceeds more than 3 times the RCS reference value corresponding to the background noise). At this time, the multipath reflection of sound waves may cause distortion of the sound field in front. The system automatically adjusts the phase and amplitude of the left and right channels to compensate for the interference caused by the reflection.
[0075] like Figure 1 As shown, according to one embodiment of the present invention, in step S3, within a dynamic time window after the output of the first warning signal, the pedestrian's feedback behavior is detected to determine whether the pedestrian's attention has been successfully captured. In this step, the dynamic time window and the collision time TTC have a monotonically decreasing negative correlation; therefore, the smaller the collision time TTC, the shorter the window, ensuring rapid response in high-risk scenarios. Furthermore, when the vehicle speed is less than or equal to a first speed threshold (e.g., 30 km / h), the dynamic time window is calculated using a linear function, and the linear function is expressed as:
[0076] in, Indicates a dynamic time window. Indicates the collision time. , The coefficients are represented; in this embodiment, a = -0.02, b = 1.0.
[0077] When the vehicle speed exceeds the first speed threshold, the dynamic time window is calculated using an exponential function, which is expressed as:
[0078] in, , The coefficients are represented by c; in this embodiment, c = 0.3 and d = 0.2.
[0079] The window duration of the dynamic time window is limited to between the preset minimum window duration (e.g., 0.3s) and the maximum window duration (e.g., 1.2s).
[0080] like Figure 1 As shown, according to one embodiment of the present invention, in step S3, within a dynamic time window after the first warning signal is output, the pedestrian's feedback behavior is detected to determine whether the pedestrian's attention has been successfully captured. Within the dynamic time window, the pedestrian's head deflection angle is continuously collected, and the change in head deflection angle relative to a preset deflection angle benchmark is calculated. If the change in head deflection angle is greater than or equal to a fourth preset angle threshold (preferably 15°), it is determined that attention has been successfully captured. Alternatively, within the dynamic time window, the pedestrian's gait and walking speed are continuously collected, and the change in walking speed relative to a preset walking speed benchmark is calculated. If the change in walking speed is greater than or equal to a preset percentage (preferably 20%), it is determined that attention has been successfully captured. In this embodiment, the preset deflection angle benchmark is obtained based on the average head deflection angle within 0.5 seconds before the first warning is issued. In this embodiment, the specific implementation method for gait / walking speed change detection is as follows: the system extracts the pedestrian's lower limb key points through a human pose estimation network (such as HRNet or OpenPose), including six key points: left and right hips, left and right knees, and left and right ankles. The instantaneous moving velocity of a pedestrian is calculated based on the displacement of ankle joint key points between consecutive frames. ,in For the first Frame and the The distance between frames projected onto the road plane by the ankle joint. This is the inter-frame time interval. The instantaneous velocity sequence is taken within 0.5 seconds before the first warning is issued. Calculate the average value ; Retrieve feedback within the detection window (window duration) Instantaneous velocity sequence Calculate the average value Rate of change of walking speed .in, This indicates the initial speed and the final speed within 0.5 seconds prior to the issuance of the first warning. This indicates the start and end speeds within the detection window.
[0081] like (A significant increase in walking speed usually indicates that the pedestrian is speeding up to avoid an obstacle or to cross the road), or A significant decrease in walking speed, usually indicating hesitation or stopping, is considered a valid change in gait. In addition, a change in cadence (number of steps per unit time) exceeding 25%, or a sudden change in the curvature of the movement trajectory of key points of the lower limbs, can also be used as supplementary criteria for judgment.
[0082] Furthermore, if the change in head deflection angle is less than the fifth preset angle threshold (preferably 5°) and the pedestrian's gait remains unchanged, it is determined that the pedestrian's attention has not been successfully captured.
[0083] Furthermore, when the distance to the pedestrian exceeds a preset distance threshold (such as 15 meters), the weight of gait detection is reduced, and it mainly relies on changes in head yaw angle.
[0084] According to one embodiment of the present invention, in step S4, the step of determining the intervention strategy type based on attention state, target characteristics, and environmental characteristics according to a preset priority rule is performed in the following priority order: First priority: if the attention state is distracted and the change in the pedestrian's head deflection angle is less than the fifth preset angle threshold, then it is judged as a distracted behavior type. The second priority is to determine if the pedestrian's age exceeds the age threshold and the main frequency of the initial warning signal is higher than the preset frequency threshold and has been continuously emitted for more than the preset duration threshold, then it is determined to be a high-frequency insensitive response type; for example, if the target is an elderly person (age > 60 years old) and the main frequency of the current warning sound is > 3kHz, and any of the following conditions are met, it is determined to be a high-frequency insensitive response type: Condition A: The initial warning sound has been continuously emitted for ≥0.5 seconds, and no attention capture has been detected within this duration; Condition B: Duration of the current feedback detection window If the warning sound is emitted for 0.5 seconds (i.e., in high-risk scenarios with a smaller TTC), then there is no need to meet the requirement of continuous sound emission for 0.5 seconds. As long as the warning sound has been emitted and the main frequency is >3kHz, and attention is not captured after the feedback window ends, it is directly judged as a high-frequency insensitive response type.
[0085] The third priority is that if the energy of the ambient noise in the preset frequency band near the main frequency of the first warning signal exceeds the energy ratio threshold of the first warning signal in the preset frequency band, it is determined to be an environmental masking type; for example, the energy of the ambient noise in the ±1 / 3 octave band of the main frequency of the warning sound exceeds the energy ratio threshold of the warning sound in that frequency band (preferably 70%).
[0086] Fourth priority: If none of the above conditions are met, it is determined to be the default type.
[0087] The above settings enable the process of causal diagnosis and strategy matching from the discovery of ineffective alerts to the decision of what secondary reminder to take. By using priority rules (distraction behavior > high-frequency insensitive response > environmental masking), the most likely cause of uncaptured attention is determined, and based on the diagnostic results, a specific intervention strategy type (e.g., distraction behavior) is output. This priority then drives the next step to execute the corresponding, differentiated secondary reminder action.
[0088] According to one embodiment of the present invention, step S5, which involves adaptively selecting and executing a secondary reminder from a set of preset cross-modal reminder strategies based on the obtained intervention strategy type, includes: For distraction-related behaviors, play a low-frequency pulse sound within the first frequency range and simultaneously activate the visual alert module to flash the lights; for example, play a low-frequency pulse sound within the first frequency range (preferably 200-300Hz), with a pulse width of 0.2 seconds, an interval of 0.1 seconds, and repeat 3 times; at the same time, control the daytime running lights to flash twice at a frequency of 2Hz (0.1 seconds each time, with a light intensity 1.5 times that of normal driving lights; the normal light intensity of daytime running lights is 400-600cd, which is increased to 600-900cd when flashing, without being dazzling).
[0089] For high-frequency insensitive response types, the main frequency of the secondary reminder tone is reduced to the second frequency range, the pulse interval is shortened, and the additional light is turned off; for example, the main frequency of the secondary reminder tone is reduced to the second frequency range (preferably 1.5kHz±10%), the pulse interval is shortened to the first preset interval (preferably 0.15 seconds), and repeated 3 times; no additional light is added (to avoid visual confusion).
[0090] For environmental masking, the overall volume is increased by a first volume increment based on frequency domain masking compensation, and a rapid pulse mode is adopted; for example, based on frequency domain masking compensation, the overall volume is increased by a first preset increment (preferably 3dB) (not exceeding the upper limit of 85dB), and a rapid pulse mode is adopted (0.1 seconds on, 0.05 seconds off, repeated 4 times).
[0091] For the default type, the volume is increased based on the first warning signal and the first volume increment is used to maintain the tone of the first warning signal. For example, the overall volume is increased by 3dB based on the first warning signal without changing the tone, and this is repeated twice.
[0092] In this embodiment, if the collision time (TTC) is less than 2 seconds and the pedestrian's attention is still not captured after a second warning, a short horn warning at the highest level (85dB, once every 0.3 seconds, limited to one time) is triggered. Simultaneously, a warning is sent to the driver's instrument panel, and the driver takes over braking control (the system itself does not perform braking). Furthermore, a maximum of two second warnings are triggered in a single event, with a cumulative audible duration not exceeding 3 seconds. If the pedestrian still does not react, the system stops audible, maintaining only the maximum permissible volume output, and records this event as a high-risk unresponsive sample. The complete record (including raw sensor data, recognition results, warning parameters, and feedback results) is uploaded to the cloud for model optimization. In this embodiment, after the second warning, step S3 is executed again to detect the pedestrian's attention, thereby determining whether the pedestrian's attention has been successfully captured.
[0093] According to one embodiment of the present invention, the closed-loop warning method based on pedestrian attention further includes: graded degrading processing for abnormal operating conditions; used to automatically degrade the system's operation in extremely noisy environments, sensor failures, or insufficient computing power to ensure that the basic warning function is not interrupted; and smooth recovery after the abnormality is resolved. This includes: When ambient noise exceeds the first noise threshold, the upper limit constraint on hearing is lifted. Specifically, when the A-weighted ambient noise exceeds 85 dB(A), the system directly outputs an alarm sound at the upper limit of 85 dB(A), no longer constrained by the crowd's hearing threshold (because even if the threshold is exceeded at this point, no additional disturbance will occur). When the ambient noise is between 75 dB(A) and 85 dB(A), the system enters a pre-degradation state, and the upper limit of volume is relaxed to 80 dB. When the ambient noise drops below 70 dB(A) and remains below this level for 1 second, the system smoothly returns to the aforementioned constraint state.
[0094] When a sensor fails, the system switches to a mode that relies solely on the remaining sensors and lowers the risk threshold. When the detection frame rate falls below the frame rate threshold or the inference latency exceeds the latency threshold, non-core identification modules are shut down. Specifically, a sensor is considered faulty if its data is continuously lost for more than 100 milliseconds or its data error exceeds 50%. In the event of a single sensor failure, the system switches to a mode that relies solely on the remaining sensors and lowers the risk threshold by 30%. Once the faulty sensor data returns to normal (no data loss for multiple consecutive frames, error ≤ preset value), the system switches back to multi-source fusion mode with a smooth transition time of 0.5 seconds (weights gradually recover).
[0095] When computing power is insufficient (i.e., inference latency exceeds 100 milliseconds or detection frame rate is below 30Hz), the system automatically shuts down non-core modules (such as pose recognition and attention recognition), running only object detection and basic alert functions. When the detection frame rate recovers to above 50Hz and the inference latency drops below 80 milliseconds for 2 seconds, the system automatically restores the full-function network.
[0096] When multiple anomalies occur simultaneously, the highest priority degradation strategy is executed in the order of extreme noise environment, sensor failure, and insufficient computing power, and a smooth recovery is achieved after the anomaly is resolved. For example, if sensor failure and extreme noise environment occur simultaneously, the extreme noise degradation strategy is executed first. After the highest priority anomaly is resolved, the corresponding strategies are executed in order of priority for the remaining anomalies until all are recovered.
[0097] According to one embodiment of the present invention, the closed-loop warning method based on pedestrian attention further includes: cloud data closed-loop iteration, specifically including: Each warning event and attention feedback result is uploaded to the cloud, marking high-risk unresponsive samples. In this embodiment, the cloud stores the uploaded high-risk unresponsive samples in real time, and the storage format is: ID of each target, classification label, attention status, comprehensive confidence level, risk score, output volume, environmental noise, avoidance feedback result, and attention capture result. The data is stored in a circular queue, retaining the most recent 30 days. Moreover, the cloud automatically detects the following situations and marks them as abnormal samples: pedestrian missed detection rate >10%, special target misjudgment rate >12%, volume control deviation >8dB(A), and avoidance judgment error ≥3 times / hour. In addition, a new category of "high-risk unresponsive samples" is added: events where the collision time TTC is less than 3 seconds and attention is still not captured after a second reminder, and there is no physical avoidance behavior.
[0098] The warning policy parameters are updated using a cloud-maintained policy optimization model, and the updated policy parameters are then sent to the vehicle via OTA. In this embodiment, the policy optimization model is a reinforcement learning model. A deep Q-network (DQN) can be maintained in the cloud to optimize the following policy parameters (replacing fixed rules). It should be noted that the reinforcement learning policy optimization described in this step is only a preferred embodiment of the present invention. Those skilled in the art can use other policy optimization models (such as PPO, A3C, evolutionary strategies, etc.) to replace the aforementioned DQN network, as long as they can achieve adaptive adjustment of the warning policy parameters based on historical data.
[0099] In this implementation, the reinforcement learning model is configured as follows: Specifically, the state space (6-dimensional, normalized to [0,1]) includes: ambient noise level (dB / 100), risk score (0-1), attention state (one-hot encoding), historical attention capture rate, TTC(s) / 5, and whether it is a quiet area (0 / 1). The action space (discrete, 24 combinations) includes: volume compensation level (0 / 3 / 6 / 9dB) × secondary reminder mode (no / low-frequency pulse / reduced frequency / rapid pulse) × whether to add light (yes / no). The reward function is: attention capture +10, physical avoidance action (speed reduction ≥30% or heading deviation ≥15°) +5, volume exceeding 85dB or user complaint flag -2, and ineffective after secondary reminder -5. The network structure and training are: a three-layer fully connected network (128-64-32), ε-greedy exploration (initial ε=0.3, decaying by 0.01 every 1000 steps). Training is triggered once every 500 valid events, using experience replay (buffersize=2000), and the target network is updated every 100 steps. The trained policy parameters are then delivered to the vehicle via OTA.
[0100] According to one embodiment of the present invention, the present invention provides a closed-loop warning system for the aforementioned pedestrian attention-based closed-loop warning method, comprising: The perception module is used to collect visual images, radar point clouds, environmental noise, and lighting data around the vehicle. The processing module is used to perceive pedestrians around the vehicle and identify their attention state, which is one of the following: focused, distracted, hesitant, or preparing to avoid collision. Based on the attention state and the risk of physical collision, it determines the initial warning strategy and outputs the initial warning signal. Within a dynamic time window after the initial warning signal is output, it detects the pedestrian's feedback behavior to determine whether the pedestrian's attention has been successfully captured. If it is determined that the pedestrian's attention has not been successfully captured, it determines the intervention strategy type according to the attention state, target characteristics, and environmental characteristics, based on a preset priority rule. The intervention strategy type is one of the following: distracted behavior type, high-frequency insensitive response type, environmental masking type, or default type. Based on the obtained intervention strategy type, it adaptively selects and executes a secondary warning from a variety of preset cross-modal reminder strategies. The cross-modal reminder strategies corresponding to different intervention strategy types use at least one of the following methods: acoustic and auxiliary sensory modalities. The execution module includes an onboard speaker and a lighting unit that support independent amplitude adjustment, used to output corresponding warning signals according to the control instructions of the processing module; The communication and storage module is used to cyclically store runtime data and support data interaction with the cloud.
[0101] For specific limitations regarding the closed-loop warning system, please refer to the limitations of the pedestrian attention-based closed-loop warning method mentioned above, which will not be repeated here. Each module in the aforementioned closed-loop warning system can be implemented entirely or partially through software, hardware, or a combination thereof. These modules can be embedded in or independent of the processor in a computer device, or stored in the memory of a computer device as software, so that the processor can call and execute the corresponding operations of each module.
[0102] In this embodiment, the memory may be, but is not limited to, Random Access Memory (RAM), Read Only Memory (ROM), Programmable Read-Only Memory (PROM), Erasable Programmable Read-Only Memory (EPROM), Electrically Erasable Programmable Read-Only Memory (EEPROM), etc.
[0103] In this embodiment, the processor can be an integrated circuit chip with signal processing capabilities. The processor can be a general-purpose processor, including a central processing unit (CPU), a network processor (NP), an in-vehicle domain controller, or an embedded AI processor; it can also be a digital signal processor (DSP), an application-specific integrated circuit (ASIC), a field-programmable gate array (FPGA), or other programmable logic devices, discrete gate or transistor logic devices, or discrete hardware components.
[0104] In this embodiment, visual images are acquired using a 1080P monocular / dual-lens camera (frame rate ≥30fps, field of view 120°), preferably with a recognition accuracy ≥90% and image distortion rate ≤1% under normal lighting conditions. Radar point clouds are acquired using millimeter-wave radar (close-range (0.5-5m) detection error ±0.2m, medium-to-long-range (5-150m) detection error ±0.5m, refresh rate 50Hz) and ultrasonic radar (detection range 0.1-5m, used for near-range blind spot compensation). Environmental noise is acquired using a four-channel microphone array (sampling rate ≥48kHz, A-weighted environmental noise acquisition error ≤2dB(A), self-noise filtering efficiency ≥80%). Illumination data is acquired using an illumination sensor (measurement range 0-1000lux, accuracy ±10lux, response delay ≤100ms).
[0105] In this embodiment, the execution module can be a vehicle speaker (left and right channels, supporting independent amplitude adjustment), an AVAS low-speed warning unit, a power amplifier module, and a daytime running light / LED light group (for cross-modal light reminder).
[0106] like Figure 1 As shown, according to one embodiment of the present invention, the present invention provides an electronic device, including a memory and a processor. The memory stores a computer program, and the processor executes the computer program to implement the steps of the aforementioned closed-loop warning method based on pedestrian attention.
[0107] To further illustrate this plan, further examples will be provided.
[0108] Example 1 Scene: 6 PM, at a zebra crossing near a middle school, three students are crossing together. Two of them are looking at their phones, and one is wearing headphones. Ambient noise level is 55 dB(A), illuminance is 300 lux (under streetlights at night), and the vehicle is approaching at 30 km / h.
[0109] System Operation: The system completed spatiotemporal synchronization and operated a full-function network. Three pedestrian targets were detected, identified as "minors" by age and "distracted" by attention status. Illumination at 300 lux was considered poor, prioritizing radar signaling. Due to the distracted state, the system automatically reduced the visual sensor's trust in the pedestrian's intention to cross by 50%.
[0110] Risk score calculation: TTC approximately 2.5 seconds, relative speed 30km / h, clear crossing intent, weighted total score of approximately 7.5 points, plus a "distraction" correction item, the final risk level is high risk. Ambient noise 55dB, high risk compensation 15dB, base volume 70dB, hearing threshold for minors 70dB, final volume after interlocking 70dB. Frequency domain masking compensation matches the urban traffic noise template, specifically boosting 1-2kHz components; simultaneously, the warning sound main frequency is switched to 4kHz.
[0111] High-risk level, outputs standard warning sound, and achieves frontal near-field enhancement through a 40% amplitude difference between the left and right channels.
[0112] The detection window is calculated at TTC=2.5s, and the linear function yields a window of approximately 0.95s. Head deflection is detected within the window—two students with their heads down deflect less than 5°, and one student wearing headphones deflects approximately 8° (less than 15°). "Attention not captured" is determined. Cause determination: Priority 1 is met (distracted behavior), output "distracted behavior type". A secondary reminder is triggered: a 200Hz low-frequency pulse sound is played 3 times, while the daytime running light flashes rapidly 2 times.
[0113] After a second reminder and a second check, both students looked up and turned their heads more than 15 degrees, indicating that their attention was successfully captured; the other student removed their headphones and also made a dodge movement. The system then terminated the alert.
[0114] Example 2 Scenario: Construction site on a main urban road, ambient noise level 85 dB(A) (piling machine, truck noise), an elderly person in a wheelchair attempts to cross the road. The vehicle is traveling slowly at 15 km / h.
[0115] System operation process: The system identifies the user as an "elderly person in a wheelchair" with an "attentive" status. The risk score is 6 (medium risk), but the ambient noise level of 85dB exceeds the extreme noise threshold, triggering an abnormal degradation. The system enters degradation mode, releases auditory constraints, and outputs volume at the maximum limit of 85dB(A).
[0116] Frequency domain masking compensation was applied to match the construction noise template. It was found that the low-frequency band was completely masked. Therefore, the main frequency of the warning sound was selected as 2kHz (to avoid the main frequency of construction noise). The warning sound was then output.
[0117] Inside the detection window, the elderly person's head turned about 20°, and their attention was successfully captured, requiring no further reminder.
[0118] After the vehicle leaves the construction area, the ambient noise drops to 68dB(A) and remains there for 1 second, at which point the system automatically resumes the normal interlock rules.
[0119] Example 3 Scenario: 11 PM, on a residential road, ambient noise level 35 dB(A), an adult is walking normally with no intention of crossing. This vehicle is following slowly at 10 km / h.
[0120] System operation: The system identifies the user as an adult with a "focused" attention state. A risk score of 1 (low risk) is assigned, and no sound is triggered. However, in quiet zone mode, a soft 40dB sound effect is output as a reminder.
[0121] The pedestrian's head turned within the detection window, indicating attention was captured, and no secondary alert was required.
[0122] Example 4 Scenario: A 70-year-old elderly person is looking down at their phone while crossing the street. The ambient noise level is moderate (55dB). The initial warning sound is at a frequency of 4kHz and lasts for 0.5 seconds. After that, a head turn of less than 5° is detected, indicating that the person's attention has not been captured.
[0123] Cause determination process: (1) Check priority 1: Attention state is "distracted" and head deflection <5° → satisfied → output distracted behavior type.
[0124] (2) Although priority 2 (elderly and no response to high-frequency warnings) is partially met, priority 1 takes precedence according to the conflict resolution rules.
[0125] (3) Implement secondary reminders: Play low-frequency pulse sounds (200-300Hz) and add flashing lights according to the distraction behavior strategy.
[0126] Result: The elderly person's attention was successfully captured by the low-frequency pulse sound, which caused them to lift their head. This was because the low-frequency pulse sound was within the elderly person's hearing sensitivity range, thus successfully attracting their attention.
[0127] The above description is merely an example of a specific solution of the present invention. For any devices and structures not described in detail herein, it should be understood that they are implemented using common devices and methods already available in the art.
[0128] The above description is merely one embodiment of the present invention and is not intended to limit the invention. Various modifications and variations can be made to the invention by those skilled in the art. Any modifications, equivalent substitutions, or improvements made within the spirit and principles of the present invention should be included within the scope of protection of the present invention.
Claims
1. A closed-loop alerting method based on pedestrian attention, characterized in that, Includes the following steps: S1. Sensing pedestrians around the vehicle and identifying their attention state; where attention state is one of the following: focused, distracted, hesitant, or preparing to avoid. S2. Based on the attention state and physical collision risk, determine the initial warning strategy and output the initial warning signal; S3. Within a dynamic time window after the first warning signal is output, detect the pedestrian's feedback behavior to determine whether the pedestrian's attention has been successfully captured; S4. If it is determined that the traveler's attention has not been successfully captured, then based on the attention state, target characteristics, and environmental characteristics, the intervention strategy type is determined according to the preset priority rules. The intervention strategy type is one of the following: distraction behavior type, high-frequency insensitive response type, environmental masking type, and default type. S5. Based on the obtained intervention strategy type, adaptively select and execute a secondary reminder from a variety of preset cross-modal reminder strategies, wherein the cross-modal reminder strategies corresponding to different intervention strategy types use at least one of acoustic and auxiliary sensory modalities for reminder.
2. The attention-based closed-loop alerting method of claim 1, wherein, Step S1, which involves sensing pedestrians around the vehicle and identifying their attentional states, includes: Acquire pedestrians' head posture, gaze direction, and handheld device status; Attention state determination is performed using a weighted voting fusion mode and / or a single feature determination mode, and the determination result is output. If the determination result satisfies both hesitation and avoidance preparation, then avoidance preparation shall prevail.
3. The attention-based closed-loop alerting method of claim 2, wherein, Step S2, which involves determining the initial warning strategy and outputting the initial warning signal based on the attention state and physical collision risk, includes: The initial physical risk score is calculated based on three indicators: time to collision (TTC), relative velocity, and crossing intention. The initial physical risk score is corrected based on the pedestrian's attention state to obtain the final physical risk score, and a first mapping relationship is established between the final physical risk score and the volume compensation value; wherein, the higher the final physical risk score, the larger the volume compensation value. The lower limit of the warning sound volume is determined based on the final physical risk score, the first mapping relationship, and the environmental noise, while the upper limit of the warning sound volume is determined based on the pedestrian type and the pedestrian hearing protection threshold. The warning sound characteristics are determined based on the final physical risk score and the pedestrian's attention state. An initial warning strategy is constructed using the warning sound characteristics, the lower limit of volume, and the upper limit of volume. The initial warning signal is generated using the initial warning strategy. The warning sound characteristics include: warning audio segment, timbre, rhythm, and duration of a single sound.
4. The attention-based closed-loop alerting method of claim 3, wherein, In step S3, within a dynamic time window after the initial warning signal is output, the pedestrian's feedback behavior is detected to determine whether the pedestrian's attention has been successfully captured. In this step, the dynamic time window has a monotonically decreasing negative correlation with the collision time TTC. Specifically, when the vehicle speed is less than or equal to a first speed threshold, the dynamic time window is calculated using a linear function, and the linear function is expressed as: wherein denotes a dynamic time window, denotes a collision time, , denotes a coefficient; When the vehicle speed exceeds the first speed threshold, the dynamic time window is calculated using an exponential function, which is expressed as: wherein , denotes a coefficient; The duration of the dynamic time window is limited to between the preset minimum and maximum window duration.
5. The attention-based closed-loop alerting method of claim 4, wherein, In step S3, within a dynamic time window after the first warning signal is output, the pedestrian's feedback behavior is detected to determine whether the pedestrian's attention has been successfully captured. Within the dynamic time window, the pedestrian's head deflection angle is continuously collected, and the change in head deflection angle relative to the preset deflection angle benchmark is calculated. If the change in head deflection angle is greater than or equal to the fourth preset angle threshold, it is determined that the attention has been successfully captured. Alternatively, within the dynamic time window, the pedestrian's gait and walking speed are continuously collected, and the change in walking speed relative to the preset walking speed benchmark is calculated. If the change in walking speed is greater than or equal to the preset percentage, it is determined that the attention has been successfully captured. If the change in head deflection angle is less than the fifth preset angle threshold and the pedestrian's gait remains unchanged, it is determined that the pedestrian's attention has not been successfully captured.
6. The closed-loop warning method based on pedestrian attention according to claim 5, characterized in that, In step S4, the intervention strategy type is determined according to a preset priority rule based on attention state, target characteristics, and environmental characteristics. The determination of the intervention strategy type is performed in the following priority order: First priority: if the attention state is distracted and the change in the pedestrian's head deflection angle is less than the fifth preset angle threshold, then it is judged as a distracted behavior type. The second priority is that if the pedestrian's age exceeds the age threshold and the main frequency of the first warning signal is higher than the preset frequency threshold and has been continuously sounding for more than the preset duration threshold, then it is determined to be a high-frequency insensitive response type. The third priority is that if the energy of the environmental noise in the preset frequency band near the main frequency of the first warning signal exceeds the energy ratio threshold of the first warning signal in the preset frequency band, it is determined to be an environmental masking type. Fourth priority: If none of the above conditions are met, it is determined to be the default type.
7. The closed-loop warning method based on pedestrian attention according to claim 5, characterized in that, Step S5, which involves adaptively selecting and executing a secondary reminder from a set of preset cross-modal reminder strategies based on the obtained intervention strategy type, includes: For distraction-related behaviors, play low-frequency pulse sounds within the first frequency range and simultaneously activate the visual alert module to flash lights; For high-frequency insensitive response types, reduce the main frequency of the secondary reminder tone to the second frequency range, shorten the pulse interval, and turn off the additional lights; For environmental masking type, the overall volume is increased by the first volume increment based on frequency domain masking compensation, and a rapid pulse mode is adopted; For the default type, the volume is increased based on the first warning signal and by the first volume increment, maintaining the tone of the first warning signal.
8. The closed-loop warning method based on pedestrian attention according to claim 7, characterized in that, In step S1, the process of perceiving pedestrians around the vehicle and identifying their attention status involves simultaneously collecting visual images, radar point clouds, environmental noise, and illumination data around the vehicle as perception inputs to achieve the perception of pedestrians around the vehicle. Specifically, the fusion weights of the visual images and radar point clouds are dynamically adjusted based on the illumination data, and the partial recognition confidence based on the visual images is corrected according to the pedestrians' attention status. Step S5 also includes: After the second reminder is completed, step S3 is executed again to detect the pedestrian's attention. If the pedestrian's attention is still not captured, the event is marked as a high-risk unresponsive sample, and the complete record is uploaded to the cloud.
9. A closed-loop warning system for the pedestrian attention-based closed-loop warning method according to any one of claims 1 to 8, characterized in that, include: The perception module is used to collect visual images, radar point clouds, environmental noise, and lighting data around the vehicle. The processing module is used to perceive pedestrians around the vehicle and identify their attention state, which is one of the following: focused, distracted, hesitant, or preparing to avoid a collision. Based on the attention state and the risk of physical collision, it determines an initial warning strategy and outputs an initial warning signal. Within a dynamic time window after outputting the initial warning signal, it detects the pedestrian's feedback behavior to determine whether the pedestrian's attention has been successfully captured. If it is determined that the pedestrian's attention has not been successfully captured, it determines an intervention strategy type according to a preset priority rule based on the attention state, target characteristics, and environmental characteristics. The intervention strategy type is one of the following: distracted behavior type, high-frequency insensitive response type, environmental masking type, or default type. Based on the obtained intervention strategy type, it adaptively selects and executes a secondary warning from a variety of preset cross-modal reminder strategies. The cross-modal reminder strategies corresponding to different intervention strategy types use at least one of the following methods: acoustic and auxiliary sensory modalities. The execution module includes a vehicle speaker and a lighting unit that support independent amplitude adjustment, and is used to output corresponding warning signals according to the control instructions of the processing module; The communication and storage module is used to cyclically store runtime data and support data interaction with the cloud.
10. An electronic device, characterized in that, The method includes a memory and a processor, the memory storing a computer program, characterized in that the processor executes the computer program to implement the steps of the closed-loop warning method based on pedestrian attention as described in any one of claims 1 to 8.