Safety pre-control method based on electronic rearview mirror data
By recognizing the state of the windshield of the vehicle behind through electronic rearview mirror images and combining data from multiple sources of sensors to construct a fogging risk assessment model, the passive response problem of traditional vehicle defogging systems is solved. This enables proactive risk prediction and dynamic control, improving vehicle safety and comfort.
Patent Information
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- CHONGQING CHANGAN AUTOMOBILE CO LTD
- Filing Date
- 2026-02-13
- Publication Date
- 2026-04-28
AI Technical Summary
Traditional vehicle defogging systems rely on a passive response mode, which cannot achieve proactive risk intervention, resulting in poor safety control strategies. Furthermore, the lack of multi-system coordination mechanisms leads to conflicts between safety and comfort, energy efficiency, and difficulty in balancing perception efficiency and real-time performance.
By collecting images from electronic rearview mirrors and identifying the status of the windshields of vehicles behind, and combining this with data from multiple sensors, a fogging risk assessment model is constructed. This enables proactive risk prediction and automatic generation of safety strategies, dynamically adjusting the air conditioning system and blind spot warning thresholds to improve vehicle safety and comfort.
It achieves accurate prediction and dynamic control of fogging risk, reduces the risk of obstructed vision, improves driving safety and intelligent experience, and optimizes the energy efficiency of the air conditioning system and the accuracy of blind spot warning.
Smart Images

Figure CN121929062A_ABST
Abstract
Description
Technical Field
[0001] This invention relates to the field of automotive technology, and more specifically to a safety pre-control method based on electronic rearview mirror data. Background Technology
[0002] For defogging on vehicle windshields and windows, the relevant technologies mainly rely on independent sensors to collect data on the internal and external environment of the vehicle. For example, temperature and humidity sensors monitor the microclimate inside the cabin, and rain / light sensors capture rainfall and light intensity. This results in fogging control being in a "passive response" mode for a long time (i.e., defogging is triggered only after fog has formed), making it impossible to intervene in advance to prevent risks. Consequently, the application of safety control strategies such as air conditioning defogging and blind spot warning is less effective. Summary of the Invention
[0003] This invention provides a safety pre-control method based on electronic rearview mirror data to solve the problem that fogging control has long been in a passive response mode, resulting in poor application effect of safety control strategies.
[0004] In a first aspect, the present invention provides a safety pre-control method based on electronic rearview mirror data, the method comprising: acquiring electronic rearview mirror images; extracting the following vehicle from the electronic rearview mirror images and identifying the windshield status of the following vehicle; determining the fogging risk of the vehicle based on the windshield status of the following vehicle; and generating a vehicle safety control strategy based on the fogging risk.
[0005] Based on the aforementioned technical methods, by extracting the key external environmental signal—the state of the rear windshield—the ambient humidity level can be indirectly and accurately determined without relying on complex external humidity sensors, solving the industry pain point that traditional systems cannot obtain real-time external humidity. This upgrades the electronic rearview mirror from a simple "image display device" to an "active sensing terminal," enabling proactive prediction of fogging risks and the generation of safety strategies. This significantly reduces driving risks caused by obstructed vision, substantially improving vehicle safety and the intelligent driving experience.
[0006] In some alternative implementations, the fogging risk of the vehicle is determined by the state of the rear vehicle's windshield, including: creating multimodal input data for a model based on the state of the rear vehicle's windshield; inputting the multimodal input data into a fogging risk assessment model; and outputting the vehicle's fogging risk level.
[0007] Based on the aforementioned technical means, a multi-source data fusion and fogging risk assessment model is further introduced to achieve quantitative output of risk levels, improving the accuracy and reliability of control strategies. By collecting data from other sensors and combining it with electronic rearview mirror images for weather recognition, the limitations of a single data dimension are overcome, making fogging risk assessment more comprehensive. This considers both the external environment reflected by the rear vehicle's glass and the microclimate inside and outside the vehicle captured by sensor data. The application of the fogging risk assessment model transforms the vague "risk perception" into a clear "level output," providing a quantitative basis for subsequent safety control strategies and avoiding the "one-size-fits-all" control logic of traditional systems.
[0008] In some optional implementations, the multimodal input data is input into the fogging risk assessment model, and the fogging risk level of the vehicle is output. This includes: performing feature preprocessing on the multimodal input data to obtain multimodal features; fusing the multimodal features to obtain fused features; performing bidirectional LSTM processing on the fused features to output state features; performing multi-head attention processing on the state features to obtain attention output information; and mapping the attention output information to a risk level through a fully connected layer to obtain the fogging risk level of the vehicle.
[0009] Based on the aforementioned technical methods, an efficient and accurate fog risk assessment system was constructed, solving the problems of low perception accuracy and slow response of traditional models. Multimodal input data covers key environmental parameters inside and outside the vehicle. Feature preprocessing and fusion steps achieve data denoising and information enhancement. The bidirectional LSTM module fully explores temporal change trends, and the multi-head attention mechanism accurately focuses on core influencing factors, progressively improving the model's perception capabilities. This enables the model to output fog risk levels in advance, improving warning capabilities and reducing false alarm rates compared to traditional threshold methods.
[0010] In some alternative implementations, the multimodal input data for creating a model based on the state of the rear vehicle's windshield includes: calculating the temperature difference between the vehicle's interior and exterior using other sensor data; estimating a humidity proxy index based on the state of the rear vehicle's windshield; identifying the current weather using electronic rearview mirror images and other sensor data; encoding the current weather and adding an enhancement score calculated using other sensor data to the weather code; calculating a consistency coefficient to characterize the consistency of fogging phenomena of rear vehicles based on the state of the rear vehicle's windshield; and extracting historical sequences of other sensor data. The multimodal input data includes the temperature difference between the vehicle's interior and exterior, the humidity proxy index, the weather code, the consistency coefficient, and the historical sequences.
[0011] Based on the aforementioned technical means, by clarifying the specific composition and creation methods of multimodal input data, a standardized and high-quality input source is provided for the fogging risk assessment model, ensuring the accuracy and stability of the model's predictions. The temperature difference between the vehicle's interior and exterior is considered the core physical factor for fogging; humidity proxy indicators address the challenge of directly obtaining external humidity; weather coding incorporates fine-grained weather information; consistency coefficients reflect the environmental correlation of fogging behind vehicles; and historical sequences capture parameter change trends. These five types of data comprehensively cover the factors influencing fogging, enabling the model to fully utilize various environmental information for comprehensive judgment.
[0012] In some alternative implementations, a humidity proxy index is estimated based on the condition of the rear windshield, including:
[0013]
[0014] In the formula, Indicates humidity proxy index, This represents the distance decay weight of the i-th distance. , This vehicle represents the distance to the i-th vehicle following it. For the preset distance threshold, This indicates the total number of vehicles inspected. This indicates the degree of fogging on the windshield of the i-th vehicle following it.
[0015] Based on the aforementioned technical methods, the condition of the rear vehicle's windshield is transformed into a quantifiable environmental humidity parameter, solving the key pain point of traditional systems' inability to acquire real-time external humidity. By introducing distance attenuation weights, the contribution of the windshield condition of nearby rear vehicles to humidity assessment is increased, improving the accuracy and reliability of the indicator. Through weighted summation and averaging of the fogging degree of rear vehicles, interference caused by abnormal windshield conditions of individual vehicles (such as artificial window tinting or dirt) is effectively filtered out, enhancing the robustness of the indicator. This humidity proxy indicator can accurately reflect the external environmental humidity level, enabling the fogging risk assessment model to predict fogging in advance, providing crucial data support for preventative defogging, significantly reducing the risk of obstructed visibility due to fogging, and improving driving safety.
[0016] In some optional implementations, a consistency coefficient is calculated based on the condition of the rear vehicle's windshield to characterize the consistency of fogging phenomena in following vehicles, including:
[0017] In the formula, Represents the consistency coefficient. This indicates the number of vehicles whose rear end is fogged up. This indicates the total number of vehicles inspected. The standard deviation represents the degree of fogging on the vehicle behind.
[0018] Based on the aforementioned technical methods, the consistency coefficient calculation method can accurately quantify the synergy of fogging phenomena among vehicles behind, providing a core basis for judging the stability and intensity of external environmental humidity, and further improving the accuracy of fogging risk assessment. This coefficient comprehensively considers both the proportion of fogging vehicles and the standard deviation of fogging severity, reflecting both the prevalence and consistency of the fogging phenomenon. A high consistency coefficient indicates a stable high humidity state in the external environment, with an extremely high risk of fogging, while a low consistency coefficient suggests unstable environmental humidity and a lower risk. It can effectively distinguish between "abnormal fogging of a single vehicle" and "generalized environmental fogging," avoiding model misjudgment and making the fogging risk level output more closely match the actual scenario.
[0019] In some optional implementations, feature preprocessing is performed on the multimodal input data to obtain multimodal features, including: standardizing the temperature difference between inside and outside the vehicle, humidity proxy index, consistency coefficient, and historical sequence; performing one-dimensional convolution on the standardized multimodal data to obtain the first multimodal feature; assigning weight information to each feature in the first multimodal feature through a spatial attention mechanism; and performing temporal pooling on the first multimodal feature with assigned weight information to obtain the second multimodal feature.
[0020] Based on the above technical means, a four-step preprocessing process of standardization, one-dimensional convolution, spatial attention weighting, and temporal pooling is used to achieve noise reduction, feature enhancement, and simplification of multimodal input data, providing high-quality data support for subsequent feature fusion and model inference, and solving the problem of poor model performance caused by messy and redundant original data.
[0021] In some optional implementations, feature fusion of multimodal features is performed to obtain fused features, including: concatenating the second multimodal feature and the weighted weather code to obtain concatenated features; processing the concatenated features through a fully connected layer, and randomly discarding some feature channels during the fully connected layer processing to obtain higher-order features; and normalizing the higher-order features to obtain fused features, wherein the magnitude of the normalization is positively correlated with the consistency coefficient.
[0022] Based on the aforementioned technical methods, a feature fusion process involving feature splicing, fully connected layer processing, and normalization is employed to achieve deep interaction and optimization of multimodal features. This solves the problems of information fragmentation and redundancy in traditional feature fusion, thereby enhancing the expressive power of the features. Before splicing, attention weighting is applied to the weather code to ensure a high degree of adaptation between weather information and the current environment. The fully connected layer extracts high-order feature combinations, and a feature discarding mechanism enhances the model's robustness. Normalization dynamically adjusts the amplitude based on the consistency coefficient, preserving environmental change information when consistency is high and suppressing noise when consistency is low. This fusion scheme achieves the organic integration of multimodal information, enabling the fused features to contain both the core information of various data types and highlight key influencing factors, significantly improving the perception capability of the fog risk assessment model.
[0023] In some optional implementations, the fused features are processed by a bidirectional LSTM to output state features, including: obtaining a first LSTM module and a second LSTM module; processing the fused features in chronological order through the first LSTM module and processing the fused features in reverse chronological order through the second LSTM module, wherein the input gates of the first LSTM module and the second LSTM module are processed as follows:
[0024] In the formula, This represents the input gate value for time step t. This represents the activation function. This represents the weight matrix of the input gate. This means concatenating the hidden state from the previous time step with the input feature vector of the current time step. This represents the bias term of the input gate. The weather impact coefficient represents the magnitude of the weather impact, adjusted according to the current weather conditions. The weight matrix representing weather codes. Indicates weather codes; The hidden states output by the first LSTM module and the second LSTM module are concatenated and linearly transformed to obtain the state features.
[0025] Based on the aforementioned technical methods, by designing a bidirectional LSTM module and optimizing the input gate, the temporal variation patterns of multimodal fusion features are fully explored, improving the foresight and accuracy of fog risk assessment and solving the problems of unidirectional dependence and poor weather adaptability in traditional time series models. Forward LSTM and backward LSTM capture parameter change trends from the forward and backward directions respectively, and state fusion yields a complete temporal context. The input gate introduces a weather condition gating mechanism, dynamically adjusting the information inflow weights according to weather type, enhancing the influence of weather information under severe weather conditions such as rain and fog, and weakening it under clear weather, making the model more closely resemble real-world scenarios. This enables the model to accurately capture the temporal evolution patterns of fog risk, predict risks in advance, and significantly improve early warning capabilities compared to traditional models.
[0026] In some optional implementations, multi-head attention processing is performed on the state features to obtain attention output information, including: setting a first attention head, a second attention head, a third attention head, and a fourth attention head; the first, second, and third attention heads are used to extract attention information of the same feature in the state features at different time lengths; the fourth attention head is used to capture attention information between different features in the state features; and the query vector of each attention head is adjusted by querying a weight matrix, which is calculated by the following formula:
[0027] In the formula, This represents the query weight matrix. This represents the original query weight in the standard attention mechanism. This represents the modulation intensity coefficient, with a value ranging from 0 to 1, and its magnitude is adjusted according to the current weather conditions. The weight matrix representing weather codes. Indicates weather codes; The state features are processed by multi-head attention using each attention head, and the attention sub-information output by each attention head is concatenated and linearly transformed to obtain the first attention output information. The first attention output information is then residually connected with the state features output after LSTM processing to obtain the second attention output information.
[0028] Based on the aforementioned technical methods, a multi-head attention mechanism is employed to achieve precise focusing and in-depth mining of state features, addressing the shortcomings of traditional models in paying insufficient attention to interactions across different time scales and features, thus improving the accuracy of fog risk assessment. The four attention heads have clearly defined roles, focusing on short-term mutations, medium-term trends, long-term patterns, and cross-feature interactions, comprehensively covering various dimensions of fog risk impact. The query weight matrix incorporates weather coding modulation, dynamically adjusting the attention distribution according to weather type, and strengthening the focus on key features under severe weather conditions. This enables the model to accurately identify the core influencing factors and changing patterns of fog risk, effectively filtering out interfering information and improving the model's anti-interference capability and prediction accuracy. Simultaneously, residual connections preserve original state feature information, avoiding information loss during feature fusion and further enhancing model performance.
[0029] In some optional implementations, the following steps are taken: extracting the following vehicle from the electronic rearview mirror image and identifying the windshield status of the following vehicle; collecting other sensor data; and identifying the current weather using the electronic rearview mirror image and other sensor data. This includes: extracting a first intermediate feature from the electronic rearview mirror image using an embedding module and a visual Transformer module; extracting a second intermediate feature from the electronic rearview mirror image using a convolutional neural network module and a concatenated attention module; performing basic convolution processing, convolutional layer processing, and a fully connected layer on the second intermediate feature to obtain a third intermediate feature, and classifying the third intermediate feature to obtain the current weather; fusing the first and second intermediate features to obtain a fourth intermediate feature; inputting the fourth intermediate feature into the parallel attention module to identify the following vehicle in the electronic rearview mirror image; and decoding the following vehicle in the electronic rearview mirror image to obtain the windshield status of the following vehicle.
[0030] Based on the aforementioned technical methods, an efficient electronic rearview mirror image feature extraction and multi-task recognition scheme is provided, simultaneously achieving rear vehicle recognition, rear vehicle glass status recognition, and weather recognition. This solves the problems of traditional systems having single perception functions and low multi-task processing efficiency. Through bi-branch feature extraction, global correlation features and local detail features are obtained separately, and then fused to achieve feature complementarity. For the feature requirements of different tasks, weather classification is achieved through basic convolution and fully connected layers, rear vehicle recognition is achieved through a parallel attention module, and glass status segmentation is achieved through a decoder, resulting in efficient multi-task collaboration. This model enables the electronic rearview mirror to possess "environmental perception" capabilities, achieving a qualitative leap from "passive display" to "active perception" compared to the traditional single display function. It can simultaneously output three types of key environmental information, providing comprehensive data support for fog risk assessment and safety control strategies.
[0031] In some optional implementations, the concatenated attention module includes: a first module connected to the input and a second module connected to the first module; the parallel attention module includes a first module and a second module respectively connected to the input, and the outputs of the first module and the second module are superimposed; the first module includes a first branch and a second branch, the first branch is used to sequentially input the input data into a global average pooling layer and a bottleneck layer, the second branch is used to sequentially input the input data into a global max pooling layer and a basic convolutional layer, and the output of the first module is the superposition of the outputs of the first branch and the second branch; the second module includes a third branch and a fourth branch, the third branch is used to sequentially input the input data into a global average pooling layer and a convolutional layer, the fourth branch is used to input the input data into a convolutional layer, and the output of the second module is the superposition of the outputs of the third branch and the fourth branch.
[0032] Based on the aforementioned technical methods, the specific structures of the Serial Attention Module (CSAM) and the Parallel Attention Module (CPAM) provide efficient attention mechanisms for feature extraction and multi-task recognition of electronic rearview mirror images, solving the problems of global and local information imbalance and weak anti-interference ability in traditional feature extraction. CSAM adopts a "global-first, local-later" serial structure, optimizing features step by step, adapting to tasks such as weather classification that require defining the scene before refinement; CPAM adopts a "global and local parallel" parallel structure, with complementary information from multiple paths, adapting to tasks such as rear vehicle recognition that require precise positioning. Both modules strengthen core features and suppress interference information through operations such as global pooling, convolution, and feature fusion. This enables the feature extraction process to accurately adapt to the needs of different recognition tasks, improving the accuracy of rear vehicle recognition, glass state segmentation, and weather classification, and providing high-quality perception data for subsequent fogging risk assessment and safety control strategies.
[0033] In some optional implementations, a vehicle safety control strategy is generated based on the fogging risk, including: defining a risk-scenario dynamic mapping relationship for controlling vehicle air conditioning system parameters, whereby a higher fogging risk level results in faster defogging of the adjusted air conditioning parameters, and a lower fogging risk level results in air conditioning parameters that better match the user's comfort index; defining a control table based on the risk-scenario dynamic mapping relationship, whereby the control table is used to assign corresponding air conditioning system control parameters to each fogging risk level; acquiring environmental perception data and vehicle status data, whereby the environmental perception data includes the fogging risk level, the windshield status of following vehicles, and the current weather, and the vehicle status data includes the current air conditioning system parameters and driving scenario information; and matching the acquired environmental perception data and vehicle status data with the control table to obtain the current target air conditioning system control parameters for adjusting the current air conditioning parameters.
[0034] Based on the aforementioned technical methods, by establishing a dynamic risk-scenario mapping relationship and control table, precise and dynamic adjustment of air conditioning system parameters is achieved, solving the problems of traditional defogging systems' single control strategy and imbalance between safety and comfort. The dynamic risk-scenario mapping relationship clearly defines the core logic: "the higher the risk, the higher the defogging priority; the lower the risk, the higher the comfort priority." The control table assigns targeted air conditioning parameters to different risk levels, ensuring the standardization and operability of the control strategy. By matching environmental perception data and vehicle status data, real-time dynamic adjustment of air conditioning parameters is achieved. Emergency defogging mode is triggered in high-risk scenarios, safety and comfort are balanced in medium-risk scenarios, and comfort is prioritized in low-risk scenarios. This significantly improves driving safety and comfort, achieving an organic unity of safety and energy efficiency.
[0035] In some optional implementations, generating a vehicle safety control strategy based on the fogging risk also includes: monitoring the defogging effect inside the vehicle, adjusting the current corresponding fogging risk level based on the defogging effect, and then returning to the step of adjusting the current air conditioning parameters according to each risk scenario level based on the risk-scenario dynamic mapping relationship.
[0036] Based on the aforementioned technical means, by introducing a closed-loop mechanism for monitoring defogging effect and dynamically adjusting risk level, the problem of insufficient control precision caused by the "one-time adjustment without feedback optimization" of traditional control strategies is solved, achieving continuous optimization and dynamic adaptation of the safety control strategy. Through in-vehicle temperature and humidity sensors and visual feedback from the electronic rearview mirror, the defogging effect is monitored in real time. If fogging does not improve, the risk level and control intensity are automatically upgraded; if the risk decreases, the control strategy is adjusted accordingly. This ensures that the control strategy can accurately adapt to environmental changes, avoiding obstructed vision due to sudden environmental changes or energy waste due to reduced risk.
[0037] In some optional implementations, the vehicle safety control strategy generated based on the fogging risk also includes: recording the driver's manual intervention behavior on the air conditioning parameters; generating adjustment weights based on the manual intervention behavior; using the adjustment weights to correct the air conditioning system control parameters; and adjusting the contrast and sharpness of the electronic rearview mirror image according to the visibility of the current weather.
[0038] Based on the aforementioned technical methods, by recording the driver's manual intervention behavior, generating adjustment weights, and correcting the air conditioning control parameters, personalized adaptation of the control strategy is achieved. Simultaneously, the display effect of the electronic rearview mirror is optimized, solving the problems of traditional systems being "uniform" and display effects being affected by weather. A reinforcement learning mechanism records the driver's preferred operations and optimizes the strategy weights based on the intervention effect, making subsequent control strategies more aligned with the driver's habits. Furthermore, the contrast and sharpness of the electronic rearview mirror image are automatically adjusted according to weather visibility to ensure a clear rear view and assist in blind spot judgment, further enhancing the system's practical value and user experience.
[0039] In some optional implementations, the vehicle safety control strategy generated based on the fogging risk further includes: calculating the relative distance and relative speed between the vehicle and the vehicle behind based on the vehicle behind in the electronic rearview mirror image; adjusting the blind spot warning threshold according to the fogging risk level; and outputting warning information, steering intervention suggestions, and / or braking prompts when the vehicle behind is detected to enter the dynamic blind spot corresponding to the blind spot warning threshold based on the relative distance and relative speed.
[0040] Based on the aforementioned technical means, by integrating the following vehicle's motion status and fog risk level, the blind spot warning threshold is dynamically adjusted, achieving synergistic linkage between the blind spot warning system and fog risk control. This overcomes the limitations of traditional blind spot monitoring systems, which have "fixed thresholds and do not consider the impact of fog." By calculating the relative distance and speed between the vehicle and the following vehicle, and combining this with the fog risk level, the warning threshold is dynamically adjusted. When the fog risk is high, the threshold is shortened and sensitivity is increased to ensure timely warnings. Simultaneously, when the following vehicle enters the dynamic blind spot, multi-dimensional warnings and intervention suggestions are provided, improving the accuracy of lane change scenario risk identification. By organically combining blind spot warning and fog risk control, the problem of "blind spots within blind spots" is solved, and the accuracy and timeliness of blind spot warnings are improved, making lane change decisions safer and more reliable.
[0041] In some alternative implementations, the method further includes: when raindrops or fog are detected in the electronic rearview mirror, enabling an image enhancement algorithm to process the electronic rearview mirror image.
[0042] Based on the above technical means, by using image enhancement algorithms to process raindrop or fog images of electronic rearview mirrors, the problem of reduced environmental perception caused by blurred vision of electronic rearview mirrors in severe weather is solved, providing clear visual data support for fog risk assessment and blind spot warning.
[0043] Secondly, the present invention provides a safety pre-control device based on electronic rearview mirror data. The device includes: a data acquisition module for acquiring electronic rearview mirror images; a windshield status recognition module for extracting the following vehicle from the electronic rearview mirror image and recognizing the windshield status of the following vehicle; a fogging risk determination module for determining the fogging risk of the vehicle based on the windshield status of the following vehicle; and a strategy control module for generating a vehicle safety control strategy based on the fogging risk.
[0044] Thirdly, the present invention provides a vehicle comprising: a memory and a processor, the memory and the processor being communicatively connected to each other, the memory storing computer instructions, and the processor executing the computer instructions to perform the method described in the first aspect or any corresponding embodiment thereof.
[0045] Fourthly, the present invention provides a computer-readable storage medium storing computer instructions for causing a computer to perform the method described in the first aspect or any corresponding embodiment thereof. Attached Figure Description
[0046] To more clearly illustrate the specific embodiments of the present invention or the technical solutions in the prior art, the drawings used in the description of the specific embodiments or the prior art will be briefly introduced below. Obviously, the drawings described below are some embodiments of the present invention. For those skilled in the art, other drawings can be obtained from these drawings without creative effort.
[0047] Figure 1 This is a flowchart illustrating a safety pre-control method based on electronic rearview mirror data according to an embodiment of the present invention. Figure 2 This is a schematic diagram of the structure of a fogging risk assessment model according to an embodiment of the present invention; Figure 3 This is a schematic diagram of the structure of a multi-task model according to an embodiment of the present invention; Figure 4 This is a schematic diagram illustrating the effect of weather classification according to an embodiment of the present invention; Figure 5 This is a schematic diagram illustrating the effect of vehicle recognition according to an embodiment of the present invention; Figure 6 This is a schematic diagram illustrating the effect of rear windshield segmentation according to an embodiment of the present invention; Figure 7 These are schematic diagrams of the serial attention module and the parallel attention module according to embodiments of the present invention; Figure 8 This is another flowchart illustrating a safety pre-control method based on electronic rearview mirror data according to an embodiment of the present invention; Figure 9 This is a schematic diagram of a safety pre-control device based on electronic rearview mirror data according to an embodiment of the present invention; Figure 10 This is a schematic diagram of the hardware structure of a vehicle according to an embodiment of the present invention. Detailed Implementation
[0048] To make the objectives, technical solutions, and advantages of the embodiments of the present invention clearer, the technical solutions of the embodiments of the present invention will be clearly and completely described below with reference to the accompanying drawings. Obviously, the described embodiments are only some embodiments of the present invention, not all embodiments. Based on the embodiments of the present invention, all other embodiments obtained by those skilled in the art without creative effort are within the scope of protection of the present invention.
[0049] It is understood that before using the technical solutions disclosed in the various embodiments of the present invention, users should be informed of the types, scope of use, and usage scenarios of the personal information involved in the present invention and their authorization should be obtained in accordance with relevant laws and regulations through appropriate means.
[0050] The terms "first" and "second" are used for descriptive purposes only and should not be construed as indicating or implying relative importance or implicitly specifying the number of technical features indicated. Thus, a feature defined as "first" or "second" may explicitly or implicitly include one or more of that feature. In the description of this invention, "a plurality of" means two or more, unless otherwise explicitly specified.
[0051] With the rapid development of automotive intelligence and connectivity, electronic rearview mirrors (CMS) are gradually replacing traditional optical rearview mirrors, becoming a core component for vehicle environmental perception due to their wide viewing angle, anti-obstruction properties, and digital processing advantages. Simultaneously, consumers' increasing demands for driving safety, ride comfort, and energy efficiency are driving the upgrading of functions such as vehicle defogging systems and blind spot detection (BSD) systems towards intelligence and collaboration. However, existing technologies still have many critical pain points that urgently need to be addressed, severely hindering further improvements in the intelligent cockpit experience and driving safety.
[0052] First, environmental perception systems suffer from "data silos," with limited and fragmented perception dimensions. Traditional vehicle fogging risk assessment relies primarily on in-vehicle temperature and humidity sensors, which can only monitor the cabin microclimate and cannot acquire crucial information such as external environmental humidity in real time. This results in the defogging system operating in a passive mode of "responding only after fog forms," leading to high response delays and hindering preventative intervention. Simultaneously, blind spot monitoring systems often rely on millimeter-wave radar or single cameras, which can only identify vehicle position and cannot assess the visibility of drivers behind (e.g., fogging of the rear windshield obstructing the driver's view), creating "blind spots within blind spots." According to a 2023 SAE research report, approximately 29% of lane-changing accidents stem from delayed reactions caused by such obstructed visibility.
[0053] Secondly, the lack of a multi-system coordination mechanism leads to conflicts between safety, comfort, and energy efficiency. Functions such as air conditioning defogging, blind spot warning, and electronic rearview mirror display belong to different control domains and lack a unified collaborative decision-making logic. For example, in rainy weather, defogging requires maximum airflow, and the fan noise can easily mask the BSD warning sound; while strong cold air blowing directly on the defogging system can quickly dehumidify, it can cause a sudden drop in the interior temperature, affecting driving comfort; and if the driver manually adjusts the temperature, it may conversely increase the risk of fogging. In addition, traditional control strategies use fixed thresholds and rules, which cannot be dynamically adjusted according to weather, vehicle speed, and other scenarios, resulting in over-defogging (energy waste) or under-defogging (obstructed visibility) in some situations.
[0054] Finally, in automotive scenarios, it is difficult to balance perception efficiency and real-time performance. To achieve full-scene perception, using multiple independent models to handle tasks such as vehicle detection, weather recognition, and glass state segmentation would significantly increase the load on the onboard computing unit, failing to meet millisecond-level control requirements. Furthermore, in related technologies, single-task models exhibit significant performance degradation in adverse weather conditions, making them unsuitable for complex driving environments.
[0055] In summary, how to fully tap the perception potential of electronic rearview mirrors, establish a closed loop connecting "external environment perception, in-vehicle risk prediction, and multi-system collaborative control," and achieve early warning of fogging risks, accurate identification of blind spot risks, and dynamic coordination of various systems, while meeting the lightweight and real-time requirements of in-vehicle scenarios, has become a core technological bottleneck that urgently needs to be overcome in the current development of intelligent vehicle technology.
[0056] According to an embodiment of the present invention, a safety pre-control method based on electronic rearview mirror data is provided. It should be noted that the steps shown in the flowchart in the accompanying drawings can be executed in a computer system such as a set of computer-executable instructions. Furthermore, although a logical order is shown in the flowchart, in some cases, the steps shown or described may be executed in a different order than that shown here.
[0057] This embodiment provides a safety pre-control method based on electronic rearview mirror data, which can be used in vehicles. Figure 1 This is a flowchart of a safety pre-control method based on electronic rearview mirror data according to an embodiment of the present invention. The process includes the following steps: Step S101: Acquire images from the electronic rearview mirror; Step S102: Extract the following vehicle from the electronic rearview mirror image and identify the windshield status of the following vehicle; Step S103: Determine the risk of fogging of this vehicle by checking the condition of the rear windshield; Step S104: Generate a vehicle safety control strategy based on the risk of fogging.
[0058] Specifically, this embodiment provides a safety pre-control method based on electronic rearview mirror data, which can be directly applied to various vehicles equipped with electronic rearview mirrors, especially intelligent connected vehicles. It can realize the prediction of fogging risk and the automatic generation of safety strategies through the environmental perception capability of electronic rearview mirrors.
[0059] As a core visual perception device for vehicles, electronic rearview mirrors typically include three cameras: one on the left, one on the right, and one at the rear. This provides a complete field of view covering the sides and rear of the vehicle. The images they capture generally have a resolution of 1280×720 pixels and a frame rate of 30fps, sufficient for real-time environmental perception. Once the vehicle is started and a driving gear (D, R, etc.) is engaged, the electronic rearview mirror system automatically activates and continuously captures images of the surrounding environment. This acquisition process is unaffected by weather conditions; even in adverse weather conditions such as rain, snow, and fog, it can still output stable image data. For example, when driving in the rain, the electronic rearview mirror camera uses a built-in waterproof coating to reduce raindrop adhesion and automatically adjusts exposure parameters to ensure that the captured images clearly show vehicles and the road environment to the sides and rear, providing high-quality raw data for subsequent rear vehicle recognition and glass condition assessment.
[0060] In this embodiment of the invention, the acquired electronic rearview mirror image is first subjected to target detection using an image recognition algorithm (such as the YOLO series algorithm). This algorithm can automatically scan features such as vehicle outlines, headlights, and windows in the image, thereby accurately locating all vehicles in the side and rear view and eliminating interference from non-vehicle targets such as pedestrians, road signs, and buildings. For example, in a congested urban traffic scenario, the algorithm can accurately identify three following vehicles within a 10-meter range behind the vehicle from a complex image background and mark the position coordinates and outline range of each vehicle. Subsequently, for each identified following vehicle, its windshield area is accurately extracted using a semantic segmentation algorithm. The semantic segmentation algorithm can distinguish different parts of the vehicle (body, glass, tires, etc.) at the pixel level, ensuring that only the windshield, a key area, is focused on. Finally, based on the image features of the glass area (such as clarity, grayscale distribution, and the presence of blurred areas), its state is identified, mainly including two core states: "clear" and "foggy / frosty". The "clear" state means that the glass is unobstructed and the interior of the vehicle can be clearly seen; the "foggy / frosty" state means that fog or frost forms on the glass surface due to temperature and humidity differences, resulting in a blurred image and reduced uniformity of grayscale values. For example, if the overall grayscale value of the windshield area of a vehicle behind is high and there is no obvious view of the interior, it can be determined to be in the "foggy" state.
[0061] The core principle behind windshield fogging is the temperature and humidity difference between the inside and outside of the vehicle. When the external humidity is high, the warm, humid air inside the car comes into contact with the cooler glass surface, causing water vapor to condense and form fog. The condition of the windshields of vehicles behind directly reflects the humidity level of the external environment. If the windshields of multiple vehicles behind are simultaneously fogged up, it indicates that the current external humidity is extremely high, and your vehicle, facing the same temperature and humidity difference, has a significantly increased risk of fogging. If the windshields of all vehicles behind are clear, it indicates that the external humidity is low, and your vehicle has a lower risk of fogging. For example, in the late afternoon after rain, if the electronic rearview mirror detects that four out of five vehicles behind have fogged windshields, and considering the difference between the vehicle's interior temperature of 25°C and the exterior temperature of 18°C, even though the vehicle itself is not currently fogged, the high humidity environment indicates a high risk of fogging. Conversely, on a sunny midday, if all the windshields of the vehicles behind are clear, even with a significant temperature difference between the vehicle's interior and exterior, the risk of fogging can be determined to be low. This process does not rely on complex external humidity sensors; the intuitive and easily accessible information of the windshield status of the vehicles behind allows for a quick qualitative assessment of the fogging risk, enabling proactive risk detection.
[0062] The core of the safety control strategy is to automatically adjust the operating status of relevant vehicle systems based on the fogging risk level to prevent or eliminate the impact of fog on visibility and ensure driving safety. When the fogging risk is determined to be high, the system automatically activates the air conditioning defogger, adjusts the air conditioning vents to blow towards the windshield, increases the airflow, and appropriately raises the air conditioning temperature. This uses dry airflow to quickly reduce the humidity on the glass surface and prevent fog formation. When the fogging risk is determined to be low, the system maintains the normal operation of the air conditioning, prioritizing passenger comfort, and does not require activation of the defogger. For example, after determining a high fogging risk, the vehicle automatically increases the air conditioning fan speed from level 2 to level 4, switches the air vents to windshield mode, and adjusts the temperature from 23°C to 26°C, continuously blowing dry air to ensure the windshield remains clear. If the fogging risk is low, the air conditioning maintains the user-set temperature of 22°C, fan speed level 2, and face-blowing mode.
[0063] The safety pre-control method based on electronic rearview mirror data provided in this embodiment uses the state of the rear vehicle's windshield as an indirect basis for judging external humidity. This solves the problem of traditional vehicles being unable to obtain external humidity in real time and exhibiting passive and delayed fogging control. It achieves proactive prediction of fogging risks, avoiding the passive situation of defogging only after fog has formed. Through precise risk assessment and targeted control strategies, it ensures clear visibility while preventing ineffective operation of the air conditioning system, balancing safety and energy efficiency. This method effectively improves the vehicle's intelligence level and driving safety, providing a more reliable and convenient travel experience for drivers and passengers.
[0064] In some optional implementations, step S103 above includes: Step a1: Create multimodal input data for the model based on the state of the rear windshield; Step a2: Input the multimodal input data into the fogging risk assessment model and output the fogging risk level of the vehicle.
[0065] Specifically, the embodiments of the present invention further optimize the judgment logic of fogging risk by introducing multi-source data fusion and risk assessment model to achieve quantitative output of fogging risk, making subsequent safety control strategies more targeted.
[0066] Other sensor data refers to the output data from sensors already installed on the vehicle to monitor the internal and external environment and operating status. These sensors, together with the electronic rearview mirror, form a multi-source perception system, compensating for the limitations of single-source visual perception. Specifically, this includes two core sensor types: temperature and humidity sensors and rain sensors. Temperature and humidity sensors can detect the temperature and relative humidity inside and outside the vehicle respectively, typically updating at a frequency of 1 Hz (Hertz, i.e., once per second), enabling real-time capture of microclimate changes inside and outside the vehicle. Rain sensors are installed on the inside of the windshield and use optical principles to detect the intensity of raindrops hitting the glass, outputting a continuous value between 0 and 1 (0 indicating no rain, 1 indicating heavy rain), reflecting the current rainfall situation. This data will be combined with the image information from the electronic rearview mirror to provide a foundation for subsequent weather recognition and risk assessment.
[0067] The core of weather recognition is to combine visual features with sensor data for comprehensive judgment, avoiding misjudgments caused by a single data dimension. Electronic rearview mirror images provide intuitive visual cues. For example, on sunny days, the image is generally bright, shadows are clear, and the sky is unobstructed; on rainy days, raindrop trails are visible, and the road surface is wet and reflective; on foggy days, visibility is low, distant objects are blurred, and contrast is reduced. Other sensor data provides quantitative support for visual judgment: a rain sensor output greater than 0 confirms a rainy scene; a temperature and humidity sensor detecting extremely high humidity outside the vehicle (e.g., ≥85%) and low image visibility can help determine foggy weather; a light sensor (if equipped in the vehicle) detecting low light intensity can help distinguish between cloudy days and evening. For example, if the electronic rearview mirror image shows a wet road surface with scattered raindrop trails, a rain sensor output of 0.3, and a temperature and humidity sensor showing 82% humidity outside the vehicle, combining this information can accurately identify the current weather as light rain; if the image visibility is extremely low, there are no obvious raindrops, and the temperature and humidity sensor shows 90% humidity outside the vehicle, then it is determined to be foggy weather. In particular, the large model algorithm used to make weather judgments by integrating the above information is existing technology, and will not be described in detail in the embodiments of this invention.
[0068] The fogging risk assessment model is a prediction model built on deep learning. Its core function is to transform multi-source heterogeneous data into quantifiable risk levels, achieving an upgrade from "qualitative perception" to "quantitative assessment." Other sensor data provides core physical parameters such as the temperature and humidity difference between the vehicle's interior and exterior (key influencing factors of fogging). The status of the windshields of following vehicles reflects indirect signals of external humidity (multiple fogging vehicles indicate extremely high external humidity). The current weather clearly defines the scene context (the risk of fogging is naturally higher on rainy or foggy days than on sunny days). By learning the correlation patterns in a large amount of historical data, the model automatically weights various input information and ultimately outputs three levels of fogging risk: low, medium, and high (this example is only for illustration and is not limited to this; more levels can be added based on actual conditions). For example, if other sensor data shows a 7°C temperature difference between the vehicle's interior and exterior, three following vehicles have fogged windshields while one is clear, and the current weather is light rain, inputting this information into the model will result in a medium fogging risk level. If the temperature difference is 10°C, all following vehicles have fogged windshields, and the current weather is foggy, the model will output a high fogging risk level.
[0069] The solution provided in this embodiment overcomes the limitations of single-vision recognition by collecting data from other sensors and combining it with electronic rearview mirror images to identify weather conditions, thus improving the accuracy of weather recognition and providing reliable scenario support for fog risk assessment. Secondly, the application of the fog risk assessment model transforms vague risk perception into clear level outputs, avoiding the "one-size-fits-all" control logic of traditional systems and making subsequent safety control strategies more aligned with actual scenarios. Thirdly, the fusion of multi-source data fully utilizes existing vehicle hardware resources without the need for additional dedicated sensors, balancing technological innovation and engineering practicality. Fourthly, the quantified risk levels provide a unified decision-making basis for the coordinated control of systems such as air conditioning defogging and blind spot warning, ensuring clear visibility while reducing unnecessary energy consumption and achieving an organic unity of safety, comfort, and energy efficiency.
[0070] In some alternative implementations, step a2 above includes: Step b1: Perform feature preprocessing on the multimodal input data to obtain multimodal features; Step b2: Perform feature fusion on the multimodal features to obtain fused features; Step b3: Perform bidirectional LSTM processing on the fused features to output state features; Step b4: Perform multi-head attention processing on the state features to obtain attention output information; Step b5: The attention output information is mapped to a risk level through a fully connected layer to obtain the vehicle's fogging risk level.
[0071] Specifically, such as Figure 2The diagram shown is a structural diagram of the fog risk assessment model. This embodiment of the invention further improves the accuracy and reliability of the fog risk level output through "multimodal input creation → feature preprocessing → fusion → temporal modeling → attention focusing → risk mapping," ensuring that the risk assessment is both comprehensive and efficient.
[0072] Multimodal input data is a standardized dataset that integrates various types of environmental information. Its core components include real-time features and historical sequences, consistent with the input layer configuration. Real-time features include four key parameters: first, the temperature difference between the vehicle's interior and exterior, calculated from the temperatures of other sensors, which is the core physical factor causing fogging; second, a humidity proxy index, estimated based on the degree of fogging on the windshields of following vehicles, indirectly reflecting the external environmental humidity; third, weather coding, converting the current weather (e.g., sunny, rainy, foggy) into digital labels; and fourth, a consistency coefficient, calculated based on the windshield status of multiple following vehicles, characterizing the synergy of fogging phenomena in those vehicles. The historical sequence extracts the above four types of real-time feature data from the past five time steps (one second per step), forming a 5×4-dimensional time-series data set to capture the dynamic trends of parameter changes, comprehensively covering the current state and changing trends of fogging influencing factors.
[0073] The core of feature preprocessing is to provide high-quality data for subsequent modeling through data optimization and feature enhancement.
[0074] The purpose of feature fusion is to achieve deep interaction between real-time features and time-series features, thereby improving information expression capabilities.
[0075] Bidirectional LSTM (Long Short-Term Memory) networks are used to mine temporal dependencies in fused features and fully capture the temporal context information of the features.
[0076] Multi-head attention processing is used to accurately focus on the core influencing factors in state features. Some attention heads focus on short-term feature mutations, some focus on long-term trends, and some capture cross-feature interaction relationships. Finally, the features output by each attention head are concatenated, and attention output information is obtained through linear transformation, thereby strengthening the core features and suppressing redundant information.
[0077] The output layer maps the attention output information to a risk level through a fully connected layer, probability calibration, and dynamic threshold adjustment, ultimately obtaining the vehicle's fogging risk level. First, the attention output information is input into the fully connected layer to complete feature transformation and spatial mapping. This fully connected layer has 64 neurons and uses the LeakyReLU activation function to map the 128-dimensional attention output feature vector to the risk assessment space, while simultaneously performing feature compression and nonlinear transformation to enhance the expression of fogging risk-related features. Specifically, this embodiment adds a dynamic threshold adjustment layer to the fully connected layer, which can obtain the current weather type (sunny, rainy, foggy, winter, etc.) in real time and automatically adjust the risk judgment decision boundary accordingly. This solves the problem of poor adaptability of traditional fixed boundaries in complex weather conditions, laying a solid foundation for subsequent accurate classification.
[0078] Subsequently, a fogging risk probability is output through a Sigmoid activation function. This Sigmoid activation function, combined with a temperature scaling mechanism, calibrates the fogging risk probability, and this probability value directly quantifies the magnitude of the vehicle fogging risk. Temperature scaling, as a core innovative design, is used to correct the Sigmoid output distribution to improve prediction reliability. The calculation formula is as follows:
[0079] In the formula Indicates the probability of fogging risk. This represents the output feature value of the fully connected layer, and T represents the dynamic temperature parameter, which is adaptively adjusted according to the consistency index CI. For example, when CI ≥ 0.7, T = 0.8, making the probability distribution more concentrated and strengthening the determinism of the judgment; when CI < 0.3, T = 1.2, making the probability distribution smoother and enhancing the anti-interference ability.
[0080] Finally, based on the calibrated Risk levels are determined by combining a three-tiered risk mapping rule with a weather-adaptive dynamic threshold. The basic rule is: low risk (P_fog < 0.3), medium risk (0.3 ≤ P_fog < 0.7), and high risk (P_fog ≥ 0.7). Furthermore, in this embodiment, to adapt to different weather characteristics, a dynamic threshold adjustment mechanism is introduced to replace the fixed threshold in the basic rule, improving adaptability. The specific standards are: sunny days (low risk upper limit 0.25, high risk lower limit 0.75), rainy days (0.35, 0.65), foggy days (0.4, 0.6), and winter (0.2, 0.8). The actual classification prioritizes the use of the corresponding weather dynamic threshold.
[0081] For example, in a foggy scene, the attention output is processed by a fully connected layer to obtain the feature value z. Combined with a high CI of T=0.8, P_fog=0.62 is calculated, which is judged as high risk according to the fog threshold (high risk lower limit 0.6). In a sunny day, P_fog=0.25, which is judged as low risk according to the sunny day threshold. In winter, P_fog=0.82, which meets the winter high risk lower limit of 0.8, and is judged as high risk. In a rainy day, P_fog=0.6, which is between 0.35 and 0.65, and is judged as medium risk.
[0082] The fogging risk assessment model in this invention comprehensively covers fogging influencing factors through multimodal input data. Feature preprocessing and fusion optimize and enhance information, bidirectional LSTM fully exploits temporal dependencies, and multi-head attention precisely focuses on core features. Finally, a fully connected layer and activation function quantify the risk level. Firstly, it provides comprehensive data coverage, considering both real-time status and historical trends, avoiding biases from single-dimensional judgments. Secondly, the model architecture is lightweight, with parameters adapted to in-vehicle embedded scenarios, ensuring real-time response (latency ≤20ms). Thirdly, it offers significant advantages in accurate risk assessment compared to traditional methods. Fourthly, it is highly practical for engineering, as all data comes from existing vehicle sensors and electronic rearview mirrors, requiring no additional hardware upgrades and reducing implementation costs. This process provides a reliable and accurate quantitative basis for the dynamic adjustment of subsequent safety control strategies, significantly improving the intelligence level of vehicle fogging risk prediction.
[0083] In an optional implementation, the fogging risk assessment model constructed based on steps b1 to b5 above adopts an end-to-end architecture of "input-preprocessing-fusion-temporal modeling-attention focusing-output" to achieve accurate quantitative assessment of fogging risk. The complete structure is as follows: 1. Input layer The core components consist of two main types of inputs: real-time features and historical sequence features, covering the current state and dynamic trends of factors influencing fog formation.
[0084] Specific parameters: Real-time features: including temperature difference between inside and outside the vehicle, humidity proxy index, weather code, and consistency coefficient, totaling 4 dimensions, directly representing the core environmental state at the current moment; Historical sequence: Extract the above four types of real-time features from the past five time steps (1 second each) to form 5×4 dimensional data and capture the evolution of parameters over time.
[0085] 2. Feature Preprocessing Layer Core function: Denoising, feature enhancement, and dimensionality optimization of input data to provide high-quality data support for subsequent modeling.
[0086] Specific components and parameters: 1D Convolutional Layer: Configured with Kernel=3, Stride=1, Channels=64, outputting 5×64-dimensional features for extracting local temporal patterns and smoothing sensor noise; Spatial Attention (SE module): Without changing the feature dimension, it outputs 5×64-dimensional features and strengthens the role of core influencing factors (such as humidity proxy indicators) by learning the feature channel weights. Temporal pooling: MaxPooling (kernel=3) is used to output 3×64-dimensional features. Key time-series nodes are retained through downsampling, reducing data redundancy.
[0087] 3. Spatiotemporal Feature Fusion Layer Core functionality: Enables deep interaction between real-time features and time-series features, integrates spatiotemporal dimensional information, and enhances feature representation capabilities.
[0088] Specific components and parameters: Feature stitching: The 3×64-dimensional temporal features are fused with the 4-dimensional real-time features to output 3×68-dimensional features, thus achieving preliminary integration of spatiotemporal information; Fully connected layer: configured with 128 nodes, outputting 3×128-dimensional features to facilitate non-linear interactions between different types of features; LayerNorm (normalization layer): It does not change the feature dimension, outputs 3×128 dimensional fused features, stabilizes the data distribution, and avoids gradient explosion during training.
[0089] 4. Bidirectional LSTM layer Core function: To mine the positive and negative temporal dependencies in the fusion features and capture the dynamic evolution of fogging risk.
[0090] Specific components and parameters: Forward LSTM: Configure Hidden_size=128, Layers=2, output 128-dimensional features, and model forward temporal dependencies in chronological order; Backward LSTM: The configuration is the same as the forward LSTM (Hidden_size=128, Layers=2), outputs 128-dimensional features, and models inverse temporal dependencies in reverse time order; State fusion: The forward and backward output features are spliced together, and a 256-dimensional state feature is output through linear transformation, which fully covers the temporal context information.
[0091] 5. Multi-head attention processing layer Core Function: Accurately focus on the core influencing factors in state characteristics, suppress redundant information, and enhance the contribution of key features.
[0092] Specific components and parameters: Projection layer: Key_dim=32, which maps the 256-dimensional state features to a 4×32-dimensional feature space, laying the foundation for attention computation; Four attention heads: Employing the ScaledDot-ProductAttention mechanism, they focus on different dimensions of information in parallel (short-term mutations, long-term trends, cross-feature interactions, etc.), each outputting 4×32-dimensional features; Output splicing and transformation: The outputs of the four attention heads are spliced together, and 128-dimensional attention output information is output through linear transformation, which is then combined with the core feature focusing result.
[0093] 6. Output layer Core function: Map attention output information to a quantified fog risk level and complete the assessment result output.
[0094] Specific components and parameters: Fully connected layer: configured with 64 nodes, outputting 64-dimensional features, and compressing and optimizing the attention output information; The Sigmoid activation function maps 64-dimensional features to probability values in the [0,1] interval, divides fogging risk into three levels (low, medium, and high) using probability thresholds, and finally outputs a 1-dimensional risk assessment result.
[0095] In some alternative implementations, step a1 above includes: Step c1: Calculate the temperature difference between the inside and outside of the vehicle using data from other sensors; Step c2: Estimate humidity proxy index based on the condition of the rear windshield; Step c3: Identify the current weather using the electronic rearview mirror image and other sensor data, encode the current weather, and add the enhancement score calculated using other sensor data to the weather code; Step c4: Calculate the consistency coefficient, which characterizes the consistency of fogging phenomenon of vehicles behind, based on the condition of the windshield of the vehicle behind. Step c5 involves extracting historical sequences of data from other sensors. The multimodal input data includes the temperature difference between inside and outside the vehicle, humidity proxy indicators, weather codes, consistency coefficients, and historical sequences.
[0096] Specifically, the embodiments of the present invention clarify the specific creation process of multimodal input data, ensuring that the data input to the model not only fully covers the factors affecting fogging, but also has the characteristics of standardization and computability, laying the foundation for subsequent accurate model evaluation.
[0097] First, the temperature difference between the inside and outside of the vehicle is calculated using data from other sensors. Among these sensors, the temperature and humidity sensors can collect real-time temperature data (in degrees Celsius, °C) for both the inside and outside of the vehicle. The temperature difference (ΔT) is calculated by subtracting the outside temperature from the inside temperature, and it is the core physical factor in determining glass fogging. When the temperature difference reaches a certain threshold, the warm, humid air inside the vehicle easily condenses and forms fog when it comes into contact with the low-temperature glass surface. For example, if the temperature and humidity sensor detects that the current inside temperature is 25°C and the outside temperature is 18°C, then by calculating 25°C - 18°C = 7°C, the temperature difference ΔT = 7°C is obtained. This value directly reflects the potential physical conditions for fogging.
[0098] Second, a humidity proxy index is estimated based on the condition of the rear vehicle's windshield. A humidity proxy index is a quantitative indicator that indirectly represents ambient humidity when the external humidity cannot be directly obtained, using the condition of the rear vehicle's windshield. Its value ranges from 0 to 1 (0 represents extremely low humidity, and 1 represents extremely high humidity). The degree of fogging on the rear vehicle's windshield is the core basis for calculating this index; the more severe the fogging, the higher the external humidity, and the larger the corresponding index value.
[0099] In one optional implementation, the humidity proxy index is calculated as follows:
[0100] In the formula, Indicates humidity proxy index, This represents the distance decay weight of the i-th distance. , This vehicle represents the distance to the i-th vehicle following it. For the preset distance threshold, This indicates the total number of vehicles inspected. This indicates the degree of fogging on the windshield of the i-th vehicle following it.
[0101] The core function of the humidity proxy index is to convert the fogging status of the rear vehicle's windshield into quantifiable ambient humidity data. Its value ranges from 0 to 1; a value closer to 1 indicates higher external humidity and a greater risk of fogging, while a value closer to 0 indicates lower external humidity and a lower risk of fogging. The calculation follows a "weighted average" logic, adjusting for the contribution of different rear vehicles through distance attenuation weights to avoid interference from the status of distant vehicles in the evaluation results.
[0102] parameter The total number of vehicles detected behind is output by the vehicle detection algorithm of the electronic rearview mirror image. That is, it's the total number of vehicles identified and counted from the side and rear of the electronic rearview mirror image by the target detection model. For example, when driving on urban roads, if three vehicles are detected behind, then... =3.
[0103] parameter The degree of fogging on the windshield of the i-th following vehicle is determined by semantic segmentation and image feature analysis. The value range is 0-1: 0 means the glass is completely clear with no fogging; 1 means the glass is completely covered by fog and the interior of the vehicle cannot be seen; the intermediate values correspond to different degrees of partial fogging.
[0104] parameter The relative distance between this vehicle and the i-th following vehicle is obtained by a vehicle ranging system (such as millimeter-wave radar, lidar, or visual ranging algorithm), and the unit is meters (m).
[0105] parameter The preset distance threshold is a fixed value calibrated based on engineering practice and a large amount of test data. It is used to adjust the rate of distance decay and is usually set to 10m (which can be finely adjusted according to vehicle model and application scenario). Its core function is to give higher weight to vehicles behind at close range and lower weight to vehicles behind at long range, which is in line with the actual scenario that "close-range environment is more valuable for reference".
[0106] parameter The distance decay weight for the i-th distance is determined by the formula... Calculations show that the closer the distance, the closer the weight is to 1; the farther the distance, the closer the weight is to 0.
[0107] After obtaining all the above parameters, the humidity proxy index is calculated step by step according to the formula. This quantitative formula transforms the vague "fogging status of the following vehicle" into a precise humidity proxy index, solving the industry pain point of not being able to directly obtain external environmental humidity and providing a reliable quantitative basis for fogging risk assessment. Furthermore, the formula incorporates distance attenuation weights, making the status of nearby following vehicles more valuable as a reference, effectively filtering out interference from distant vehicle statuses, and improving the accuracy and reliability of the index. In addition, the parameters are clearly defined, the calculation logic is simple, and all parameters can be obtained through existing vehicle sensors and algorithms without additional hardware, making it highly feasible in engineering. Moreover, the standardized index value range facilitates unified processing by the subsequent fogging risk assessment model, laying a solid foundation for the model to accurately output the risk level. This upgrades fogging risk prediction from "qualitative perception" to "quantitative assessment," significantly improving the intelligence and accuracy of the entire safety pre-control system.
[0108] Third, the current weather is coded, and enhancement scores calculated using data from other sensors are added to the weather code. Weather coding converts fine-grained weather types such as "sunny," "cloudy," "rainy," "foggy," "snowy," and "dust storm" into digital labels recognizable by the model. This typically uses one-hot encoding (i.e., each weather condition corresponds to a unique numerical vector; for example, sunny weather is coded as [1,0,0,0,0,0], and rainy weather as [0,1,0,0,0,0]). Enhancement scores are supplementary values calculated based on data from other sensors (such as rain gauges and temperature / humidity sensors) to enhance the representational capability of the corresponding weather code. For example, if a rain gauge detects a rainfall intensity of 0.7 (a continuous value from 0 to 1) on a rainy day, this value is added as an enhancement score to the rainy weather code, so that the code reflects not only the weather type but also its intensity. Through the weather coding provided in this embodiment of the invention, the representation is upgraded from "only indicating type" to a comprehensive representation of "type and intensity," providing the model with more accurate weather information.
[0109] Fourth, in this embodiment of the invention, a consistency coefficient is calculated based on the windshield condition of the following vehicles to characterize the consistency of fogging phenomena. The consistency coefficient is an indicator that quantifies the synergy of fogging phenomena among multiple following vehicles, with a value range of 0-1 (0 indicates no consistency in fogging phenomena, and 1 indicates that all following vehicles have completely consistent fogging conditions). Its core logic is that "the higher the proportion of fogging vehicles and the smaller the difference in fogging degree, the stronger the consistency, indicating a more stable ambient humidity." The calculation requires counting the total number of detected following vehicles and the number of fogging vehicles, and calculating the standard deviation of the fogging degree for all following vehicles. The smaller the standard deviation, the more uniform the fogging degree.
[0110] The specific formula is as follows:
[0111] In the formula, Represents the consistency coefficient. This indicates the number of vehicles whose rear end is fogged up. This indicates the total number of vehicles inspected. The standard deviation represents the degree of fogging on the vehicle behind.
[0112] Specifically, the core logic of the consistency coefficient is that "the higher the proportion of vehicles experiencing fogging and the smaller the difference in fogging degree among vehicles, the more stable the external humidity and the clearer the risk of fogging." Its value ranges from 0 to 1. The closer the value is to 1, the stronger the consistency of fogging phenomenon of the following vehicles, the more stable the external environment is in a high humidity state, and the higher the risk of fogging for this vehicle. The closer the value is to 0, the more chaotic the fogging state of the following vehicles is, mostly caused by the vehicle's own factors, the poor stability of the external humidity, and the relatively low risk of fogging for this vehicle.
[0113] parameter The degree of fogging for each following vehicle was determined through semantic segmentation and image feature quantification analysis, with values ranging from 0 to 1: 0 indicates completely clear glass with no fog coverage; 1 indicates the glass is completely obscured by fog, making it impossible to see inside the vehicle; intermediate values correspond to different degrees of partial fogging. For example, if approximately 60% of the glass of a following vehicle is fogged, then... =0.6.
[0114] parameter The standard deviation represents the degree of fogging among all following vehicles. A larger value indicates a greater difference in the degree of fogging among the following vehicles, while a smaller value indicates a more uniform degree of fogging.
[0115] After obtaining all the above parameters, calculate the consistency coefficient step by step according to the formula. This reflects the prevalence of fogging. Reflecting the consistency of fogging degree, and Multiply to obtain the consistency coefficient. .
[0116] By quantifying the "fogging status of multiple following vehicles" into a precise consistency coefficient, the pain point of related technologies being unable to determine whether fogging of following vehicles is caused by environmental factors is solved, providing a reliable environmental correlation basis for fogging risk assessment; it can filter out interference caused by abnormal fogging of a single vehicle, making fogging risk assessment more in line with the actual environment, avoiding misjudgments caused by the special status of individual vehicles, and further enhancing the reliability and accuracy of the entire safety pre-control system.
[0117] Fifth, this embodiment of the invention extracts historical sequences of data from other sensors. The historical sequence is a time-series data set formed by extracting core data from other sensors (such as vehicle interior and exterior temperatures, rainfall, etc.) over the past five time steps (each time step is 1 second, for a total of 5 seconds). Its function is to capture the dynamic changing trends of environmental parameters. Current data alone cannot reflect the changing patterns of parameters, while the historical sequence allows the model to perceive dynamic information such as "whether the temperature difference is continuously increasing" and "whether the rainfall is gradually increasing." For example, extracting the vehicle interior and exterior temperature difference data for the past 5 seconds as 5℃, 5.5℃, 6℃, 6.5℃, and 7℃ forms a historical sequence of temperature changes. Combined with the current vehicle interior and exterior temperature difference (7℃), humidity proxy index (0.82), weather code (rainy day + enhancement score 0.7), and consistency coefficient (0.784), it constitutes complete multimodal input data, comprehensively covering the core information of "current state + dynamic trend."
[0118] The technical means provided in this invention have constructed comprehensive and standardized multimodal input data. The data dimensions are complete, covering the temperature difference between vehicle interior and exterior, humidity proxy indicators, weather codes, consistency coefficients, and historical sequences. It includes both the core physical factors of fogging and the external environmental characterization and dynamic trends, avoiding assessment biases caused by single data dimensions. The quantification logic is scientific; humidity proxy indicators and consistency coefficients are quantified through clear calculation rules, and weather codes are introduced to enhance the accuracy of characterization, transforming ambiguous environmental information into digital signals that the model can calculate. The introduction of historical sequences meets the needs of subsequent time-series modeling modules such as bidirectional LSTM, providing data support for mining parameter change patterns. This multimodal input data provides high-quality, comprehensive input support for the fogging risk assessment model, forming the core foundation for the model to achieve accurate and early warning capabilities, effectively improving the reliability and intelligence level of the entire safety pre-control system.
[0119] In some alternative implementations, step b1 above includes: Step d1 involves standardizing the temperature difference between inside and outside the vehicle, humidity proxy indicators, consistency coefficients, and historical sequences. Step d2: Perform one-dimensional convolution on the standardized multimodal data after standardization to obtain the first multimodal feature; Step d3: Assign weight information to each feature in the first feature of the multimodal model using a spatial attention mechanism; Step d4: Perform time pooling on the first multimodal feature with assigned weight information to obtain the second multimodal feature.
[0120] Specifically, the embodiments of the present invention complete the preprocessing of multimodal input data through a four-step process of standardization, one-dimensional convolution, spatial attention weighting, and temporal pooling. The core objectives are noise reduction, enhancement of core features, and simplification of data dimensions, providing high-quality and highly available multimodal features for subsequent feature fusion and model inference, and avoiding the problem of poor model performance caused by messy and redundant original data.
[0121] Standardization is a fundamental step in data preprocessing. Its core function is to eliminate dimensional differences between different types of data (e.g., the temperature difference between inside and outside a vehicle is measured in °C, while humidity is a dimensionless value between 0 and 1), mapping all data to the same numerical range. This prevents a single feature from dominating model training due to an excessively large numerical range, ensuring that each feature is treated fairly by the model. Specifically, the Z-score standardization method is used. Historical sequence data is standardized step-by-step according to the same rules to ensure the consistency of time-series data.
[0122] One-dimensional convolution (1D convolution) is a feature extraction technique for time-series data. It scans standardized time-series data through a sliding window (convolution kernel) to capture local temporal patterns (e.g., the correlation between "continuously increasing temperature difference and synchronously rising humidity proxy indicators"), while smoothing out random noise generated during sensor acquisition (e.g., minute instantaneous temperature fluctuations). In this embodiment, the parameters of the one-dimensional convolutional layer are configured as follows: Kernel=3 (kernel size is 3, i.e., scanning data from 3 consecutive time steps each time), Stride=1 (kernel slides 1 time step each time), and Channels=64 (outputting 64 feature channels to enhance feature representation). For example, a standardized historical sequence is 5×4 dimensional data (5 time steps, 4 types of features). After one-dimensional convolution processing, a 5×64 dimensional multimodal first feature is output. This feature retains the temporal correlation of the original data and enhances the discriminative power of local features through multi-channel output.
[0123] The core logic of the spatial attention mechanism (using the SE module and Squeeze-Excitation module in this implementation) is to "automatically learn the importance of each feature channel," assigning higher weights to features with a greater impact on fogging risk (such as feature channels corresponding to humidity proxy indicators) and lower weights to features with a smaller impact (such as historical temperature difference channels with no significant changes), thereby strengthening the contribution of core features and suppressing redundant information. The specific process includes:
[0124]
[0125] In the formula, The attention weight of the i-th input feature represents the importance of that feature to the current prediction task. The original score of the i-th feature; For all input features j Summing is performed, and the sum is used as the normalized denominator. The whole is a softmax function, which converts the original score into a softmax function. Convert to a probability distribution between 0 and 1, where the sum of all weights is 1. The output layer weight matrix maps the hidden representations to scalar scores; The activation function introduces non-linearity and filters out negative values. It is the hidden layer weight matrix, which maps the input features to the hidden representation space; It is the i-th input feature vector (such as the temperature difference between inside and outside the vehicle, humidity proxy, etc.); It is a bias term that increases the model's expressive power.
[0126] The core function of temporal pooling (MaxPooling in this embodiment) is to downsample temporal features, retaining key temporal node information while simplifying data dimensions and reducing model computational burden, thus adapting to the real-time requirements of automotive embedded scenarios. The temporal pooling parameters are configured as kernel=3 (pooling window size is 3, meaning the maximum value is taken from the features of 3 consecutive time steps each time), and a sliding step size of 1. For example, after temporal pooling, 5×64-dimensional weighted features output as 3×64-dimensional multimodal second features. During pooling, each pooling window retains the maximum value within that interval, corresponding to key temporal nodes such as "sudden temperature change" and "sudden increase in humidity proxy index," ensuring that core trend information is not lost. Simultaneously, the time steps are reduced from 5 to 3, effectively reducing the computational overhead of subsequent modules.
[0127] The standardized processing in this invention unifies the data scale, resolving the issue of inconsistent dimensions in multimodal data and laying the foundation for fair learning of features by the model. One-dimensional convolution accurately captures local temporal patterns, smooths sensor noise, and improves feature reliability. Spatial attention mechanism enables adaptive enhancement of core features, allowing the model to focus more on key influencing factors of fogging risk and improving feature expressiveness. Temporal pooling simplifies data dimensions while retaining key information, reducing the model's computational burden and ensuring the preprocessing process adapts to the real-time processing requirements of in-vehicle systems. The entire preprocessing process is progressive and collaboratively optimized, transforming the original multimodal input data into high-quality multimodal second features, providing solid support for subsequent feature fusion and accurate inference of the fogging risk assessment model.
[0128] In some alternative implementations, step b2 above includes: Step e1: Concatenate the multimodal second feature and the weighted weather code to obtain the concatenated feature; Step e2: Process the spliced features through a fully connected layer, and randomly discard some feature channels during the fully connected layer processing to obtain higher-order features; Step e3 involves normalizing the higher-order features to obtain the fused features. The magnitude of the normalization process is positively correlated with the consistency coefficient.
[0129] Specifically, the embodiments of the present invention achieve deep fusion of multimodal features and weather information through a three-step process of feature splicing, fully connected layer transformation, and conditional normalization. This integrates key environmental information, extracts high-order correlation features, and suppresses noise interference, providing accurate and stable fused features for subsequent time series modeling and avoiding model performance degradation caused by limitations of a single feature dimension or data redundancy.
[0130] The multimodal second feature is the core temporal feature (3×64 dimensions) retained after preprocessing, covering the dynamic changes of key parameters such as the temperature difference between inside and outside the vehicle and humidity proxy indicators. The weather code is a digital label (12 dimensions) representing the current weather type and intensity, containing fine-grained weather information such as sunny, rainy, and foggy days. In this embodiment of the invention, the weather code is attention-weighted before splicing. The correlation between the current environment and various weather features is analyzed through an attention mechanism, and higher weights are given to weather features that are highly relevant to the current scene (for example, in a rainy scene, the weight of the coding dimension corresponding to "rainy day" is increased), so that the model focuses more on key weather information and weakens the interference of irrelevant weather features. The weighted weather code is still 12 dimensions, and it is spliced with the 3×64-dimensional multimodal second feature in terms of channel dimensions, finally obtaining a 3×(64+12)=3×76-dimensional spliced feature, realizing the initial integration of temporal features and weather information. For example, the second feature of the multimodal model focuses on "the humidity proxy index continues to rise in the last 3 seconds". The weighted weather code strengthens the "light rain" feature. After splicing, the associated feature "humidity continues to rise in light rain" is formed, which lays the foundation for subsequent high-order feature extraction.
[0131] The fully connected layer is the core component for realizing nonlinear feature transformation. In this embodiment of the invention, this layer is configured with 256 neurons and uses the LeakyReLU activation function. It can still retain weak gradients when the input is negative, avoiding the gradient vanishing problem, and is more suitable for the feature transformation requirements in vehicle scenarios. The core function of the fully connected layer is to perform nonlinear mapping on the 3×76-dimensional spliced features, explore the potential correlations between different feature dimensions (such as the combined correlation of "rainy day, large temperature difference, high humidity"), and extract more expressive high-order features. At the same time, in order to improve the robustness of the model and avoid overfitting, this embodiment of the invention introduces a feature dropout regularization method, which randomly drops some feature channels during training and inference (for example, the dropout probability is preset to 0.2), simulates the feature loss situation under different scenarios, and forces the model to rely on core features for judgment. For example, after the spliced features are processed by the fully connected layer, the output is a 3×256-dimensional feature vector. After feature dropout, a 3×256-dimensional high-order feature is obtained. This feature retains the core correlation information and has stronger anti-interference ability.
[0132] The core function of normalization is to stabilize the distribution of feature data and accelerate model training convergence. This invention employs a conditional LayerNorm mechanism, specifically dynamically adjusting the normalization parameters based on the consistency coefficient (CI) to achieve "on-demand normalization." The consistency coefficient characterizes the synergy of fogging phenomena in following vehicles, and its value is positively correlated with the stability of external environmental humidity: when CI is high (≥0.7), it indicates that the external environmental humidity is stable and the fogging correlation information is reliable. In this case, the normalization intensity is reduced to retain more details of environmental changes (such as small fluctuations in humidity proxy indicators); when CI is low (<0.3), it indicates that the external environmental humidity is unstable and there may be noise in the data. In this case, the normalization intensity is increased to suppress noise interference and stabilize the data distribution; when CI is in the middle range, the normalization amplitude is adaptively adjusted to balance detail preservation and noise suppression. For example, if the CI for a high-order feature is 0.8 (high consistency), the parameter adjustment range is reduced during normalization to preserve the detailed correlation between "continuously rising humidity and rainy days"; if CI is 0.2 (low consistency), the normalization range is increased to filter out noise caused by abnormal fogging of individual vehicles. The final output is a 3×256-dimensional fusion feature, which has both high-order correlation expression capabilities and stable data distribution characteristics.
[0133] The present invention employs attention-weighted concatenation of weather codes, enabling the model to accurately focus on key weather features and resolving the issue of insufficient emphasis on core information caused by the "equal distribution" of traditional concatenation. Furthermore, the combination of fully connected layers and feature discarding achieves the dual goals of high-order feature extraction and robustness improvement, adapting to complex and ever-changing in-vehicle application scenarios. Conditional LayerNorm dynamically adjusts the normalization amplitude based on the consistency coefficient, making the normalization process more closely reflect the actual environment, preserving effective information while suppressing noise, thus improving feature reliability. This provides high-quality data support for subsequent bidirectional LSTM time-series modeling, significantly enhancing the accuracy and stability of the fog risk assessment model.
[0134] In some alternative implementations, step b3 above includes: Step f1: Obtain the first LSTM module and the second LSTM module; Step f2 involves processing the fused features sequentially using the first LSTM module and then processing them in reverse chronological order using the second LSTM module. The input gates of the first and second LSTM modules are processed as follows:
[0135] In the formula, This represents the input gate value for time step t. This represents the activation function. This represents the weight matrix of the input gate. This means concatenating the hidden state from the previous time step with the input feature vector of the current time step. This represents the bias term of the input gate. The weather impact coefficient represents the magnitude of the weather impact, adjusted according to the current weather conditions. The weight matrix representing weather codes. Indicates weather codes; Step f3 involves concatenating and linearly transforming the hidden states output by the first LSTM module and the second LSTM module to obtain the state features.
[0136] Specifically, this embodiment of the invention uses a bidirectional LSTM module for temporal modeling to deeply mine the temporal dependencies in the fused features, capture the dynamic evolution of fog risk, and improve the model's adaptability to different weather scenarios through a weather-adaptive input gate, providing comprehensive and accurate state features for subsequent attention focusing.
[0137] The LSTM module is a deep learning model specifically designed for processing time-series data. It effectively captures dependencies in long sequences and addresses the vanishing or exploding gradient problems of traditional recurrent neural networks. This embodiment of the invention uses two structurally identical LSTM modules, defined as the first LSTM module and the second LSTM module, both configured as two-layer structures with a hidden layer dimension (Hidden_size) of 128. This ensures feature representation capabilities while adapting to the computational resource constraints of automotive embedded scenarios. The two modules operate independently but collaboratively, processing and fusing features from different temporal directions to jointly construct complete temporal context information.
[0138] The first LSTM module processes the fused features in chronological order (from the 1st second to the 3rd second), models positive temporal dependencies, and captures positive change patterns such as "gradual increase in humidity proxy index and stable rain intensity". The second LSTM module processes the fused features in reverse temporal order (from the 3rd second to the 1st second), models reverse temporal dependencies, and verifies the rationality of positive changes. For example, it confirms that "high humidity in the 3rd second" is caused by "continuous accumulation of humidity in the previous 2 seconds" rather than a sudden anomaly.
[0139] The core of this invention lies in the weather-adaptive improvement of the input gate. The input gate is a core component of the LSTM module, used to control whether the input feature vector at the current time step can enter the memory unit, and its output value... The closer to 1, the more current features are allowed to enter the memory unit; the closer to 0, the more current features are rejected.
[0140] This represents the weather impact coefficient, which adjusts in magnitude according to the current weather conditions, such as rainy days, foggy days, and other high humidity weather. ≈0.8 (Impact of weather on input gate), under sunny and dry weather conditions ≈0.3 (weakening the impact of weather). The weight matrix representing the weather code has dimensions of 128×12 (corresponding to the hidden layer dimension and the 12-dimensional weather code), and is used to map the weather code to a feature space that fits the input gate. Weather codes are digital representations of the current weather type and intensity.
[0141] For example, in a rainy scenario, ≈0.8, the weather coding Wcode enhances the "rainy day" feature; the input gate calculation will focus on considering current humidity-related features. The higher value allows more humidity-related information to enter the memory unit; in a sunny scenario, With a value of approximately 0.3, the impact of weather is reduced, and the input gate focuses more on core physical factors such as temperature difference to avoid interference from irrelevant weather information.
[0142] After processing by the first LSTM module, a 128-dimensional forward hidden state is output, containing forward temporal dependency information. The second LSTM module outputs a 128-dimensional inverse hidden state, containing inverse temporal dependency information. First, the two 128-dimensional hidden states are concatenated to obtain a 256-dimensional concatenated feature, which fully covers the forward and inverse dependencies of the temporal data. Then, a linear transformation layer is used to integrate and optimize the concatenated feature, eliminating redundancy and conflicts in the output features of the two modules, ultimately outputting a 256-dimensional state feature, providing a high-quality temporal foundation for subsequent attention processing.
[0143] This invention employs a bidirectional LSTM module to model temporal dependencies from both forward and reverse time directions. Compared to traditional unidirectional LSTM, this approach more comprehensively captures the dynamic evolution of fog risk, avoiding dependency omissions caused by a single direction. The weather-adaptive input gate allows the model to dynamically adjust feature input strategies based on different weather scenarios, improving modeling accuracy in complex weather conditions. Risk prediction accuracy is significantly enhanced for scenarios such as rain and fog. The module parameter configuration is lightweight, with a dual-layer 128-dimensional hidden layer ensuring feature representation capabilities while controlling computational overhead, adapting to the real-time requirements of in-vehicle systems. State features integrate forward and reverse temporal information, providing richer contextual representation capabilities and solid support for the subsequent multi-head attention module to accurately focus on core features, further improving the overall performance of the fog risk assessment model.
[0144] In some alternative implementations, step b4 above includes: Step g1: Set up a first attention head, a second attention head, a third attention head, and a fourth attention head. The first attention head, the second attention head, and the third attention head are used to extract attention information of the same feature in the state features at different time lengths. The fourth attention head is used to capture attention information between different types of features in the state features. Step g2 involves adjusting the query vectors of each attention head by querying the weight matrix. The weight matrix is calculated using the following formula:
[0145] In the formula, This represents the query weight matrix. This represents the original query weight in the standard attention mechanism. This represents the modulation intensity coefficient, with a value ranging from 0 to 1, and its magnitude is adjusted according to the current weather conditions. The weight matrix representing weather codes. Indicates weather codes; Step g3: Perform multi-head attention processing on the state features using each attention head, and concatenate and linearly transform the attention sub-information output by each attention head to obtain the first attention output information; Step g4: Perform a residual connection between the first attention output information and the state features output after LSTM processing to obtain the second attention output information.
[0146] Specifically, this embodiment of the invention uses a multi-head attention mechanism with four heads working together to accurately focus on the core temporal information and cross-feature association information in the state features. At the same time, it introduces weather-adaptive query weight adjustment and residual connection to improve the model's adaptability to different weather scenarios and the completeness of feature representation.
[0147] The core advantage of multi-head attention lies in its ability to capture information from different dimensions of features in parallel, avoiding information focusing bias caused by a single attention head. This embodiment of the invention addresses the needs of fog risk assessment by assigning differentiated functions to four attention heads: the first, second, and third attention heads focus on extracting attention information of the same feature (e.g., extracting humidity proxy indicators and vehicle-inside / outside temperature difference) at different time lengths. The first attention head focuses on short-term time series (e.g., the most recent 1 second), the second attention head focuses on medium-term time series (e.g., the most recent 2 seconds), and the third attention head covers long-term time series (e.g., the most recent 3 seconds). Through this layered focus on the time dimension, the dynamic change patterns of features are fully captured. The fourth attention head specifically captures the correlation information between different features, such as the combined correlation between "vehicle-inside / outside temperature difference and humidity proxy indicators" and "weather type and consistency coefficient," uncovering the multi-factor synergistic influence patterns of fog risk. The output dimensions of all four attention heads are configured to 32 dimensions, ensuring the balance of feature representation and the convenience of subsequent fusion.
[0148] The query vector is the core carrier used to match key information in the attention mechanism. The innovation of this invention lies in dynamically adjusting the query weight matrix through weather coding, so that the attention head is more in line with the feature requirements of the current weather scene.
[0149] This represents the modulation intensity coefficient, ranging from 0 to 1, and its value is adjusted according to the current weather. For example, in high humidity weather (such as rain or fog), γ≈0.8 (strengthening the impact of weather on the query vector), while in low humidity weather (such as sunny or dry weather), γ≈0.3 (weakening the impact of weather), balancing the weight of weather factors and core physical factors. For example, in a foggy scenario, γ=0.8, weather coding. The "foggy day" characteristic is enhanced through calculation using a formula. This will cause the query vectors of each attention point to focus more on features that are strongly correlated with fog formation, such as humidity proxy indicators and consistency coefficients; in the clear sky scenario, γ=0.3, the weather influence is weakened, and the query vectors focus more on core physical factors such as the temperature difference between the inside and outside of the vehicle, avoiding interference from irrelevant weather information.
[0150] First, the state features are input into four attention heads. Each attention head calculates its attention weight. The first to third attention heads score the importance of a single target feature (such as a humidity proxy indicator) within the state features over different time periods. A higher weight indicates a greater impact of the feature on the risk of fogging at that time point. The fourth attention head scores the correlation strength between different features. Then, each attention head outputs attention sub-information, containing the focused core features. Finally, the sub-information from the four attention heads is concatenated to obtain the concatenated features. A linear transformation layer then integrates and optimizes the concatenated features, eliminating redundancy and conflicts in the outputs of each head, ultimately yielding the first attention output information. For example, the first attention head focuses on "a sudden increase in humidity in the last second," the second attention head focuses on "a sustained high humidity level in the last two seconds," the third attention head covers "a rising trend in humidity in the last three seconds," and the fourth attention head highlights the correlation between "high humidity and large temperature difference." After concatenation and transformation, a comprehensive and focused core feature is formed.
[0151] The core function of residual connections is to preserve the complete information of the original state features, avoiding feature loss during attention processing and enhancing the training stability of the model. Specifically, the first attention output information is adapted to the dimensionality of the state features (e.g., by reducing the dimensionality of the state features to 128 dimensions using a 1×1 convolution), and then element-wise summed to obtain the second attention output information. (Formula reference:) output = Linear(cat(heads)) +0.3·LSTM_state In the formula, cat(heads) represents the concatenation operation of the attention head outputs, Linear(cat(heads)) represents the linear transformation processing, and 0.3 is a weighting coefficient used to deeply integrate the first attention output information after multi-head attention focusing with the temporal context features of the bidirectional LSTM output, and finally generate the second attention output information for risk level mapping.
[0152] This operation ensures that the output information includes both the core features focused by the attention mechanism and the temporal context information of the original state features, making the second attention output information more comprehensive and reliable.
[0153] This invention's four-headed attention system enables parallel processing of time-dimension hierarchical focusing and cross-feature association capture. Compared to traditional single-headed attention, it more comprehensively mines key information in state features, enhancing the richness of feature representation. Weather-adaptive query weight adjustment allows the model to dynamically adjust the focus direction according to different weather scenarios, improving risk prediction accuracy by over 18% in complex scenarios such as rain and fog. Concatenation and linear transformation effectively integrate information from multiple attention heads, while residual connections ensure the integrity of feature information, preventing information loss during focusing. This provides high-quality, focused feature support for accurate mapping of fog risk levels, further strengthening the overall performance of the fog risk assessment model.
[0154] In some optional implementations, steps S102 and c3 above include: Step h1: Extract the first intermediate features from the electronic rearview mirror image using the embedding module and the visual Transformer module; Step h2: Extract the second intermediate feature from the electronic rearview mirror image using a convolutional neural network module and a cascaded attention module; Step h3: Perform basic convolution processing, convolutional layer processing, and fully connected layer processing on the second intermediate feature to obtain the third intermediate feature, and classify the third intermediate feature to obtain the current weather. Step h4: Fuse the first intermediate feature and the second intermediate feature to obtain the fourth intermediate feature; Step h5: Input the fourth intermediate feature into the parallel attention module to identify the following vehicle in the electronic rearview mirror image; Step h6: Decode the image of the rear vehicle in the electronic rearview mirror to obtain the windshield status of the rear vehicle.
[0155] Specifically, such as Figure 3 As shown, this embodiment of the invention achieves the above three capabilities simultaneously through a multi-task model based on vehicle detection, rear vehicle windshield state segmentation, and weather classification.
[0156] This invention achieves multi-task parsing of electronic rearview mirror images (weather recognition, rear vehicle detection, and glass condition recognition) simultaneously through dual-branch feature extraction, weather classification, feature fusion, rear vehicle recognition, and glass condition determination. The core objective is to improve the accuracy and efficiency of environmental perception in complex scenarios through differentiated feature extraction and collaborative fusion, providing comprehensive and reliable visual perception data for fog risk assessment.
[0157] like Figure 3 As shown, the first intermediate feature is extracted from the electronic rearview mirror image through an embedding module and a visual Transformer module (ViT Block). The electronic rearview mirror image is a 1280×720 pixel color image containing rich information such as vehicles to the side and rear, road environment, and weather. The core function of the embedding module is to transform the two-dimensional image into a one-dimensional feature vector that the model can process. First, the image is divided into multiple image blocks of a fixed size (e.g., 16×16 pixels). Then, each image block is transformed into a feature vector of a fixed dimension (e.g., 768 dimensions) through linear projection, while adding positional encoding to preserve the spatial position information of the image blocks. The visual Transformer module captures the global correlation between image blocks through a multi-head self-attention mechanism, such as the correlation between "the fog in the distant sky and the fogging state of the nearby vehicle's glass" and "the road water accumulation and the trajectory of raindrops," mining the global semantic information of the image (the ViT Block structure is existing technology and will not be elaborated further here). After the above processing, the first intermediate feature with high dimension and strong global correlation is output (such as [197,768] dimensions, corresponding to the number of image patches and classification token and feature dimension). This feature is good at capturing long-distance dependence and global scene information, providing global context support for rear vehicle recognition.
[0158] The Convolutional Neural Network (CNN) module employs multiple convolutional layers (such as a combination of Conv2d, BatchNorm, and ReLU activation functions) to extract local features from the image through sliding convolutional kernels, capturing detailed information such as edges, textures, and local shapes, including vehicle outlines, glass boundaries, and local texture features of raindrops / fog. The Concatenated Attention (CSAM) module, as a feature enhancement component, adopts a "global-then-local" concatenated structure. It first extracts global features and channel correlations, then refines and downsamples local features, enhancing the feature representation of key targets such as weather and vehicles while suppressing background interference. For example, in a rainy scene, the CNN module extracts the local texture of raindrops, while the CSAM module enhances raindrop features while suppressing road background interference, ultimately outputting a second intermediate feature (e.g., [H / 8, W / 8, 512] dimensions, where H and W are the original image height and width) that combines local details and channel optimization. This feature excels at capturing local details, providing accurate detail support for weather classification and glass condition recognition.
[0159] The second intermediate feature contains rich local details and weather-related features (such as strong sunlight texture on sunny days, raindrop trails on rainy days, and low-contrast features on foggy days), requiring a series of processing steps for weather classification. First, the second intermediate feature is dimensionally adjusted and optimized using a basic convolutional module (CBS) to unify the feature scale. Then, a convolutional layer (Conv) further extracts key weather-related features and filters out irrelevant information. Next, a fully connected layer (Fc) transforms the two-dimensional feature map into a one-dimensional feature vector (e.g., 128-dimensional), obtaining the third intermediate feature. Finally, a softmax activation function is used to classify the third intermediate feature, outputting fine-grained weather types (sunny, cloudy, rainy, foggy, snowy, dusty). For example, if the "raindrop trails" and "wet road surface texture" features are prominent in the third intermediate feature, the classification result is rainy; if the "low contrast" and "global blur" features are significant, it is classified as foggy, providing accurate weather scene labels for fog risk assessment. Figure 4 As shown in (A), this is a schematic diagram illustrating the sunny weather effect in weather classification, as follows: Figure 4 (B) shows a schematic diagram of the cloudy weather effect in the weather classification.
[0160] The first intermediate feature excels at capturing global contextual information, while the second intermediate feature focuses on local detail representation. Their fusion achieves complementary "global plus local" information, avoiding the limitations of a single feature dimension. The fusion process is implemented through a fusion module (Fusion Block). First, the first intermediate feature undergoes dimensional adaptation (transforming it into a feature map with the same spatial dimensions as the second intermediate feature through linear projection). Then, a combination of channel stitching and weighted fusion is used to integrate the two types of features. Channel stitching retains all feature information, while weighted fusion dynamically adjusts the weights of the two types of features according to task requirements (rear vehicle recognition) (e.g., global feature weight 0.4, local feature weight 0.6), strengthening the expression of vehicle-related features. Finally, a non-linear transformation using an activation function (e.g., ReLU) outputs a fourth intermediate feature (e.g., [H / 8, W / 8, 1024] dimensions). This feature contains both global scene context and retains local details of key targets such as vehicles and glass, providing high-quality feature support for rear vehicle recognition.
[0161] The Parallel Attention Module (CPAM) adopts a "global and local parallel" structure. It processes the fourth intermediate feature in parallel and then fuses the two outputs, preserving global vehicle distribution information while accurately capturing local vehicle contour details, thus improving vehicle detection capabilities in complex backgrounds (such as multiple vehicles or road obstacles). The feature map output by this module is processed by a target detection head (such as a YOLO detection head). Through bounding box regression and category classification, the position coordinates (e.g., [x1, y1, x2, y2]) and confidence score of the following vehicle are obtained. Detection boxes with confidence scores higher than a preset threshold (e.g., 0.5) are selected, representing the following vehicles identified in the electronic rearview mirror image. The recognition effect is as follows: Figure 5 As shown.
[0162] Next, the rear vehicle in the electronic rearview mirror image is decoded to obtain the windshield status of the rear vehicle. First, based on the output rear vehicle position bounding box, a local image of each rear vehicle is cropped from the original electronic rearview mirror image. Then, the decoder upsamples and refines the features corresponding to the rear vehicle position in the fourth intermediate feature. The decoder uses transposed convolution to gradually restore the feature map resolution and combines skip connections to fuse shallow detail features, achieving pixel-level instance segmentation of the rear vehicle's windshield area, and obtaining a binary mask of the glass area (the rear vehicle's windshield segmentation effect is shown in the image). Figure 6 (As shown). Finally, the state of the glass area is determined based on its image characteristics (such as the variance of gray values, sharpness, and texture uniformity of the segmented glass area). If the variance of gray values is large and the sharpness is high, it is judged as "sharp"; if the gray values are uniform, the sharpness is low, and there is a blurry texture, it is judged as "fogging / frost".
[0163] The embodiments of this invention utilize dual-branch feature extraction to achieve differentiated capture of global and local information, adapting to the needs of different perception tasks and improving environmental perception accuracy compared to single feature extraction schemes. Secondly, the targeted application of serial / parallel attention modules (CSAM for weather classification, CPAM for vehicle detection) strengthens key target features, suppresses interference, and improves robustness in complex scenarios. Thirdly, multi-task collaborative processing (weather classification, rear vehicle detection, glass status recognition) shares the feature extraction module, reducing computational overhead and providing comprehensive and accurate visual perception data for fog risk assessment.
[0164] In some optional implementations, the serial attention module includes: a first module connected to the input and a second module connected to the first module; the parallel attention module includes a first module and a second module respectively connected to the input, and the outputs of the first module and the second module are superimposed. The first module includes a first branch and a second branch. The first branch is used to feed the input data into the global average pooling layer and the bottleneck layer in sequence. The second branch is used to feed the input data into the global max pooling layer and the basic convolutional layer in sequence. The output of the first module is the sum of the outputs of the first branch and the second branch. The second module includes a third branch and a fourth branch. The third branch is used to feed the input data into the global average pooling layer and the convolutional layer in sequence, and the fourth branch is used to feed the input data into the convolutional layer. The output of the second module is the sum of the outputs of the third branch and the fourth branch.
[0165] Specifically, such as Figure 7 As shown, the Serial Attention Module (CSAM) and Parallel Attention Module (CPAM) constructed in this embodiment of the invention are both based on the core structures of the first and second modules, and adapt to different perception task requirements through differentiated connection methods. CSAM focuses on "step-by-step optimization" feature extraction, and is used for tasks such as weather classification that require defining the scene first and then refining it; CPAM focuses on "multi-path complementarity" feature fusion, and is used for tasks such as rear vehicle detection that require precise positioning.
[0166] The first module (CA_SCA module) and the second module (CA_Conv module) are the core components of the two types of attention modules. Both adopt a dual-branch parallel structure, which strengthens core information and suppresses redundant interference by extracting and superimposing features from different branches.
[0167] Module 1 Structure and Workflow: The first module focuses on global feature and channel correlation extraction, and includes a first branch and a second branch. The two branches process the input data in parallel and then superimpose the output.
[0168] The first branch: The input data first enters the Global Average Pooling (GAP) layer, which performs global averaging calculation on each channel of the feature map, extracts global statistical information at the channel level, and weakens local noise interference; then, the pooled features are compressed and non-linearly transformed through the Bottle Neck layer (for example, compressing high-dimensional features to 1 / 4 of the original dimension and then restoring them), which strengthens the key correlations between channels and reduces computational overhead.
[0169] The second branch: The input data first passes through a Global Max Pooling (GMP) layer to capture local peak features within each channel (such as edge peaks of vehicle contours and strong response areas of weather textures); then, it passes through a basic convolutional layer (CBS) for feature mapping and dimensionality unification, preserving key local details. The structure of the basic convolutional layer is existing technology and will not be elaborated further in this embodiment of the invention.
[0170] Output fusion: The global statistical features of the first branch and the local peak features of the second branch are element-wise added and superimposed, and then processed by ordinary convolution (Conv) to obtain the output features of the first module. This feature contains both global channel correlation information and preserves local peak details, providing a comprehensive foundation for subsequent feature processing.
[0171] Module Two Structure and Workflow: The second module focuses on local feature refinement and downsampling, and also includes a third and fourth branch, which work together to optimize feature representation.
[0172] The third branch: The input data first passes through a global average pooling layer to extract global channel information and filter out local abnormal interference; then it passes through a regular convolutional layer (Conv) to perform feature downsampling (with a stride of 2), and generates a weight map through the convolutional layer, which is multiplied with the original feature map to achieve channel attention.
[0173] Fourth branch: Input data is directly processed through a regular convolutional layer (Conv) for local feature extraction and downsampling (stride set to 2), preserving the spatial correlation of local details (such as the continuous texture of glass boundaries and the local distribution of raindrops).
[0174] Output fusion: The global optimization features from the third branch and the local detail features from the fourth branch are added element-wise to obtain the output features of the second module. This feature combines the spatial integrity of local details with the anti-interference properties of global optimization, making it suitable for accurate target localization and detail recognition.
[0175] The concatenated attention module adopts a "first module → second module" concatenated connection method to achieve hierarchical feature optimization from "global to local" and is suitable for tasks that require scenario prior guidance, such as weather classification. The input features first enter the first module, and through dual-branch fusion, features containing global channel correlation and local peaks are obtained, completing the initial extraction of global scene information (such as global raindrop distribution features in rainy scenes and low-contrast global features in foggy scenes).
[0176] The output features of the first module are directly input into the second module. After bi-branch downsampling and local refinement, the feature representation of key targets (such as weather-related textures and vehicle outlines) is enhanced, while the interference of the background (such as road-irrelevant textures and distant non-target objects) is suppressed. At the same time, the feature dimension is reduced to improve computational efficiency.
[0177] The final output features undergo two levels of optimization: "global refinement" and "local refinement". They possess both scene-level global relevance and target-level local details, enabling them to accurately support the scene feature requirements of tasks such as weather classification.
[0178] The parallel attention module adopts a connection method of "parallel input and output superposition of the first module and the second module" to achieve complementary features of multiple paths.
[0179] The input features are simultaneously split into the first module and the second module, which process them independently and in parallel. The first module focuses on global channel correlation and local peak feature extraction, while the second module focuses on local detail refinement and downsampling.
[0180] After obtaining the global-local fusion features of the first module and the local-downsampling features of the second module respectively, the two types of features are added and superimposed element by element to achieve multi-path information complementarity. The features of the first module ensure global correlation and local peak response, while the features of the second module enhance the spatial integrity of local details and receptive field coverage.
[0181] The final output features combine the advantages of multi-path processing, which not only avoids information loss in a single path, but also enhances the adaptability to multi-scale targets (such as large vehicles at close range and small vehicles at long range), and can accurately support the target localization requirements of tasks such as rear vehicle detection.
[0182] This invention unifies the core structures of the first and second modules through modular design, reducing module design complexity. Simultaneously, it forms CSAM and CPAM through differentiated connection methods, adapting to the feature requirements of different perception tasks. The dual-branch structure enables a single basic module to capture both global and local features, enhancing the richness of feature representation.
[0183] By employing a concatenated structure, CA_SCA first captures global semantic information, and then CA_Conv focuses on local details, forming a feature extraction process from macro to micro. Convolutions with a stride of 2 replace pooling layers, reducing the number of parameters while maintaining feature resolution and avoiding information loss. Experiments show that while the SCA module improves recall, it may introduce false positives, while CSAM, through local correction by subsequent CA_Conv, can effectively suppress redundant responses, making it suitable for guided tasks such as weather classification.
[0184] The parallel structure of CPAM allows CA_SCA and CA_Conv to process features independently, avoiding information loss in a single path and enabling more comprehensive capture of target features, especially in complex scenes. No additional transformations are performed before adding the two feature paths, preserving the integrity of the original features and enhancing the model's adaptability to multi-scale targets. Fusion is achieved through simple addition, resulting in low computational overhead, making it suitable for real-time detection tasks and target recognition tasks such as rear vehicle detection, providing efficient and reliable feature enhancement support for multi-task perception.
[0185] In some optional implementations, step S104 above includes: Step I1: Define the risk-scenario dynamic mapping relationship for controlling the parameters of the vehicle air conditioning system. The risk-scenario dynamic mapping relationship is used to indicate that the higher the fogging risk level, the faster the defogging effect of the adjusted air conditioning parameters, and the lower the fogging risk level, the more the adjusted air conditioning parameters are adapted to the user's comfort index. Step I2: Define a control table based on the risk-scenario dynamic mapping relationship. The control table is used to assign corresponding air conditioning system control parameters to each fogging risk level. Step I3: Obtain environmental perception data and vehicle status data. Environmental perception data includes fogging risk level, windshield status of following vehicles, and current weather. Vehicle status data includes current air conditioning system parameters and driving scenario information. Step I4: Match the acquired environmental perception data and vehicle status data with the control table to obtain the current target air conditioning system control parameters, so as to adjust the current air conditioning parameters.
[0186] Specifically, the risk-scenario dynamic mapping relationship is the core logical principle of the air conditioning control strategy, clearly defining the correspondence between "fogging risk level" and "air conditioning parameter adjustment priority." The higher the fogging risk level, the higher the priority of defogging effect, and the air conditioning parameter adjustment must focus on quickly eliminating or preventing fog. Conversely, the lower the fogging risk level, the higher the priority of driving comfort, and the air conditioning parameter adjustment must conform to the user's daily usage habits to avoid unnecessary energy consumption and physical discomfort. Specifically, high-risk levels (often with a fogging probability ≥70%) correspond to a "safety first" strategy, prioritizing clear visibility and allowing for a moderate sacrifice of some comfort; medium-risk levels (typically with a fogging probability of 30%~70%) correspond to a "safety-comfort balance" strategy, maintaining a good driving experience while preventing fogging; and low-risk levels (typically with a fogging probability <30%) correspond to a "comfort first" strategy, maintaining the user-set air conditioning parameters and only making minor adjustments when potential risks occur. For example, in foggy weather, if multiple vehicles behind are identified as foggy (determined to be high-risk) by the electronic rearview mirror, the mapping relationship will trigger the "quick defogging" logic; while in sunny weather, if none of the vehicles behind are foggy (determined to be low-risk), the mapping relationship will maintain the "comfortable operation" logic and no additional defogging intervention will be performed.
[0187] Then, a control table is defined based on the risk-scenario dynamic mapping relationship, assigning corresponding air conditioning system control parameters to each fogging risk level. The control table serves as a standardized and executable carrier of the mapping relationship, clearly defining the specific configuration of core air conditioning parameters under different risk levels. It covers four key parameters: target temperature setting, airflow distribution ratio, internal / external circulation mode, and defogging intensity, ensuring the consistency and operability of the control strategy. The internal / external circulation mode refers to the air circulation method of the air conditioning system. Internal circulation uses only the air inside the vehicle, blocking the entry of external air, suitable for scenarios with high external humidity; external circulation introduces fresh air from outside the vehicle, suitable for scenarios with dry external air.
[0188] For example, in one optional implementation, the specific control table is configured as follows: At high risk levels, the target temperature is set to 20~22℃ (the optimal dehumidification range, slightly lower than the normal comfort temperature), the airflow distribution is 70% for the windshield and 30% for the side windows (focusing on key fogging areas), forced internal circulation (blocking high humidity air from the outside), and the defogging intensity is set to the strong mode (maximum airflow for the windshield + linkage of rearview mirror / rear window heating). At medium risk level, the target temperature is set to 23~24℃ (balancing dehumidification and comfort), the airflow distribution is 60% for the front windshield and 40% for the side windows (dynamically adapting to the fogging area), intelligent circulation (internal circulation when external humidity is >80%, external circulation when it is <60%), and the defogging intensity is set to enhanced mode (airflow increased by 30% + local heating as needed). At low risk levels, the target temperature is maintained at the user-set value (e.g., 22~26℃), the air volume is evenly distributed (default comfort mode), the user-set circulation mode is the main mode (automatically switching to external circulation when the outside is dry), and the defogging intensity is set to standard mode (basic defogging guarantee).
[0189] The above configuration ensures rapid defogging capability in high-risk scenarios while avoiding excessive intervention in low-risk scenarios.
[0190] Next, environmental perception data and vehicle status data are acquired to provide real-time basis for parameter matching. Environmental perception data is the core input for judging the current scene, including the fogging risk level (output by the fogging risk assessment model), the status of the rear vehicle's windshield (electronic rearview mirror image recognition results, such as "3 vehicles fogged, 2 clear"), and the current weather (fine-grained classification results, such as rain / fog / sunny). These three types of data jointly characterize the fogging-related conditions of the external environment. Vehicle status data is the basic reference for adjusting air conditioning parameters, including the current air conditioning system parameters (user-set temperature, airflow, circulation mode, etc.) and driving scenario information (vehicle speed, turn signals, window opening / closing status), such as vehicle speed 60km / h, windows closed, current air conditioning temperature 25℃, and airflow level 2. This data is collected in real time through the vehicle bus to ensure that the data is synchronized with the actual scene, providing accurate and real-time input support for subsequent parameter matching.
[0191] Finally, the environmental perception data and vehicle status data are matched with the control table to obtain and adjust the current target air conditioning system control parameters. First, the fogging risk level is extracted from the environmental perception data to determine the target parameter range in the control table. Then, the parameters are optimized by combining the current weather and driving scenario information. For example, in rainy or foggy weather, the contrast of the electronic rearview mirror is enhanced on top of the control table, and in dusty weather, dry air is introduced through external circulation. Finally, the parameter adjustment range is calculated with reference to the current air conditioning system parameters (e.g., if the current temperature is 25℃ and the target temperature is 21℃, then the temperature is reduced by 4℃). The adjustment command is executed through the air conditioning controller to achieve a smooth transition of parameters (avoiding sudden changes in temperature and airflow that could cause discomfort). For example, if the environmental perception data shows a high fogging risk level and the current weather is foggy, and the vehicle status data shows the current air conditioning temperature is 24℃, fan speed is level 2, and the windows are closed, the target parameters obtained after matching the control table are: temperature 21℃, fan speed level 4, internal circulation, and strong defogging. The air conditioning system will adjust gradually according to these parameters to quickly reduce the risk of glass fogging. If the fogging risk level is low and the current weather is sunny, the user-set 25℃, fan speed level 2, and external circulation mode will be maintained, retaining only basic defogging protection.
[0192] This invention upgrades air conditioning control from a traditional "passive response" (intervening only after fog formation) to "active prediction and precise regulation" by defining a dynamic risk-scenario mapping relationship and control table. This significantly improves defogging efficiency and effectively avoids obstructed visibility due to untimely defogging. Control parameters are configured differently according to risk levels, balancing driving safety and passenger comfort, and solving the problems of untimely defogging or excessive dehumidification (such as discomfort caused by strong cold air blowing directly) resulting from the "one-size-fits-all" control of traditional air conditioning systems. The control logic is implemented based on the vehicle's existing hardware, requiring no additional dedicated control module, ensuring high engineering feasibility and achieving an organic unity of safety, comfort, and energy efficiency, significantly improving the vehicle's intelligence level and driving experience.
[0193] In some optional implementations, step S104 above further includes: Step I5: Monitor the defogging effect inside the vehicle, adjust the current fogging risk level according to the defogging effect, and then return to the step of adjusting the current air conditioning parameters according to each risk scenario based on the risk-scenario dynamic mapping relationship; Step I6: Calculate the relative distance and relative speed between the vehicle and the vehicle behind based on the image of the vehicle behind in the electronic rearview mirror; Step I7: Adjust the blind spot warning threshold according to the fog risk level; Step I8: When a dynamic blind spot corresponding to the blind spot warning threshold is detected based on the relative distance and relative speed, output warning information, steering intervention suggestion, and / or braking prompt.
[0194] Specifically, this embodiment of the invention adds a closed-loop adjustment function for defogging effect and a blind spot collaborative warning function, further improving driving safety and system intelligence.
[0195] Defogging effectiveness monitoring is a core closed-loop component ensuring the effectiveness of the control strategy. Data is collected collaboratively by an in-vehicle camera and temperature and humidity sensors. The in-vehicle camera captures real-time images of the windshield and side windows, and image clarity algorithms (such as edge detection and grayscale variance calculation) are used to determine whether fog still exists on the glass. The temperature and humidity sensors continuously collect the glass surface temperature and the relative humidity inside the vehicle to help verify the defogging effect (for example, if the glass surface temperature is higher than the dew point temperature and the humidity is lower than 60%, the defogging is considered to be up to standard). For example, if the air conditioning is adjusted to medium-risk parameters for 5 seconds, and the camera detects that there are still partially blurred areas on the windshield, and the humidity sensor shows that the humidity inside the vehicle is 75%, then the defogging effect is determined to be insufficient. The current fogging risk level is then raised from "medium" to "high," and the control table is re-matched to further increase the defogging intensity (e.g., increasing the fan speed to the maximum setting and activating the glass heater). If the glass is detected to be clear and the humidity is below 55%, then the defogging is determined to be sufficient, and the risk level is lowered by one level (e.g., from "medium" to "low"). The air conditioning parameters are adjusted accordingly to balance comfort and energy efficiency, ensuring that the air conditioning control strategy can dynamically adapt to environmental changes and avoid insufficient or excessive defogging due to fixed parameters. For another example, when the risk level drops to "low" and remains so for 30 seconds, the defogging intensity is gradually reduced (strong → enhanced → standard), restoring the comfort-first air conditioning setting (e.g., the temperature rises back to 25°C).
[0196] Furthermore, in the control strategy for linked blind spot warning, relative distance refers to the straight-line distance between the vehicle and the vehicle behind, and relative speed is the speed of the vehicle behind relative to the vehicle itself (the speed of the vehicle behind minus the speed of the vehicle itself; a positive value indicates that the vehicle behind is approaching, and a negative value indicates that the vehicle behind is moving away). The calculation process is based on the target tracking algorithm of the electronic rearview mirror image. First, the position of the same vehicle behind in consecutive image frames is locked through inter-frame matching technology. Combining the camera parameters of the electronic rearview mirror (focal length, installation height) and the image pixel ratio, the pixel coordinates of the vehicle behind in the image are converted into the actual physical distance (i.e., relative distance). Then, by the change in the position of the vehicle behind in adjacent frames and the frame interval time (30fps corresponds to a frame interval of approximately 33ms), the moving speed of the vehicle behind is calculated, and then the driving speed of the vehicle itself (obtained through the vehicle's CAN bus) is subtracted to obtain the relative speed.
[0197] A blind spot refers to an area behind or to the side of a vehicle that the driver cannot directly observe through the rearview mirror. The blind spot warning threshold is the critical distance at which a blind spot warning is triggered (i.e., when a following vehicle enters this distance range, the system activates the warning). Traditional vehicle warning thresholds are mostly fixed values and do not consider the impact of fog on visibility. This invention links the fog risk level with the blind spot warning threshold. The higher the fog risk, the more severely the driver's visibility is obstructed, and the warning threshold needs to be increased accordingly (to improve warning sensitivity) to allow more reaction time. The lower the fog risk, the better the visibility conditions, and the warning threshold can be appropriately reduced (to avoid excessive warnings interfering with driving). For example, at a low risk level (clear visibility), the blind spot warning threshold is set to 3 meters (warning is issued when the following vehicle is ≤3 meters away from the vehicle); at a medium risk level, the threshold is increased to 5 meters; and at a high risk level (severely obstructed visibility), the threshold is further increased to 8 meters to ensure that the driver has enough time to react to following vehicles in the blind spot.
[0198] When a vehicle following the driver enters a dynamic blind spot corresponding to the blind spot warning threshold, the system outputs warning information, steering intervention suggestions, and / or braking alerts. The dynamic blind spot refers to the area defined by the adjusted blind spot warning threshold. The system continuously compares the relative distance between the vehicle and the following vehicle with the warning threshold to determine if the following vehicle has entered the blind spot: when the relative distance is less than or equal to the warning threshold, the system determines that the following vehicle has entered the dynamic blind spot, and at this time, the system activates multi-dimensional safety alerts. Warning information includes visual warnings on the in-vehicle dashboard (such as a flashing blind spot indicator) and voice warnings from the audio system (such as "Caution: Vehicle in the left blind spot"), ensuring the driver quickly perceives the risk; steering intervention suggestions apply a slight counter-torque through the electric power steering system to remind the driver that there is a risk in the current steering (e.g., when the driver activates the left turn signal and there is a vehicle in the left blind spot, the steering system applies slight resistance to the right); braking alerts are simultaneously communicated to the driver via a pop-up window on the dashboard and voice prompts, suggesting "slow down." If the relative speed is too high (e.g., a following vehicle approaches at a speed ≥20 km / h), the system can trigger slight braking (requires driver-permitted driving assistance mode) to reduce the risk of collision. In addition, when a following vehicle enters the dynamic blind spot at a relatively high speed, the urgency of the warning is increased. For example, the flashing yellow light is changed to a flashing red light to indicate a higher risk of collision. For instance, in a high-risk fog scenario, the blind spot warning threshold is 8 meters. When a following vehicle is detected at a relative distance of 7 meters and a relative speed of 15 km / h (approaching the vehicle), the system immediately activates the left indicator light to flash and issues a voice warning, "Beware of a vehicle approaching in the left blind spot." At the same time, when the driver attempts to activate the left turn signal, steering resistance is applied, and a suggestion to "reduce speed to 40 km / h" is given, ensuring comprehensive safety during lane changes or driving.
[0199] This invention addresses the issue of "one-time adjustments without feedback" through closed-loop adjustment of the defogging effect. By dynamically calibrating the fogging risk level, it ensures that air conditioning parameters are always adapted to actual defogging needs. Precise calculation of relative distance and relative speed provides a quantitative basis for blind spot warnings, avoiding the limitations of traditional warnings that "only detect existence, not assess danger." The linkage and adaptation between the fogging risk level and the blind spot warning threshold solves the problem of "blind spots within blind spots" caused by obstructed visibility in severe weather. Warning sensitivity is adjusted as needed, reducing the risk of blind spot collisions by 41%. Finally, the combination of multi-dimensional warnings and intervention suggestions provides both rapid driver alerts through audible and visual warnings and substantial safety support through steering resistance and braking indicators, while preserving driver control and balancing safety and driving experience.
[0200] In some optional implementations, step S104 above further includes: Step I9: Record the driver's manual intervention behavior regarding air conditioning parameters; Step I10: Generate adjustment weights based on manual intervention behaviors; Step I11: Adjust the control parameters of the air conditioning system by adjusting the weights; Step I12: Adjust the contrast and sharpness of the electronic rearview mirror image according to the visibility of the current weather.
[0201] Specifically, manual intervention refers to actions performed by the driver, after the system has automatically adjusted the air conditioning parameters, through the vehicle's air conditioning control panel, central control screen, or voice commands, to manually modify parameters such as temperature, airflow, circulation mode, and defogging intensity. This includes actions that overshadow system decisions, such as forcibly shutting down the defogging function or significantly adjusting defogging-related parameters. This invention uses a reinforcement learning framework to collect drivers' personalized usage habits, providing data support for subsequent strategy weight adjustments.
[0202] This invention monitors air conditioning control commands in real time via the vehicle's CAN bus, accurately distinguishing between "automatic system commands" and "manual driver commands." Based on a reinforcement learning-based experience replay mechanism, it records manual intervention actions in multiple dimensions, including intervention time, current fogging risk level, real-time environmental perception data (weather, in-vehicle humidity), system automatic parameters before intervention, driver-set parameters after intervention, intervention magnitude, and whether it covers core system decisions (e.g., the system determines high risk and activates powerful defogging, the driver forcibly disables the defogging function). For example, if the system automatically adjusts the air conditioning temperature to 23°C, fan speed to level 3, and activates enhanced defogging at a medium risk level, and the driver manually raises the temperature to 25°C, lowers the fan speed to level 2, and forcibly disables the defogging function, the system records this intervention as follows: intervention magnitude +2°C, fan speed -1 level, defogging disabled; associated scenario: cloudy day, medium fogging risk level, 68% in-vehicle humidity; labeled as "covering core system decisions." All intervention data is stored in the reinforcement learning experience pool, categorized and archived by driver identity, providing high-quality samples for subsequent weight calculations.
[0203] The adjustment weight is a coefficient that quantifies the driver's intervention preference, ranging from 0 to 1. It represents the driver's willingness to intervene and the intensity of their intervention overriding system decisions. A higher weight indicates a stronger demand from the driver for adjustments to the system's automatic parameters, requiring subsequent parameter corrections to better align with this intervention habit. Weight generation employs an algorithm combining reinforcement learning and statistical analysis, based on the frequency of driver intervention overriding system decisions in similar scenarios, the magnitude of intervention, and the feedback effects after intervention (defogging effect, ride comfort). If a user frequently overrides system decisions in similar scenarios (same weather, fog risk level) (e.g., 3 or more consecutive interventions all forcibly shutting down defogging or significantly modifying defogging parameters), the driver's preference is considered clear, and the adjustment weight is adjusted towards 1. The larger the magnitude of a single intervention, especially the larger the adjustment of defogging-related parameters, the higher the weight assigned. If the defogging effect is satisfactory after intervention (e.g., clear glass, interior humidity ≤60%) and lasts for a relatively long time (e.g., ≥1 minute), the intervention is considered reasonable, and the weight is further increased through a positive reward mechanism of reinforcement learning. Conversely, the weight is reduced through a negative penalty mechanism to weaken the impact of unreasonable intervention. For example, in a cloudy, medium-risk scenario, if the driver manually intervenes three times, increasing the temperature by 2°C and decreasing the fan speed by one level, and forcibly shuts down the system's automatically activated enhanced defogging function twice, and the defogging effect still meets the standard after the intervention, the system will generate an adjustment weight of 0.8 for this scenario through reinforcement learning iterative calculation. This indicates that the system's automatic parameters need to be adapted to the driver's preference for "increasing temperature, decreasing fan speed, and weakening defogging" by 80%. If the driver only intervenes occasionally in a certain scenario, and the intervention causes the glass to fog up and affect visibility, the weight will be set to 0.2 through negative penalty to weaken the impact of the intervention on the automatic parameters.
[0204] Subsequently, the driver's personalized preferences are integrated into the system's automatic control strategy. The revised parameters ensure both effective defogging and driver comfort, while also respecting the preferences formed by the driver's frequent interactions with system decisions. The revision logic is: revised parameters = automatic system parameters + (driver intervention amplitude × adjustment weight). For core decisions frequently made by the driver (such as forcibly turning off the defogging), a separate preference coefficient is added to prioritize adapting to such habits. At the same time, parameter boundary thresholds are set (such as temperature range 16~32℃, airflow range 1~7) to prevent the revised parameters from exceeding the equipment's operating range or severely affecting the defogging effect. For example, the system automatically generates parameters for high-risk levels: temperature 21℃, fan speed 4, internal circulation, and strong defogging. Combining the adjustment weight of 0.8 for this scenario and the driver's historical intervention habits (multiple forced shutdown of defogging, temperature increase of 2℃, fan speed reduction of 1 level), the corrected parameters are calculated as follows: temperature = 21℃ + (2℃ × 0.8) = 22.6℃ (rounded to 23℃), fan speed = 4 levels + (-1 level × 0.8) = 3.2 levels (rounded to 3 levels). The circulation mode maintains internal circulation, and the defogging intensity is adjusted to standard mode (instead of strong defogging) to accommodate the driver's preference for frequently shutting down the defogging.
[0205] The contrast and sharpness of an electronic rearview mirror image directly affect the driver's ability to clearly observe the environment to the sides and rear. Contrast refers to the difference in brightness between bright and dark areas of the image, while sharpness refers to the clarity of edge details. Traditional electronic rearview mirrors often use fixed parameters, which cannot adapt to the needs of different visibility scenarios. This invention dynamically adjusts parameters to maintain the best display effect of the electronic rearview mirror image under different visibility conditions. First, the visibility data of the current weather is obtained through the environmental perception module. Visibility is comprehensively determined by the electronic rearview mirror image sharpness algorithm combined with the weather type, and then the image parameters are adapted according to the visibility level. For example, in high visibility, the contrast is set to the default value (e.g., 50%) and the sharpness is set to medium (e.g., 40%) to avoid the image being too sharp and causing visual fatigue; in medium visibility, the contrast is increased to 65% and the sharpness is increased to 55% to enhance the distinction of image details and clearly present the outline of the following vehicle; in low visibility, the contrast is further increased to 80% and the sharpness is adjusted to 70%, while the edge enhancement algorithm is activated to enhance key information such as the edges of the following vehicle and road markings, and to counteract the blurring caused by fog and raindrops.
[0206] This invention utilizes reinforcement learning to record driver intervention behaviors, particularly optimizing the weight generation logic for behaviors that overlay system decisions. This enables the air conditioning control strategy to possess personalized adaptive capabilities, addressing the problems of traditional systems where "uniform parameters cannot adapt to different driver habits" and "frequent overlay of user decisions," thereby improving driver satisfaction. Based on visibility, the electronic rearview mirror image parameters are dynamically adjusted to specifically optimize the field of vision under different weather conditions, resolving issues of blurred images and loss of detail in inclement weather. This further enhances the integrated capability of vehicle safety perception and comfort control, improving the overall intelligent driving experience.
[0207] In some optional embodiments, the safety pre-control method based on electronic rearview mirror data provided by the present invention further includes: Step k1: When raindrops or fog are detected in the electronic rearview mirror, the image enhancement algorithm is activated to process the electronic rearview mirror image.
[0208] Specifically, embodiments of the present invention solve the problems of blurred electronic rearview mirror images and obstructed field of vision caused by raindrops and fog by real-time status detection and targeted image enhancement processing.
[0209] During vehicle operation, electronic rearview mirror cameras may experience raindrops or fogging on their lens surface due to environmental factors such as rain and high humidity. This can lead to blurry images, light spots, and loss of detail. Therefore, it is necessary to first detect the image status of the electronic rearview mirror in real time, and then perform targeted image enhancement processing based on the detection results. The image status detection of the electronic rearview mirror is achieved through a visual feature analysis algorithm. The system analyzes each frame of the image captured in real time by the electronic rearview mirror. If irregular semi-transparent light spots, streaks of watermark texture, and a sudden drop in brightness in a local area are detected, it is determined that raindrops are present on the electronic rearview mirror. If the overall image contrast is reduced, edge details are blurred, there are no obvious local texture features, and the image grayscale values are uniformly distributed, it is determined that the electronic rearview mirror lens is fogged. The above detection process is performed simultaneously with the electronic rearview mirror image acquisition to ensure rapid response to raindrops and fogging.
[0210] When raindrops or fog are detected in the electronic rearview mirror, the system immediately activates an image enhancement algorithm to process the currently acquired and subsequently acquired electronic rearview mirror images in real time. This image enhancement algorithm restores clear image details through multiple processing steps. First, image defogging / deraining preprocessing is performed, using a dark channel prior algorithm to suppress scattered light caused by fog and raindrops, reducing the homogenization blur and highlighting the image's edge contours. Next, contrast enhancement processing is performed, using an adaptive histogram equalization algorithm to adjust the image's grayscale distribution, improving the distinction between bright and dark areas and addressing insufficient image contrast caused by raindrops and fog. Finally, detail sharpening processing is performed, enhancing the image's edge features and restoring details of key environmental information such as the outline of following vehicles and road markings. Simultaneously, Gaussian filtering is used to eliminate noise generated during the processing, ensuring overall image quality. The electronic rearview mirror image processed by this image enhancement algorithm effectively eliminates interference from raindrops and fog, clearly presenting key information such as following vehicles and the road environment to the side and rear of the vehicle. If the electronic rearview mirror image detects no raindrops or fog, the system directly outputs the original image, avoiding unnecessary algorithmic processing that consumes computing resources and ensuring the operational efficiency of the vehicle system. Furthermore, dynamic defogging progress indicators (such as "Windshield defogging 80% complete") can be overlaid on the in-vehicle screen to enhance user perception.
[0211] This invention enables the on-demand activation of image enhancement algorithms by real-time detection of raindrops and fogging in electronic rearview mirrors. This ensures effective restoration of blurred images while avoiding ineffective algorithm operation, balancing image processing performance with the computational efficiency of the vehicle system, and providing stable support for the overall vehicle safety perception system.
[0212] This embodiment also provides a safety pre-control device based on electronic rearview mirror data. This device is used to implement the above embodiments and preferred embodiments, and details already described will not be repeated. As used below, the term "module" can refer to a combination of software and / or hardware that implements a predetermined function. Although the device described in the following embodiments is preferably implemented in software, hardware implementation, or a combination of software and hardware, is also possible and contemplated.
[0213] In one specific embodiment, please refer to Figure 8 The present invention provides a safety pre-control method based on electronic rearview mirror data, the complete process of which is as follows: 1. Collect images from the electronic rearview mirror while simultaneously collecting data from other sensors.
[0214] 2. Input the electronic rearview mirror image into the multi-task model, and simultaneously identify the following vehicle, the current weather, and the windshield status of the following vehicle.
[0215] 3. Input the current weather, the windshield status of the vehicle behind, and other sensor data into the fog risk assessment model, process the data, and output the fog risk level.
[0216] 4. Finally, the fogging risk level is used to achieve defogging control of the air conditioning system, blind spot warning control, and electronic rearview mirror image enhancement control.
[0217] This embodiment provides a safety pre-control device based on electronic rearview mirror data, such as... Figure 9 As shown, it includes: Data acquisition module 901 is used to acquire images from the electronic rearview mirror; The windshield status recognition module 902 is used to extract the following vehicle from the electronic rearview mirror image and recognize the windshield status of the following vehicle. The fogging risk assessment module 903 is used to determine the fogging risk of this vehicle by observing the state of the windshield of the vehicle behind it. The strategy control module 904 is used to generate a vehicle safety control strategy based on the risk of fogging.
[0218] The safety pre-control device based on electronic rearview mirror data provided in this embodiment of the invention can execute the safety pre-control method based on electronic rearview mirror data provided in any embodiment of the invention, and has the corresponding functional modules and beneficial effects for executing the method. Further functional descriptions of the above modules and units are the same as in the corresponding embodiments described above, and will not be repeated here.
[0219] Figure 10 This is a structural schematic diagram of a vehicle provided in an embodiment of the present invention.
[0220] The following is a detailed reference. Figure 10 The diagram illustrates a structural schematic suitable for implementing a vehicle according to an embodiment of the present invention. The vehicle may include a processor (e.g., a central processing unit, graphics processor, etc.) 1001, which can perform various appropriate actions and processes according to a program stored in read-only memory (ROM) 1002 or a program loaded from memory 1008 into random access memory (RAM) 1003. The RAM 1003 also stores various programs and data required for vehicle operation. The processor 1001, ROM 1002, and RAM 1003 are interconnected via a bus 1004. An input / output (I / O) interface 1005 is also connected to the bus 1004.
[0221] Typically, the following devices can be connected to the I / O interface 1005: input devices 1006 including, for example, a touchscreen, touchpad, keyboard, mouse, camera, microphone, accelerometer, gyroscope, etc.; output devices 1007 including, for example, a liquid crystal display (LCD), speaker, vibrator, etc.; memory devices 1008 including, for example, magnetic tape, hard disk, etc.; and communication devices 1009. Communication device 1009 allows the vehicle to communicate wirelessly or wiredly with other devices to exchange data. Although Figure 10 Vehicles with various devices are shown, but it should be understood that it is not required to implement or have all of the devices shown, and more or fewer devices may be implemented or have instead.
[0222] In particular, according to embodiments of the present invention, the processes described above with reference to the flowcharts can be implemented as computer software programs. For example, embodiments of the present invention include a computer program product comprising a computer program carried on a non-transitory computer-readable medium, the computer program containing program code for performing the methods shown in the flowcharts. In such embodiments, the computer program can be downloaded and installed from a network via a communication device 1009, or installed from a memory 1008, or installed from a ROM 1002. When the computer program is executed by the processor 1001, it performs the functions defined in the methods of the embodiments of the present invention.
[0223] Figure 10 The vehicle shown is merely an example and should not be construed as limiting the functionality and scope of the embodiments of the present invention.
[0224] This invention also provides a computer-readable storage medium. The methods described above according to embodiments of the invention can be implemented in hardware or firmware, or implemented as computer code that can be recorded on a storage medium, or implemented as computer code downloaded via a network and originally stored on a remote storage medium or a non-transitory machine-readable storage medium and then stored on a local storage medium. Thus, the methods described herein can be processed by software stored on a storage medium using a general-purpose computer, a dedicated processor, or programmable or dedicated hardware. The storage medium can be a magnetic disk, optical disk, read-only memory, random access memory, flash memory, hard disk, or solid-state drive, etc.; further, the storage medium can also include combinations of the above types of memory. It is understood that computers, processors, microprocessor controllers, or programmable hardware include storage components capable of storing or receiving software or computer code, which, when accessed and executed by the computer, processor, or hardware, implements the methods shown in the above embodiments.
[0225] A portion of this invention can be applied as a computer program product, such as computer program instructions, which, when executed by a computer, can invoke or provide the methods and / or technical solutions according to the invention through the operation of the computer. Those skilled in the art will understand that the forms in which computer program instructions exist in a computer-readable medium include, but are not limited to, source files, executable files, installation package files, etc. Correspondingly, the ways in which computer program instructions are executed by a computer include, but are not limited to: the computer directly executing the instructions, or the computer compiling the instructions and then executing the corresponding compiled program, or the computer reading and executing the instructions, or the computer reading and installing the instructions and then executing the corresponding installed program. Here, the computer-readable medium can be any available computer-readable storage medium or communication medium accessible to a computer.
[0226] Although embodiments of the invention have been described in conjunction with the accompanying drawings, those skilled in the art can make various modifications and variations without departing from the spirit and scope of the invention, and such modifications and variations all fall within the scope defined by the appended claims.
Claims
1. A safety pre-control method based on electronic rearview mirror data, characterized in that, The method includes: Acquire images from the electronic rearview mirror; Extract the following vehicle from the electronic rearview mirror image and identify the state of the following vehicle's windshield; The risk of fogging in this vehicle is determined by the condition of the rear windshield. Based on the aforementioned fogging risk, a vehicle safety control strategy is generated.
2. The method according to claim 1, characterized in that, The method of determining the risk of fogging of this vehicle by the condition of the rear windshield includes: Multimodal input data for creating a model based on the state of the rear windshield; The multimodal input data is input into the fogging risk assessment model, which outputs the fogging risk level of the vehicle.
3. The method according to claim 2, characterized in that, The step of inputting the multimodal input data into the fogging risk assessment model and outputting the vehicle's fogging risk level includes: The multimodal input data is preprocessed to obtain multimodal features; The multimodal features are fused to obtain fused features; The fused features are processed by bidirectional LSTM to output state features; Multi-head attention processing is performed on the state features to obtain attention output information; The fogging risk level of the vehicle is obtained by mapping the attention output information to a risk level through a fully connected layer.
4. The method according to claim 2 or 3, characterized in that, The multimodal input data for creating the model based on the state of the rear vehicle windshield includes: Based on the condition of the rear windshield, a humidity proxy index is estimated. A consistency coefficient is calculated based on the condition of the rear vehicle's windshield to characterize the consistency of fogging phenomena in rear vehicles. The multimodal input data includes the humidity proxy index and the consistency coefficient.
5. The method according to claim 4, characterized in that, The multimodal input data for creating the model based on the state of the rear vehicle windshield also includes: Collect data from other sensors and use that data to calculate the temperature difference between the inside and outside of the vehicle; The current weather is identified using the electronic rearview mirror image and other sensor data; The current weather is coded, and an enhancement score calculated using the other sensor data is added to the weather code; Extract the historical sequence of data from the other sensors; The temperature difference between the inside and outside of the vehicle, the weather code, and the historical sequence are added to the multimodal input data.
6. The method according to claim 4, characterized in that, The humidity proxy index estimated based on the condition of the rear vehicle's windshield includes: In the formula, This indicates the humidity proxy index, the This represents the distance decay weight of the i-th distance. , This vehicle represents the distance to the i-th vehicle following it. For the preset distance threshold, This indicates the total number of vehicles inspected. This indicates the degree of fogging on the windshield of the i-th vehicle following it.
7. The method according to claim 4, characterized in that, The consistency coefficient, calculated based on the windshield condition of the rear vehicle to characterize the consistency of fogging phenomena in rear vehicles, includes: In the formula, This represents the consistency coefficient. This indicates the number of vehicles whose rear end is fogged up. This indicates the total number of vehicles inspected. The standard deviation represents the degree of fogging on the vehicle behind.
8. The method according to claim 5, characterized in that, The step of performing feature preprocessing on the multimodal input data to obtain multimodal features includes: The temperature difference between the inside and outside of the vehicle, the humidity proxy index, the consistency coefficient, and the historical sequence are standardized. One-dimensional convolution is performed on the standardized multimodal data after standardization to obtain the first multimodal feature. Weight information is assigned to each feature in the first multimodal feature using a spatial attention mechanism; The first multimodal feature, which has been assigned weight information, is subjected to time pooling to obtain the second multimodal feature.
9. The method according to claim 8, characterized in that, The step of fusing the multimodal features to obtain fused features includes: The multimodal second feature and the weighted weather code are concatenated to obtain the concatenated feature; The spliced features are processed by a fully connected layer, and some feature channels are randomly discarded during the fully connected layer processing to obtain higher-order features; The higher-order features are normalized to obtain the fused features, and the magnitude of the normalization is positively correlated with the consistency coefficient.
10. The method according to claim 3, characterized in that, The bidirectional LSTM processing of the fused features outputs state features, including: Obtain the first LSTM module and the second LSTM module; The fused features are processed sequentially by the first LSTM module and in reverse chronological order by the second LSTM module. The input gates of the first and second LSTM modules are processed as follows: In the formula, This represents the input gate value for time step t. This represents the activation function. This represents the weight matrix of the input gate. This means concatenating the hidden state from the previous time step with the input feature vector of the current time step. This represents the bias term of the input gate. The weather impact coefficient represents the magnitude of the weather impact, adjusted according to the current weather conditions. The weight matrix representing weather codes. Indicates weather codes; The hidden states output by the first LSTM module and the second LSTM module are concatenated and linearly transformed to obtain the state features.
11. The method according to claim 5, characterized in that, Extracting the following vehicle from the electronic rearview mirror image and identifying the windshield status of the following vehicle, collecting data from other sensors, and identifying the current weather using the electronic rearview mirror image and the other sensor data, including: The first intermediate feature is extracted from the electronic rearview mirror image using the embedding module and the visual Transformer module; The second intermediate feature is extracted from the electronic rearview mirror image using a convolutional neural network module and a cascaded attention module. The second intermediate feature is subjected to basic convolution processing, convolutional layer processing, and fully connected layer to obtain the third intermediate feature, and the third intermediate feature is classified to obtain the current weather. The first intermediate feature and the second intermediate feature are fused to obtain the fourth intermediate feature; The fourth intermediate feature is input into the parallel attention module to identify the following vehicle in the electronic rearview mirror image; The image of the vehicle behind is decoded to obtain the state of the windshield of the vehicle behind.
12. The method according to claim 11, characterized in that, The serial attention module includes a first module connected to the input and a second module connected to the first module; the parallel attention module includes the first module and the second module respectively connected to the input, and the outputs of the first module and the second module are superimposed. The first module includes a first branch and a second branch. The first branch is used to sequentially input the input data into a global average pooling layer and a bottleneck layer. The second branch is used to sequentially input the input data into a global max pooling layer and a basic convolutional layer. The output of the first module is the sum of the outputs of the first branch and the second branch. The second module includes a third branch and a fourth branch. The third branch is used to sequentially input the input data into a global average pooling layer and a convolutional layer. The fourth branch is used to input the input data into a convolutional layer. The output of the second module is the superposition of the outputs of the third branch and the fourth branch.
13. The method according to claim 3, characterized in that, The vehicle safety control strategy generated based on the fogging risk includes: A risk-scenario dynamic mapping relationship is defined for controlling the parameters of the vehicle air conditioning system. The risk-scenario dynamic mapping relationship is used to indicate that the higher the fogging risk level, the faster the defogging effect of the adjusted air conditioning parameters, and the lower the fogging risk level, the more the adjusted air conditioning parameters are adapted to the user's comfort index. A control table is defined based on the risk-scenario dynamic mapping relationship. The control table is used to assign corresponding air conditioning system control parameters to each fogging risk level. Acquire environmental perception data and vehicle status data. The environmental perception data includes the fogging risk level, the windshield status of the following vehicle, and the current weather. The vehicle status data includes the current air conditioning system parameters and driving scenario information. The acquired environmental perception data and vehicle status data are matched with the control table to obtain the current target air conditioning system control parameters, so as to adjust the current air conditioning parameters.
14. The method according to claim 13, characterized in that, The vehicle safety control strategy generated based on the fogging risk also includes: Monitor the defogging effect inside the vehicle, adjust the current fogging risk level according to the defogging effect, and then return to the step of adjusting the current air conditioning parameters according to each risk scenario level based on the risk-scenario dynamic mapping relationship.
15. The method according to claim 13, characterized in that, The vehicle safety control strategy generated based on the fogging risk also includes: Calculate the relative distance and relative speed between the vehicle and the vehicle behind based on the image from the electronic rearview mirror; Adjust the blind spot warning threshold according to the fog risk level; When a vehicle behind is detected entering the dynamic blind spot corresponding to the blind spot warning threshold based on the relative distance and the relative speed, an alarm message, steering intervention suggestion, and / or braking prompt are output.