A multi-modal rearview mirror automatic adjustment method and system
By collecting and processing multimodal data, a multimodal feature fusion model is constructed to dynamically adjust the rearview mirror angle, solving the problems of cumbersome rearview mirror adjustment and blind spots. This enables intelligent driver status and environmental perception, improving driving safety and comfort.
Patent Information
- Application Number
- CN202510607828.1
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Filing Date
- 2025-05-13
- Publication Date
- 2025-12-30
- Estimated Expiration
- 2045-05-13
AI Technical Summary
In existing technologies, rearview mirror adjustment is cumbersome and cannot fully reflect the driver's state and environmental changes, resulting in blind spots and poor safety. Multimodal data fusion processing is insufficient, and there is a lack of effective risk assessment models.
Multimodal data is collected, synchronized and preprocessed in time, driver state features and environmental perception features are extracted, a multimodal feature fusion model is constructed, and the rearview mirror angle is dynamically adjusted through dynamic coupling equations and risk field equations, combined with a PID controller to achieve precise adjustment.
It achieves intelligent automatic adjustment of the rearview mirror, making full use of the correlation and redundancy of multimodal data to ensure driving safety and comfort, and avoids the cumbersome and blind spot problems of traditional manual adjustment.
Smart Images

Figure CN120116850B_ABST
Abstract
Description
Technical Field
[0001] This invention relates to the field of rearview mirror adjustment, and more particularly to a multimodal automatic rearview mirror adjustment method and system. Background Technology
[0002] In existing technologies, rearview mirror adjustments are typically performed manually by the driver, a cumbersome process prone to blind spots. As automotive intelligence continues to advance, rearview mirror adjustment methods based on single-modal data can no longer meet the actual needs of drivers. Single-modal data cannot comprehensively reflect the driver's state and environmental changes, making it difficult to achieve intelligent rearview mirror adjustment.
[0003] To address these issues, a multimodal data fusion approach is needed to acquire driver status and environmental information from multiple dimensions. Driver status characteristics and environmental perception characteristics are fused into a single multimodal feature, which outputs the optimal angle for the rearview mirror, thereby enabling intelligent adjustment of the rearview mirror.
[0004] However, existing technologies have shortcomings in the fusion and processing of multimodal data, making it difficult to fully utilize the correlation and redundancy between different modal data, resulting in information loss. Simultaneously, the lack of an effective risk assessment model prevents the dynamic adjustment of the rearview mirror angle based on driver status and environmental changes, hindering the assurance of driving safety. Therefore, a novel automatic adjustment method for multimodal rearview mirrors is needed to address these technical problems. Summary of the Invention
[0005] One of the objectives of this invention is to provide a method for automatic adjustment of multimodal rearview mirrors to solve the problems existing in the fusion processing of multimodal data in the prior art.
[0006] This invention is achieved through the following technical solution: a multimodal rearview mirror automatic adjustment method, comprising the following steps: S100, collecting multimodal data, synchronizing the collected data in time, and adding a unified time tag; S200, preprocessing the multimodal data, including data noise reduction and coordinate alignment; S300, extracting driver state features and environmental perception features from the preprocessed multimodal data, wherein the driver state features are used to quantify the driver's attention, posture, and intention, and the environmental perception features are used to quantify the dynamic changes of the surrounding environment; S400, fusing the driver state features and environmental perception features into a multimodal feature, and constructing a multimodal feature fusion model to output the optimal angle of the rearview mirror.
[0007] Furthermore, the automatic adjustment method may also include S500: a terminal adjustment step, which controls the motor in the rearview mirror to adjust the angle of the rearview mirror by using the calculated optimal angle.
[0008] Furthermore, the terminal adjustment step may include the following sub-step: S510, calculating the optimal angle Decomposed into the horizontal angle that determines the coverage of the left and right fields of view of the rearview mirror. And the pitch angle that determines the vertical field of view coverage of the rearview mirror S520, Decompose the horizontal angles obtained from the decomposition. and pitch angle The signal is converted into a pulse signal, which, combined with a PID controller, adjusts the motor speed to enable the rearview mirror to accurately track the optimal angle.
[0009] Furthermore, the multimodal data is divided into driver-side data and surrounding environment-side data. Driver-side data includes: head posture information collected by a monocular RGB-D camera located in the front-view mirror, dynamic motion data of the driver collected by an IMU sensor located in the steering wheel / seat, and pressure data collected by a pressure sensor located in the seat, used for posture analysis. Surrounding environment-side data includes: radar data collected by a millimeter-wave radar located at the rear of the vehicle, used to detect vehicles or obstacles behind and provide high-precision distance and relative speed information, and LiDAR point cloud data collected by LiDAR sensors located on both sides of the rearview mirror, used to provide high-precision environmental perception capabilities and detect static and dynamic objects behind.
[0010] Furthermore, time synchronization can be achieved using PTP. A time synchronization server is deployed in the vehicle network, and all sensors are connected to the server via the network. At the same time, the deployed sensors have PTP-compatible timestamp generation modules inside, which attach a uniform timestamp to each data collection.
[0011] Furthermore, data denoising and coordinate alignment include: inferring the degree of blur in the camera image by analyzing the frequency of vehicle vibration during driving, then restoring image clarity through deconvolution technology, dynamically compensating the image, and removing high-frequency noise from the sensor through Kalman filtering; coordinate alignment involves transforming data from different sensors to the same coordinate system, unifying the data reference frame, obtaining a vehicle-based coordinate system, and mapping the pre-calibrated in-vehicle camera position and RGB-D camera intrinsic and extrinsic parameters to the vehicle coordinate system using the PnP algorithm to obtain the accurate driver's head position and orientation.
[0012] Furthermore, the driver state features include: extracting key points of the driver's head using a deep learning framework, estimating the 3D pose of the head using the PnP algorithm to calculate the head's pitch, yaw, and roll, and determining the direction of the driver's gaze by combining the calibration information from the RGB-D camera; the environmental perception features include: performing target detection using the YOLO model to identify and locate targets around the vehicle, combining the target speed detected by radar with the target contour data provided by LiDAR, and constructing complete target information using Kalman filtering technology.
[0013] Furthermore, posture stability analysis can be performed using data from pressure sensors. The specific steps are as follows: a pressure distribution heatmap of the seat is generated using pressure sensor data to show the pressure changes in different areas of the seat at different times; an SVM (Support Vector Machine) classifier is used to analyze the pressure change patterns in the heatmap to classify the driver's posture (such as leaning forward, leaning back, or leaning to the side); and the temporal changes in pressure distribution are analyzed to identify abnormal movements.
[0014] Furthermore, the modal feature fusion model includes: a dynamic coupling equation describing the evolution of features in space and time, a risk field equation for evaluating the risk level of different modalities, and a cost function for minimizing the risk field and outputting the optimal rearview mirror angle. The modal feature fusion model is constructed through the following steps:
[0015] S410. The multimodal feature is a joint state tensor, representing the feature intensity of each mode in the spatiotemporal domain, and is represented by a column vector indicating the feature intensity of different modes at different locations and times; S420. A reaction-diffusion equation is used to construct a dynamic coupling equation, which includes a diffusion term to simulate the propagation of features in space, a reaction term to describe the interaction between modes, and a source term to represent the input of an external signal or source, as shown in the following equation: ,in, The symbol is for partial differentials; Here is the diffusion matrix. This is the Laplace operator, used to describe the diffusion of features in space; For diffusion terms; This is a reaction term used to describe the nonlinear coupling relationship between modes; The source term is used to represent the input of an external signal or source; S430, the risk field equation is expressed by the following formula.
[0016] ,in, The risk field depends on time t and spatial location x. To sense the number of modes, The weighting coefficients for the i-th modal sensing mode are... This represents the importance weight of each modality. For at a point in spacetime The perceptual features of the i-th modality. This is the regularization coefficient, used to balance the weights between dynamic risks and environmental constraints; S440 is an environmental constraint term used to describe the inherent risks of the physical space; the cost function, S440, is used to minimize the comprehensive risk of the risk field and is expressed by the following formula:
[0017] ,in, For the field of view function, For angle, For smoothness penalty weight parameters, For the rearview mirror angle, To account for the rate of change of the rearview mirror angle This is a smoothing term used to penalize abrupt changes in angle, ensuring a smooth adjustment process for the rearview mirror. For all spaces inside the car, The internal integral represents the position of all spaces within the vehicle. x Summation is performed; T is the total time. For external integral, it represents the accumulation over time, indicating the vehicle's position within a time interval [0, ...]. T Operations within ]
[0018] Furthermore, a column vector can be represented by the following formula: ,in, For time variables, , representing the time domain of the driving process. T is the total driving time; For spatial coordinates, This indicates the three-dimensional spatial position inside the vehicle. It is a subset in three-dimensional space, representing a specific area inside the vehicle; Represents three-dimensional Euclidean space; Modal numbers are assigned to distinguish different modes. , It represents the total number of modes.
[0019] Furthermore, the column vector can also be represented by the following formula:
[0020] , of which each It is the characteristic intensity of mode i at time t and position x, with a value in the interval [0, 1].
[0021] Furthermore, the reaction term is expressed by the following formula:
[0022] ,in, This is the reaction matrix, describing the coupling strength between modes. The Hadamard product represents the mutual influence between modes;
[0023] Furthermore, the source term can be expressed by the following formula:
[0024] ,in, The symbol for the Dirac function. For a specific location, For a specific location The signal introduced.
[0025] Furthermore, to ensure that the optimization decision of the rearview mirror angle matches the actual visible range, a spatial visibility function can be defined. The spatial visibility function is used to describe whether a position x is within the field of view of the rearview mirror at a given angle. The formula for this spatial visibility function is:
[0026] Specifically, the field of view of the rearview mirror can be modeled as a three-dimensional viewing cone, the geometry of which is determined by the curvature of the mirror surface and the adjustment angle. and The decision is made by using a ray tracing algorithm to determine whether a point x in space lies within the viewing cone. The boundary of the viewing cone can be calculated based on the curvature and angle of the rearview mirror. Ray tracing is then used to determine if x falls within this region. If point x lies within the viewing cone, then... This indicates that the position is within the field of view of the rearview mirror; if it is not within the visual cone, then... .
[0027] Another aspect of the invention provides a multimodal rearview mirror automatic adjustment system, including a memory, a processor, and a computer program stored in the memory and executable on the processor, wherein the processor executes the program to implement the multimodal rearview mirror automatic adjustment method as described above.
[0028] Compared with the prior art, the present invention has the following advantages and beneficial effects:
[0029] 1. This invention, by collecting multimodal data and synchronizing it in time, can obtain driver status and environmental information from multiple dimensions, comprehensively reflect changes in the driver's status and environment, and provide data support for intelligent adjustment of rearview mirrors.
[0030] 2. This invention constructs a multimodal feature fusion model, which integrates driver state features and environmental perception features into a single multimodal feature. This fully utilizes the correlation and redundancy between various modal data, avoids information loss, and effectively describes the evolution of features in time and space, evaluates the risk level of different modalities, minimizes the optimal angle of the rearview mirror output by the risk field, and dynamically adjusts the rearview mirror angle according to changes in driver state and environment to ensure driving safety.
[0031] 3. By introducing a smoothing term, this invention can punish abrupt changes in the rearview mirror angle, ensuring the smoothness of the rearview mirror adjustment process and improving driving comfort. It realizes intelligent adjustment of the rearview mirror, eliminating the need for manual operation by the driver. The adjustment process is automated, avoiding the cumbersome and blind spot problems of traditional manual adjustment, and improving driving convenience and safety. Attached Figure Description
[0032] The accompanying drawings, which are included to provide a further understanding of embodiments of the invention and form part of this application, do not constitute a limitation thereof. In the drawings:
[0033] Figure 1 This is a flowchart of the method provided in Embodiment 1 of the present invention. Detailed Implementation
[0034] To make the objectives, technical solutions, and advantages of the embodiments of the present invention clearer, the technical solutions of the embodiments of the present invention will be clearly and completely described below with reference to the accompanying drawings. Obviously, the described embodiments are only some embodiments of the present invention, and not all embodiments. The components of the embodiments of the present invention described and shown in the accompanying drawings can generally be arranged and designed in various different configurations.
[0035] Example 1
[0036] This embodiment discloses a multimodal rearview mirror automatic adjustment method and system to address the shortcomings of existing technologies in the fusion processing of multimodal data. These technologies struggle to fully utilize the correlation and redundancy between different modal data, leading to information loss. Furthermore, the lack of an effective risk assessment model prevents dynamic adjustment of the rearview mirror angle based on driver status and environmental changes, thus hindering driving safety.
[0037] Figure 1 The flowchart of the method in this embodiment is shown. As can be seen from the flowchart, this embodiment includes the following steps:
[0038] S100: Acquires multimodal data and establishes a foundational layer for synchronous acquisition of multimodal data, ensuring that data from all sensors can be aligned under a unified time reference, thereby providing high-quality input for subsequent data fusion and analysis.
[0039] In this embodiment, the perception of the driver's head posture and the rear environment is coordinated. The driver's head posture is captured by a camera, the IMU sensor provides dynamic changes in the steering wheel and seat, while millimeter-wave radar and LiDAR provide the position and motion information of objects behind. By precisely synchronizing the collected data using the PTP protocol, all sensors provide a unified timestamp, and the collected data is stored in a circular buffer. This allows the high-frequency data from the IMU sensor to be aligned with the low-frequency data from the camera through interpolation, ensuring that each frame of sensor data can be processed on the same time basis.
[0040] Through the above hardware deployment and multimodal data synchronization, accurate and consistent driver behavior and environmental perception data can be obtained, providing a solid foundation for subsequent analysis and decision-making.
[0041] Specifically, the collection of multimodal data can be divided into driver-side and surrounding environment-side data.
[0042] Multimodal data on the driver's side can be acquired by a monocular RGB-D camera installed at the position of the in-vehicle front mirror and by IMU sensors arranged in the steering wheel / seat.
[0043] A monocular RGB-D camera captures the driver's head posture information, providing comprehensive head position and posture information through the acquired RGB images and depth data. During installation, to ensure wide-angle coverage, it is best to choose a monocular RGB-D camera with a 120-degree field of view to ensure that the driver's head movements, especially dynamic changes in the face and head, are captured.
[0044] IMU sensors within the steering wheel / seat are used to capture the driver's dynamic motion data, particularly through a combination of accelerometers and gyroscopes, to monitor the displacement and angular velocity of the steering wheel and seat. IMU sensors can be mounted on the driver's seat back or integrated into the steering wheel to capture the driver's hand movements and operational intentions, while the seat IMU provides information about the driver's posture and seat adjustments.
[0045] Pressure sensor data can also be installed inside the seat for posture analysis.
[0046] Multimodal data on the surrounding environment can be obtained through millimeter-wave radar installed on the rear bumper of the vehicle and LiDAR sensors deployed on both sides of the rearview mirror.
[0047] Millimeter-wave radar is used to detect vehicles or obstacles behind, providing high-precision distance and relative speed information.
[0048] LiDAR sensors are used to provide high-precision environmental perception capabilities, helping to detect static and dynamic objects behind the camera. LiDAR can provide higher spatial resolution, making it suitable for accurate perception in complex scenes.
[0049] After acquiring multimodal data, to ensure that data from different sensors can be aligned under a unified time reference, PTP is used for time synchronization in this embodiment. A time synchronization server is deployed in the vehicle network, and all sensors connect to this server via the network. The deployed sensors contain PTP-compatible timestamp generation modules that attach a unified timestamp to each acquired data set.
[0050] Given the different sampling frequencies of various sensors, ensuring data from different sources are aligned at the same time point is a critical issue. In this embodiment, a circular buffer can be used to buffer sensor data, especially for aligning high-frequency data (such as IMU) with low-frequency data (such as camera data). In complex environments, IMU sampling frequencies are typically high, potentially reaching 100Hz, while camera sampling frequencies are lower, typically 50Hz or lower. This is achieved by importing all data streams (e.g., IMU, camera, radar) into a shared buffer. For high-frequency sampled IMU data, interpolation algorithms (such as linear interpolation or spline interpolation) are used to adjust it to the time point corresponding to the low-frequency data. This ensures all data can be aligned at the same time point, thus avoiding errors caused by time differences.
[0051] S200: Preprocesses the acquired multimodal data, including data denoising and coordinate alignment. Data from different sensors often contains noise, especially in dynamic or high-frequency vibration environments. Therefore, effective denoising is necessary to improve data quality and accuracy.
[0052] Specifically, in this embodiment, considering that vibrations during vehicle operation may cause blurring in the images captured by the camera, affecting the accuracy of visual perception, the degree of blurring in the camera image can be inferred by analyzing the frequency of vibrations during vehicle operation. Then, deconvolution technology can be used to restore the image clarity and dynamically compensate for the image.
[0053] Data acquired by IMU sensors is often affected by high-frequency noise from vehicle vibrations or steering wheel hand operations, especially high-frequency vibrations in gyroscope data. In this embodiment, Kalman filtering can be used to remove this high-frequency noise. Kalman filtering is a recursive estimation method based on linear dynamic systems. By filtering the acceleration and angular velocity data from the IMU sensor, it can effectively remove random noise from the sensor and retain true dynamic information.
[0054] Typically, in multimodal data fusion, data from different sensors are represented using different coordinate systems. To unify the data reference framework, it is necessary to transform the data from different sensors to the same coordinate system. In this embodiment, pixel coordinates (for images), polar coordinates (for radar), point cloud coordinates (for LiDAR), etc., need to be transformed into a unified vehicle coordinate system (based on the right front upper reference).
[0055] Specifically, the coordinates of camera data are typically pixel coordinates of the image. During the transformation, the conversion from pixel coordinates to the vehicle coordinate system can be performed using the camera's intrinsic and extrinsic parameters (such as the camera matrix and rotation / translation matrices). This is usually achieved through camera calibration, and in this embodiment, the calibration process can use the PnP algorithm to calculate the camera position and orientation.
[0056] Radar data is typically represented in polar coordinates (distance, angle), which needs to be converted to a Cartesian coordinate system (X, Y, Z coordinates) for fusion with other sensor data. LiDAR data, on the other hand, is usually point cloud data, requiring conversion of the point cloud coordinates to the vehicle coordinate system based on the LiDAR's installation location and orientation. This coordinate transformation can be achieved using rotation and translation matrices based on the LiDAR's rotation angle and in-vehicle calibration information.
[0057] Finally, relying on the pre-calibrated in-vehicle camera position and the RGB-D camera's internal and external parameters, the driver's head posture information is mapped to the vehicle coordinate system using the PnP algorithm to obtain the accurate driver's head position and orientation, which is convenient for subsequent analysis.
[0058] S300: Extract driver state features and environmental perception features from preprocessed multimodal data to better quantify driver behavior, intentions, and dynamic changes in the surrounding environment.
[0059] Specifically, driver state characteristics are used to quantify driver attention, posture, intentions, and other state information. This process involves extracting key points of the driver's head, posture information, and gaze direction, and combining this with pressure sensor data to perform posture analysis.
[0060] Key points of the driver's head can be extracted using existing deep learning frameworks, and the 3D head pose can be estimated using the PnP algorithm. Combining the estimated 3D pose information, we can calculate: head pitch, i.e., the angle of rotation of the head up or down, representing the degree of head tilt; head yaw, the angle of rotation of the head left or right, representing the deviation of the head to the left or right; and head roll, the angle of rotation of the head around its own axis, typically representing head tilt. By extracting the head key points and combining them with the calibration information of the RGB-D camera and the iris position of the eyes, the direction of the driver's gaze can be determined.
[0061] The following are the specific steps for analyzing sitting posture stability using data from seat pressure sensors:
[0062] A pressure distribution heatmap of the seat is generated using pressure sensor data, which shows the pressure changes in different areas of the seat at different times.
[0063] The driver's sitting posture (such as leaning forward, leaning back, or leaning to the side) is classified by analyzing the pressure change patterns in the heat map using an SVM (Support Vector Machine) classifier.
[0064] Analyzing the temporal changes in pressure distribution can help identify abnormal movements (such as frequent changes in sitting posture, rapid forward leaning, etc.), which may indicate some abnormal behavior of the driver (such as anxiety, fatigue, or improper driving behavior).
[0065] Environmental perception features are used to capture dynamic environmental information behind the vehicle and in blind spots, extract dynamic information about risk points and targets in the surrounding environment, and provide timely warnings to the driver.
[0066] Use the YOLO model for object detection to identify and locate targets around the vehicle (such as other vehicles, pedestrians, etc.).
[0067] By combining the target velocity detected by radar with the target contour data provided by LiDAR, Kalman filtering technology can be used to construct complete information about the target (such as target position, shape, velocity, etc.).
[0068] S400: After extracting driver state features and environmental perception features, these two features are fused into a multimodal feature, and a multimodal feature fusion model is constructed to describe the dynamic coupling of multimodal features (such as vision, pressure, radar, etc.).
[0069] Meanwhile, a decision model based on risk field optimization combines multimodal characteristics with decision objectives to assess the risk level at different spatial locations.
[0070] Specifically, the multimodal feature fusion model is constructed through the following steps:
[0071] 1) First, define the multimodal feature fusion feature representation form, clarify the spatiotemporal evolution law of its mathematical expression and remodel the features, and describe how the tensor changes in time and space through partial differential equations.
[0072] The multimodal feature fusion feature is a joint state tensor that represents the feature intensity of each modality in the spatiotemporal domain.
[0073] Specifically, it can be expressed as This is a column vector representing the feature intensity of different modalities at a given location and time. For example, the perceived values of the visual modality, pressure modality, and radar modality at that location and time. Specifically:
[0074] For time variables, , represents the time domain of the driving process. T is the total driving time. For spatial coordinates, This indicates the three-dimensional spatial position inside the vehicle. It is a subset in three-dimensional space, representing a specific area inside the vehicle; This represents three-dimensional Euclidean space, which contains all three-dimensional coordinate points and can be represented by three real numbers. Modal numbers are assigned to distinguish different modes. , It represents the total number of modes.
[0075] This column vector can also be represented by the following formula:
[0076] , of which each It is the characteristic intensity of mode i at time t and position x, with a value in the interval [0, 1].
[0077] 2) In this embodiment, the dynamic coupling of multimodal features is described using a reaction-diffusion equation. This equation includes a diffusion term to simulate the propagation of features in space, a reaction term to describe the interaction between modes, and a source term to represent the input of an external signal or source. This equation describes how the tensor changes in time and space. Specifically, the dynamic coupling relationship of the multimodal features can be expressed by the following equation:
[0078] ,
[0079] in, The symbol is for partial differentials; The diffusion matrix represents the propagation capability of each modal feature in space. The propagation coefficients in the diffusion matrix can be preset with empirical values based on the physical characteristics of the sensor (e.g., radar data has a higher diffusion coefficient, while LiDAR has a lower one). , Represents a diagonal matrix. The spatial propagation coefficient for a certain mode; This is the Laplace operator, used to describe the diffusion of features in space; This is a diffusion term.
[0080] The reaction term describes the nonlinear coupling relationship between modes. The reaction term can be expressed as follows: ,in, This is the response matrix, describing the coupling strength between modes. For example, A 1,2 A value of >0 can indicate a positive correlation between visual and stress modalities, such as the joint correction of gaze credibility by head posture and sitting posture. The Hadamard product (element-wise product) represents the mutual influence between modes.
[0081] The source term represents an external signal or input to a source, such as the raw signal captured by a sensor. A source term can be represented by the following formula:
[0082] ,
[0083] in, k The symbol representing the position of the source term The symbol for the Dirac function. The position represented by the source term. The position represented by the source term The introduced signal, this formula is used at a specific location Signal introduced at the location This represents the sensor's position. The input signal.
[0084] 3) The reaction-diffusion equation describes the evolution of features in space and time, providing dynamic coupling relationships of dynamic and multimodal modes. A risk field is constructed using data of dynamic coupling relationships, and the overall risk is minimized through optimization decisions.
[0085] The risk field, by fusing multimodal features, transforms abstract sensor data into a spatial risk heatmap, quantifying real-time risk values in different areas inside and outside the vehicle. This includes dynamic risks arising from direct threats from driver state (e.g., distraction, frequent changes in seating position) and environmental perception (e.g., approaching vehicles from behind, blind spot obstacles), as well as inherent risks represented by environmental constraints (e.g., blind spots in rearview mirrors, lane edges).
[0086] Specifically, the risk field can be represented by the following formula:
[0087] ,
[0088] in, The risk field depends on time t and spatial location x, and combines features from multiple perceptual modalities with some spatial constraints. To detect the number of modes, each mode has a corresponding sensing information. The weighting coefficient for the i-th modality perception mode represents the importance of that mode in the overall decision-making process. This represents the importance weight of each modality, which can be adaptively adjusted based on the fusion features of the modalities using the Sigmoid function. For at a point in spacetime The perceptual features of the i-th modality. For example: for the visual modality, it might be image data captured by a camera; for the pressure modality, it might be the magnitude of pressure sensed on the seat; for the radar modality, it might be distance or object detection information from radar; these perceptual features It describes the perceptual information of different modalities at specific locations and times. This is the regularization coefficient, used to balance the weights between dynamic risks and environmental constraints. This is an environmental constraint term used to describe the inherent risks of a physical space. For example, the risk of a blind spot in a rearview mirror can be very high, indicating that there is a high risk in that area.
[0089] The goal of optimization decision-making is to minimize the overall risk of the risk field. In the scenario of rearview mirror angle adjustment, the objective is to minimize the following cost function:
[0090]
[0091] in, Let be the field of view function, representing the angle . ,Location Visibility. Indicates that it is not visible. Indicates that it is fully visible. For smoothness penalty weight parameters, For the rearview mirror angle, To account for the rate of change of the rearview mirror angle, if A large size means the rearview mirror adjusts very quickly, which may cause discomfort. If A smaller size means the rearview mirror adjusts smoothly, making the driver feel more comfortable. The value is the square of the rate of change of the rearview mirror angle, which means that if the adjustment speed is fast, the cost will increase. This is a smoothing term used to penalize abrupt changes in angle, ensuring a smooth adjustment process for the rearview mirror. For all spaces inside the car, The internal integral represents the position of all spaces within the vehicle. x Summation is performed; T is the total time. For external integral, it represents the accumulation over time, indicating the vehicle's position within a time interval [0, ...]. T Operations within ]
[0092] S500: The optimal angle of the rearview mirror is obtained by calculating and minimizing the cost function. Based on this optimal angle, the motor in the rearview mirror is controlled to adjust the angle of the rearview mirror.
[0093] Specifically, this may include the following sub-steps:
[0094] 1) The optimal angle obtained in S400 It includes decomposition in the horizontal and pitch directions, namely, the horizontal angle that determines the coverage of the left and right fields of view of the rearview mirror. The pitch angle determines the vertical field of view covered by the rearview mirror. .
[0095] 2) To ensure that the optimization decision of the rearview mirror angle matches the actual visible range, a spatial visibility function needs to be defined. , is used to describe whether a certain position x is within the field of view of the rearview mirror at a given angle.
[0096] In this embodiment, the spatial visibility function is defined as follows:
[0097] ,
[0098] Specifically, the field of view of the rearview mirror can be modeled as a three-dimensional viewing cone, the geometry of which is determined by the curvature of the mirror surface and the adjustment angle. and The decision is made by using a ray tracing algorithm to determine whether a point x in space lies within the viewing cone. The boundary of the viewing cone can be calculated based on the curvature and angle of the rearview mirror. Ray tracing is then used to determine if x falls within this region. If point x lies within the viewing cone, then... This indicates that the position is within the field of view of the rearview mirror; if it is not within the visual cone, then... .
[0099] 3) Based on the optimal angle, the physical adjustment of the rearview mirror needs to be achieved through motor control, based on the horizontal angle obtained from the decomposition. and pitch angle This is converted into a pulse signal, which drives the stepper motor to adjust the angle. To ensure the rearview mirror accurately tracks the optimal angle and avoids excessive adjustment or vibration, a PID controller can be used to adjust the motor speed.
[0100] The specific embodiments described above further illustrate the purpose, technical solution, and beneficial effects of the present invention. It should be understood that the above description is only a specific embodiment of the present invention and is not intended to limit the scope of protection of the present invention. Any modifications, equivalent substitutions, improvements, etc., made within the spirit and principles of the present invention should be included within the scope of protection of the present invention.
Claims
1. A multi-modal rearview mirror automatic adjustment method, characterized in that, The rearview mirror automatic adjusting method comprises, S100, collecting multi-modal data, time synchronizing the collected data, and adding a unified time label; S200, preprocessing the multi-modal data, The preprocessing comprises data denoising and coordinate alignment; S300, extracting driver state features and environment perception features from the preprocessed multi-modal data, The driver state features are used for quantifying the attention, sitting posture, and intention of the driver, The environment perception features are used for quantifying the dynamic changes of the surrounding environment; S400, fusing the driver state features and the environment perception features into one multi-modal feature, and constructing a multi-modal feature fusion model to output the optimal angle of the rearview mirror; The multi-modal feature fusion model comprises: a dynamic coupling equation for describing the evolution of features in space-time, a risk field equation for evaluating the risk degree of different modalities, and a cost function for minimizing the risk field and outputting the optimal angle of the rearview mirror, The multi-modal feature fusion model is constructed through the following steps: S410, the multi-modal feature is a joint state tensor, representing the feature intensity of each modality in the space-time domain, and the feature intensity of different modalities at a position and time is represented by a column vector; S420, a reaction-diffusion equation is used to construct a dynamic coupling equation, which comprises a diffusion term for simulating the propagation of features in space, a reaction term for describing the interaction between modalities, and a source term for representing the input of external signals or sources, as shown in the following formula: , wherein, is a partial differential symbol; is a diffusion matrix, is a Laplacian operator for describing the diffusion of features in space; is a diffusion term; is a reaction term for describing the nonlinear coupling relationship between modalities; is a source term for representing the input of external signals or sources, is a multi-modal feature; S430, the risk field equation is represented by the following formula , wherein, is the risk field, depending on time t and spatial location , is the number of perception modalities, is the weighting coefficient of the i-th modality of perception modalities, represents the importance weight of each perception modality, is the perception feature of the i-th modality at the spatio-temporal point , is the regularization coefficient for balancing the weight between dynamic risk and environmental constraints; is the environmental constraint term for describing the inherent risk of the physical space; S440, the cost function is used to minimize the comprehensive risk of the risk field, and is represented by the following formula: , wherein, is a field of view function, is an angle, is a smoothness penalty weight parameter, is a mirror angle, is a rate of change of the mirror angle, is a smoothness term to penalize abrupt changes in the angle, ensuring smoothness of the mirror adjustment process, is a subset of a three-dimensional space, is an inner integral, representing a summation over all spatial positions in the subset of the three-dimensional space is a total driving time, is an outer integral, representing a summation over time, representing the operation of the vehicle over a time period [0, T].
2. The method of claim 1, wherein, The rearview mirror automatic adjusting method further comprises a terminal adjustment step S500, which controls the motor in the rearview mirror to adjust the angle of the rearview mirror through the calculated optimal angle.
3. The method of claim 1, wherein, The multi-modal data is divided into driver side data and surrounding environment side data, wherein, The driver side data comprises: head posture information collected by a monocular RGB-D camera arranged at the front mirror position in the vehicle, dynamic motion data of the driver collected by an IMU sensor arranged in the steering wheel or the seat, and pressure data collected by a pressure sensor arranged in the seat for sitting posture analysis; The surrounding environment side data comprises: radar data collected by a millimeter wave radar arranged at the rear of the vehicle, for detecting rear vehicles or obstacles, and providing high-precision distance and relative speed information, LiDAR point cloud data collected by LiDAR sensors arranged on both sides of the rearview mirror, for providing high-precision environment perception capability and detecting static and dynamic objects at the rear.
4. The method of claim 1, wherein, The data denoising and coordinate alignment comprises: by analyzing the frequency of vibration of the vehicle during driving, the blur degree in the camera image is inferred, and then the image clarity is restored through deconvolution technology, and the image is dynamically compensated, high-frequency noise of the sensor is removed through Kalman filtering; The coordinate alignment unifies the reference frame of data by converting data of different sensors to the same coordinate system, obtains a vehicle-based coordinate system, maps the pre-calibrated in-vehicle camera position and the RGB-D camera internal and external parameters to the vehicle coordinate system through a PnP algorithm, and obtains an accurate driver head position and orientation.
5. The method of claim 1, wherein, The driver state features include: The key points of the driver head are extracted through a deep learning framework, the 3D pose of the head is estimated through a PnP algorithm, the pitch of the head, the yaw of the head and the roll of the head are calculated, and the direction of the driver line of sight is determined in combination with the calibration information of the RGB-D camera. The environment perception features include: target detection is performed through a YOLO model, targets around the vehicle are identified and located, the target speed detected by the radar is combined with the target contour data provided by the LiDAR, and the complete information of the target is constructed through Kalman filtering technology.
6. The method of claim 2, wherein, The step S500 includes the following sub-steps: S510, the calculated optimal angle is decomposed into a horizontal angle determining the coverage of the left and right fields of view of the rearview mirror and a pitch angle determining the coverage of the up and down fields of view of the rearview mirror ; S520, Decompose the horizontal angles obtained from the decomposition and pitch angle The signal is converted into a pulse signal, which, combined with a PID controller, adjusts the motor speed to enable the rearview mirror to accurately track the optimal angle.
7. The method of claim 1, wherein, The column vector is represented by the following formula: wherein, is a time variable, , denotes the time domain of the driving process, T is the total driving time; is a spatial position, denotes a three-dimensional spatial position within the car, ; is a subset of the three-dimensional space, denotes a certain specific area within the car; denotes the three-dimensional Euclidean space; is a modal number, used to distinguish different modalities, , is the total number of modalities.
8. The multi-modal rearview mirror automatic adjustment method of claim 1, wherein, The reaction term is represented by the following formula: , wherein, is the reaction matrix, describing the coupling strength between the modes, is the Hadamard product, indicating the mutual influence between the modes; The source term is represented by the following formula: , where k is the sign of the position represented by the source term, is the Dirac function symbol, is the position represented by the source term, is the position represented by the source term is the incoming signal.
9. A multi-modal rearview mirror automatic adjusting system, characterized by, The rearview mirror automatic adjustment system includes: A processor; A memory storing a computer program, when the computer program is executed by the processor, the multi-modal rearview mirror automatic adjustment method of any one of claims 1-8 is implemented.
Citation Information
Patent Citations
Control system and method for adaptively adjusting position of outside rear-view mirror based on DMS
CN119659470A