Method and system for optimizing human-machine interface interactive experience of vehicle-mounted display terminal
By analyzing the driver's interaction behavior data and vehicle environment information, dynamically adjusting the functional module layout of the on-board display terminal interface, the driver's visual burden and attention distraction caused by the static interface design in the prior art is solved, and a more efficient and safe human-computer interaction experience is achieved.
Patent Information
- Application Number
- CN202510466428.3
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Filing Date
- 2025-04-15
- Publication Date
- 2025-06-06
- Estimated Expiration
- 2045-04-15
AI Technical Summary
The human-computer interface of the existing vehicle display terminal has a static fixed layout design, which cannot be dynamically adjusted to adapt to the driver's usage habits, resulting in increased visual burden, distracted attention, and low access efficiency of functional modules.
By obtaining the driver's interactive behavior data, an attention distribution heat map is generated, the importance score of the functional module is calculated, and combining the vehicle's operating status and environmental perception information, dynamically identify the driving scene, and adjust the display position and touch response area of the functional module.
It realizes a human-computer interactive experience that is more in line with the driver's habits, reduces the driver's operating burden and visual search time, and improves driving safety and timeliness of system response.
Smart Images

Figure CN119987609B_ABST
Abstract
Description
Technical Field
[0001] The present invention relates to the technical field of vehicle-mounted human-computer interaction, and in particular to a method and system for optimizing the human-computer interface interaction experience of a vehicle-mounted display terminal. Background Art
[0002] With the rapid development of intelligent connected vehicles, in-vehicle display terminals have become an important interactive interface for drivers to obtain information and control vehicle functions. Traditional in-vehicle display terminals provide drivers with diverse functions such as navigation, entertainment, and vehicle control through touch screens, buttons, and other interactive methods. At present, the human-computer interaction design of in-vehicle display terminals pays more and more attention to the driver's usage habits and cognitive characteristics, and optimizes the interface layout and interaction method by analyzing the driver's interactive behavior data to improve driving safety and operational convenience.
[0003] However, there are some problems with the existing human-machine interface of the in-vehicle display terminal. The interface layout often adopts a static and fixed design scheme, which cannot be dynamically adjusted according to the driver's actual usage habits; the arrangement of functional modules lacks consideration for the driver's attention allocation, which easily causes visual burden and distraction to the driver; the differences in the usage requirements of functional modules in different driving scenarios are not fully considered, resulting in inefficient access to some frequently used functional modules; the correlation between functional modules is not reasonably utilized, which reduces the consistency and fluency of interface operations.
[0004] In summary, there is an urgent need for a vehicle-mounted display terminal human-machine interface optimization method based on driver interaction behavior analysis. Through multi-dimensional interaction behavior data collection and fusion analysis, a driver's attention distribution model is established; combined with vehicle operation status and environmental perception information, dynamic recognition of driving scenes is achieved; based on the importance score of functional modules and usage probability prediction, the interface layout is optimized and adjusted in real time. The present invention can provide drivers with a more habitual, safer and more efficient human-machine interaction experience. Summary of the invention
[0005] The embodiments of the present invention provide a method and system for optimizing the human-machine interface interactive experience of an in-vehicle display terminal, which can solve the problems in the prior art.
[0006] According to a first aspect of the embodiments of the present invention,
[0007] A method for optimizing the human-machine interface interaction experience of an in-vehicle display terminal is provided, comprising:
[0008] Acquiring interactive behavior data of the driver, wherein the interactive behavior data includes eye gaze point position data, hand operation trajectory data, and head rotation angle data;
[0009] The eye gaze position data and the head rotation angle data are fused and calculated, and the weights are adjusted in combination with the hand operation trajectory data to generate a heat map of the driver's attention distribution;
[0010] Based on the heat value of each area in the attention distribution heat map, calculate the importance score of each functional module of the vehicle display terminal;
[0011] Collect vehicle operation status information and environmental perception information, build a dynamic scene recognition model, use the deep learning network to identify the current driving scene type in real time, and output the usage probability prediction value of the functional module;
[0012] The importance score and the usage probability prediction value are weighted and integrated to obtain the functional module interaction experience optimization index;
[0013] According to the preset index threshold, the corresponding functional modules are adjusted to the hot spots of the attention distribution heat map, and grouped and clustered based on the correlation of the functional modules, and the touch response area and display level of the functional modules are dynamically adjusted;
[0014] When the change in the interactive behavior data exceeds a preset change threshold, the optimization step is triggered to be re-executed.
[0015] In an optional embodiment,
[0016] The eye gaze point position data and the head rotation angle data are fused and calculated, and the weights are adjusted in combination with the hand operation trajectory data to generate a heat map of the driver's attention distribution, including:
[0017] The pupil center coordinates are obtained based on the driver's eye gaze point position data, the high-frequency jitter data in the pupil center coordinates are eliminated by Kalman filtering, and the pupil center coordinates are smoothed by using a RANSAC algorithm to obtain continuous gaze point trajectory data;
[0018] According to the yaw angle and translation vector in the head rotation angle data, the continuous gaze point trajectory data is compensated for spatial coordinates to obtain gaze point coordinate compensation data;
[0019] Through the hand positioning points in the hand operation trajectory data, the three-dimensional space coordinate sequence is extracted, the hand movement speed, movement acceleration, movement direction angle, displacement distance and gesture type are calculated, and the hand trajectory feature vector is constructed;
[0020] Using a dynamic time warping algorithm, the gaze point coordinate compensation data, the head rotation angle data, and the hand trajectory feature vector are time-series aligned to construct a multimodal feature tensor;
[0021] Calculating the weight coefficient of each feature in the multimodal feature tensor, and combining the weight coefficient with a Gaussian kernel function to generate attention density distribution data;
[0022] Within a preset time window, the density feature vector of the attention density distribution data is extracted and time series fusion is performed through exponential decay weights. A dual-domain regularization function is constructed to eliminate abnormal points. The initial hot zone boundary point set is determined based on an adaptive segmentation threshold, and the boundary is guided to dynamically expand to generate a driver's attention distribution heat map.
[0023] In an optional embodiment,
[0024] In a preset time window, the density feature vector is extracted from the attention density distribution data and time series fusion is performed through exponential decay weights, a dual-domain regularization function is constructed to eliminate abnormal points, an initial hot zone boundary point set is determined based on an adaptive segmentation threshold, and the boundary is guided to dynamically expand, and a driver's attention distribution heat map is generated, including:
[0025] The statistical features and spatial moment features of the attention density distribution data are extracted within the preset time window to construct a density feature vector. The density feature vectors of the historical window and the current window are fused through the exponential decay weight to obtain a time series fusion feature vector.
[0026] Determine the coordinates of the attention center point according to the time series fusion feature vector, construct a dual-domain regularization function containing spatial distance constraints and density value constraints, and perform regularization processing to obtain the optimized data of attention density distribution;
[0027] Calculate the adaptive segmentation threshold based on the attention density distribution optimization data and determine the initial hot zone boundary point set;
[0028] Calculate the density gradient direction and density value of the initial hot zone boundary point set, and set the main threshold and secondary threshold in the corresponding eight neighborhoods along the density gradient direction;
[0029] When the density value of the eight neighborhoods is greater than the main threshold, the corresponding points are included in the initial hot zone boundary point set;
[0030] When the density value of the eight neighborhoods is between the primary threshold and the secondary threshold, the corresponding points are marked as candidate boundary point sets; the curvature value of the initial hot zone boundary point set is calculated to construct a dynamic scaling function, and the primary threshold and the secondary threshold are adjusted; the shortest path distance from the candidate boundary point set to the initial hot zone boundary point set is calculated, and the candidate boundary point set whose shortest path distance is less than the preset connectivity threshold is merged into the initial hot zone boundary point set;
[0031] When the area growth rate of the initial hot zone boundary point set is less than the preset growth threshold, and the number of candidate boundary point sets is less than the preset total number threshold, the boundary expansion is stopped to obtain the expanded hot zone boundary point set;
[0032] The kernel density estimation and normalization processing are performed on the boundary point set of the expansion hot zone to obtain the driver's attention distribution heat map.
[0033] In an optional embodiment,
[0034] Collect vehicle operation status information and environmental perception information, build a dynamic scene recognition model, and use the deep learning network to identify the current driving scene type in real time. The output function module usage probability prediction values include:
[0035] Acquire vehicle operation status information and environmental perception information, calculate the volatility of the vehicle operation status information and environmental perception information based on a sliding time window, and determine the volatility by the ratio of the change amplitude of the characteristic values at adjacent time points to the average change amplitude in the time window; divide the scene sequence into a steady-state interval and a mutation interval according to a preset fluctuation threshold; perform feature sampling of different frequencies on the steady-state interval and the mutation interval, respectively, to obtain an adaptive sampling feature sequence;
[0036] Matching the adaptive sampling feature sequence with historical scene samples, constructing and solving the scene evolution equation based on the element coupling relationship in the historical scene samples and the stability of the scene prototype, obtaining the state association matrix, and dynamically updating the dynamic coefficient using the prediction deviation;
[0037] In the adaptive sampling feature sequence, a sampling moment whose prediction deviation is greater than a preset deviation threshold is selected as a key scene frame; the features of the key scene frame are input into a pre-trained scene classifier to obtain a scene type recognition result; and based on the scene type recognition result and the prediction deviation, a usage probability prediction value of the corresponding functional module is calculated.
[0038] In an optional embodiment,
[0039] The adaptive sampling feature sequence is matched with the historical scene samples, the scene evolution equation is constructed and solved based on the element coupling relationship in the historical scene samples and the stability of the scene prototype, the state association matrix is obtained, and the dynamic coefficient is dynamically updated using the prediction deviation, including:
[0040] Extracting scene elements from the historical scene samples, wherein the scene elements include vehicles, pedestrians, and road facilities; constructing an element coupling matrix, wherein each element of the element coupling matrix represents an interaction influence intensity value between adjacent scene elements;
[0041] Clustering is performed based on the historical scene samples to obtain a plurality of scene prototypes; obtaining the eigenvalue of each scene prototype in the element coupling matrix, and calculating the stability index of each scene prototype in combination with the spatial distribution density of the scene elements;
[0042] Constructing a scenario evolution equation including a steady-state term and a disturbance term, wherein the coefficient of the steady-state term comes from the stability index, and the coefficient of the disturbance term comes from the element coupling matrix; solving the scenario evolution equation to obtain the state transition probability between scenario prototypes, and constructing a state association matrix;
[0043] The adaptive sampling feature sequence is input into the state association matrix for feature prediction to obtain predicted features; the predicted features are compared with the actual features to calculate the prediction deviation, and based on the prediction deviation, the interaction influence intensity value in the element coupling matrix and the disturbance term coefficient in the scenario evolution equation are updated respectively.
[0044] In an optional embodiment,
[0045] According to the preset index threshold, the corresponding functional modules are adjusted to the hot spot area of the attention distribution heat map, and grouped and clustered based on the correlation of the functional modules. The touch response area and display level of the functional modules are dynamically adjusted, including:
[0046] Receive the functional module interaction experience optimization index and the attention distribution heat map, and select the functional modules to be adjusted from the functional module interaction experience optimization index based on a preset index threshold; divide the attention distribution heat map into multiple regional levels, and sort the functional modules by importance according to the functional module interaction experience optimization index; extract the interactive operation data of the functional module, including the operation time series, the operation interval duration, and the operation transfer probability;
[0047] Acquire the hierarchical relationship tree of the functional modules, calculate the shortest path distance between the functional modules, and construct the initial association strength based on the shortest path distance; calculate the interaction migration probability between the functional modules according to the interaction operation data, perform weighted combination of the interaction migration probability and the initial association strength, and generate a dynamic association matrix;
[0048] Performing eigendecomposition on the dynamic association matrix to obtain a main eigenvector; adaptively clustering the functional modules based on the main eigenvector to determine a functional module group; calculating the intra-group association density and inter-group association sparsity of each functional module group;
[0049] The functional module group is projected to the regional level of the attention distribution heat map, and the projection position is jointly determined by the intra-group association density and the functional module interaction experience optimization index; the relative layout spacing of the functional module group is determined based on the inter-group association sparsity; according to the position of the functional module in the attention distribution heat map, the touch response area and display level are dynamically adjusted.
[0050] In an optional embodiment,
[0051] Calculating the interaction migration probability between the functional modules according to the interaction operation data, and weighting and combining the interaction migration probability with the initial association strength to generate a dynamic association matrix includes:
[0052] Projecting the operation time series in the interactive operation data onto a two-dimensional plane to generate an operation trajectory point set; calculating the velocity vector and acceleration vector between adjacent trajectory points in the operation trajectory point set; and counting the residence time of the operation trajectory point set in the neighborhood of the functional module and the number of switching times between the functional modules;
[0053] Determining the type of operation intention according to the velocity vector and the acceleration vector;
[0054] Establishing an operation intention counter for each of the functional modules, and counting the counting results of each type of operation intention; and calculating the probability of interaction migration between the functional modules based on the counting results;
[0055] Setting a sliding time window, performing exponential smoothing on the interaction migration probability sequence within the sliding time window, and determining a combined weight of the interaction migration probability and the initial association strength, wherein the combined weight is positively correlated with the counting result within the sliding time window;
[0056] When new interactive operation data enters the sliding time window, the interactive migration probability and the combination weight are recalculated, and the weighted combination is re-performed to update the dynamic association matrix.
[0057] According to a second aspect of the embodiments of the present invention,
[0058] Provided is a vehicle-mounted display terminal human-machine interface interactive experience optimization system, including:
[0059] The first unit is used to obtain the interactive behavior data of the driver, wherein the interactive behavior data includes eye gaze point position data, hand operation trajectory data and head rotation angle data;
[0060] The second unit is used to fuse the eye gaze point position data with the head rotation angle data, and adjust the weights in combination with the hand operation trajectory data to generate a driver's attention distribution heat map; based on the heat values of each area in the attention distribution heat map, calculate the importance score of each functional module of the vehicle display terminal;
[0061] The third unit is used to collect vehicle operation status information and environmental perception information, build a dynamic scene recognition model, identify the current driving scene type in real time through a deep learning network, and output the usage probability prediction value of the functional module; the importance score and the usage probability prediction value are weighted and integrated to obtain the functional module interaction experience optimization index;
[0062] The fourth unit is used to adjust the corresponding functional modules to the hot spot area of the attention distribution heat map according to the preset index threshold, and group and cluster the functional modules based on the correlation degree, and dynamically adjust the touch response area and display level of the functional modules; when the change in the interactive behavior data exceeds the preset change threshold, it triggers the re-execution of the optimization step.
[0063] According to a third aspect of the embodiments of the present invention,
[0064] An electronic device is provided, comprising:
[0065] processor;
[0066] a memory for storing processor-executable instructions;
[0067] The processor is configured to call the instructions stored in the memory to execute the aforementioned method.
[0068] According to a fourth aspect of the embodiments of the present invention,
[0069] A computer-readable storage medium is provided, on which computer program instructions are stored. When the computer program instructions are executed by a processor, the aforementioned method is implemented.
[0070] In an embodiment of the present invention, by fusing eye gaze point position data, head rotation angle data and hand operation trajectory data, a driver's attention distribution heat map is generated, thereby achieving accurate capture of the driver's attention, effectively improving the accuracy and response speed of human-computer interaction, and reducing the driver's operating burden; a dynamic scene recognition model is constructed in combination with vehicle operation status information and environmental perception information, and the current driving scene type is identified in real time through a deep learning network, and the display position and touch response area of the functional module are dynamically adjusted accordingly, so that the interface layout is more in line with the driver's usage habits in different scenarios, and the intelligence level and adaptability of the interactive experience are improved; the functional module interactive experience optimization index is used as an evaluation standard, important functional modules are adjusted to the attention hotspot area, and grouping and clustering are performed based on functional relevance, thereby achieving dynamic optimization of the interface layout, reducing the driver's visual search time and cognitive load, improving driving safety, and ensuring the timeliness and stability of system response. BRIEF DESCRIPTION OF THE DRAWINGS
[0071] Figure 1 A schematic diagram of a process for optimizing a human-machine interface interactive experience of an in-vehicle display terminal according to an embodiment of the present invention;
[0072] Figure 2 This is a line chart of driving scene recognition accuracy;
[0073] Figure 3 Visualization diagram of the dimensionality reduction of scene prototype feature space;
[0074] Figure 4 Scatter heat map of action intent classification performance. DETAILED DESCRIPTION
[0075] In order to make the purpose, technical solution and advantages of the embodiments of the present invention clearer, the technical solution in the embodiments of the present invention will be clearly and completely described below in conjunction with the drawings in the embodiments of the present invention. Obviously, the described embodiments are only part of the embodiments of the present invention, not all of the embodiments. Based on the embodiments of the present invention, all other embodiments obtained by ordinary technicians in this field without creative work are within the scope of protection of the present invention.
[0076] The technical solution of the present invention is described in detail with specific embodiments below. The following specific embodiments can be combined with each other, and the same or similar concepts or processes may not be described in detail in some embodiments.
[0077] Figure 1 FIG. 1 is a flow chart of a method for optimizing the human-machine interface interactive experience of a vehicle-mounted display terminal according to an embodiment of the present invention. Figure 1 As shown, the method includes:
[0078] Acquiring interactive behavior data of the driver, wherein the interactive behavior data includes eye gaze point position data, hand operation trajectory data, and head rotation angle data;
[0079] The eye gaze position data and the head rotation angle data are fused and calculated, and the weights are adjusted in combination with the hand operation trajectory data to generate a heat map of the driver's attention distribution;
[0080] Based on the heat value of each area in the attention distribution heat map, calculate the importance score of each functional module of the vehicle display terminal;
[0081] Collect vehicle operation status information and environmental perception information, build a dynamic scene recognition model, use the deep learning network to identify the current driving scene type in real time, and output the usage probability prediction value of the functional module;
[0082] The importance score and the usage probability prediction value are weighted and integrated to obtain the functional module interaction experience optimization index;
[0083] According to the preset index threshold, the corresponding functional modules are adjusted to the hot spots of the attention distribution heat map, and grouped and clustered based on the correlation of the functional modules, and the touch response area and display level of the functional modules are dynamically adjusted;
[0084] When the change in the interactive behavior data exceeds a preset change threshold, the optimization step is triggered to be re-executed.
[0085] In a specific implementation, an infrared camera is used to track and capture the eye movement trajectory, and the two-dimensional coordinate position and dwell time of the gaze point on the display screen are recorded; the position coordinate sequence and pressure value of finger sliding and clicking are collected through a capacitive touch panel; and the rotation angle data of the driver's head in three degrees of freedom are obtained through a posture sensor.
[0086] The eye gaze point position and the head rotation angle are spatially mapped and corrected to eliminate the influence of head movement on the gaze point positioning. Then, based on the spatiotemporal characteristics of the hand operation, the degree of certainty of the operation intention is judged and used as a weight factor to adjust the credibility of the gaze behavior. Through the kernel density estimation method, the corrected gaze point position data is converted into a continuous attention distribution heat map, and the heat value indicates the degree to which the area obtains attention resources.
[0087] The display terminal interface is divided into grid units, and the overlap between the grid units occupied by each functional module and the attention distribution heat map is counted to calculate the attention coverage rate. At the same time, the frequency of use and operation duration of the functional module are considered, and these features are weighted and summed with the attention coverage rate to obtain the importance score of each functional module.
[0088] The vehicle sensor network collects vehicle status parameters such as speed, acceleration, steering wheel angle, and environmental information such as lane lines, traffic signs, and surrounding obstacles. These multimodal data are input into a pre-trained deep neural network, which contains a convolutional layer for extracting spatial features, a recurrent layer for capturing temporal dependencies, and an attention mechanism for highlighting key information. The network outputs the category discrimination results of the current scene in real time, and predicts the usage probability of each functional module in this type of scene based on historical data.
[0089] The importance score and the predicted usage probability value are weighted averaged by adaptive weights to obtain the optimization index. Functional modules whose optimization index exceeds the preset threshold are rearranged and moved to the hot spot area of the attention distribution heat map. At the same time, an affinity matrix is constructed based on the call association relationship between functional modules, and the spectral clustering algorithm is used to divide the highly associated functional modules into groups, and the spatial proximity of the modules in the group is maintained during the interface layout. The system also dynamically adjusts the touch response area size and display hierarchy priority of each functional module according to the attention distribution.
[0090] Continuously monitor the changes in interactive behavior data. When the eye gaze pattern, gesture operation characteristics or head posture change significantly, trigger a new round of optimization process to achieve real-time adaptive adjustment of the interface layout.
[0091] In an optional implementation, the eye gaze point position data and the head rotation angle data are fused and calculated, and the weights are adjusted in combination with the hand operation trajectory data to generate a driver's attention distribution heat map, including:
[0092] The pupil center coordinates are obtained based on the driver's eye gaze point position data, the high-frequency jitter data in the pupil center coordinates are eliminated by Kalman filtering, and the pupil center coordinates are smoothed by using a RANSAC algorithm to obtain continuous gaze point trajectory data;
[0093] According to the yaw angle and translation vector in the head rotation angle data, the continuous gaze point trajectory data is compensated for spatial coordinates to obtain gaze point coordinate compensation data;
[0094] Through the hand positioning points in the hand operation trajectory data, the three-dimensional space coordinate sequence is extracted, the hand movement speed, movement acceleration, movement direction angle, displacement distance and gesture type are calculated, and the hand trajectory feature vector is constructed;
[0095] Using a dynamic time warping algorithm, the gaze point coordinate compensation data, the head rotation angle data, and the hand trajectory feature vector are time-series aligned to construct a multimodal feature tensor;
[0096] Calculating the weight coefficient of each feature in the multimodal feature tensor, and combining the weight coefficient with a Gaussian kernel function to generate attention density distribution data;
[0097] Within a preset time window, the density feature vector of the attention density distribution data is extracted and time series fusion is performed through exponential decay weights. A dual-domain regularization function is constructed to eliminate abnormal points. The initial hot zone boundary point set is determined based on an adaptive segmentation threshold, and the boundary is guided to dynamically expand to generate a driver's attention distribution heat map.
[0098] In a specific embodiment, the eye tracking device is used to collect the eye gaze point position data of the driver to obtain the pupil center coordinates. The collected original pupil center coordinates usually contain high-frequency jitter. For example, for an eye tracking device with a collection frequency of 60Hz, continuous but jittery coordinate points such as (320.5, 240.3), (318.7, 241.2), (321.6, 239.8) may be recorded. These coordinates are processed using a Kalman filter, and the process noise covariance is set to 0.01, the measurement noise covariance is set to 0.1, and high-frequency jitter data is eliminated. Subsequently, the filtered coordinates are smoothed using the RANSAC algorithm, the inner point threshold is set to 2 pixels, the number of iterations is 100, and continuous gaze point trajectory data is fitted, such as the smoothed coordinate sequence (320.1, 240.5), (320.4, 240.7), (320.8, 240.9), etc.
[0099] According to the yaw angle and translation vector in the head rotation angle data, the continuous gaze point trajectory data is compensated for spatial coordinates. For example, when the head yaw angle is detected to be 15 degrees and the translation vector is (5cm, 2cm, 0cm), the gaze point coordinates are compensated through the spatial coordinate transformation matrix. In specific implementation, a conversion relationship between the head coordinate system and the fixed coordinate system in the car is established. When the head rotates, the eye gaze point is mapped to the fixed coordinate system in the car through the compensation matrix. The compensated gaze point coordinates can reflect the position that the driver actually pays attention to. For example, the original coordinates (320.8, 240.9) may be compensated to (350.2, 245.3).
[0100] Through the hand positioning points in the hand operation trajectory data, the three-dimensional space coordinate sequence is extracted, and the coordinate sequence of the driver's right hand when operating the steering wheel is recorded as {(45.2, 30.5, 10.3), (46.8, 32.1, 10.5), (48.3, 33.7, 10.8)}. Based on these coordinates, the hand motion features are calculated: the motion speed is the displacement between adjacent sampling points divided by the time interval, such as 12.5cm / s; the motion acceleration is the velocity change rate, such as 2.3cm / s²; the motion direction angle is calculated by the displacement vector of adjacent points, such as 35 degrees; the displacement distance is the cumulative path length, such as 8.7cm; the gesture type is determined by the finger joint angle, such as the "holding" state. These features are combined to form the hand trajectory feature vector [12.5, 2.3, 35, 8.7, 1], where the last digit "1" indicates the holding gesture type.
[0101] The dynamic time warping algorithm is used to align the gaze point coordinate compensation data, head rotation angle data, and hand trajectory feature vectors in time series. The time window is set to 2 seconds, the window step is 0.5 seconds, and the time correspondence between the data streams is calculated. For example, when the sampling rate of the gaze point data is 60Hz, the sampling rate of the head data is 30Hz, and the sampling rate of the hand data is 20Hz, the dynamic programming method is used to find the best alignment path to align the three types of data in the time dimension. After alignment, a multimodal feature tensor is constructed, which includes the time dimension, space dimension, and feature dimension to form a unified data structure.
[0102] Based on the characteristics of the driving scene, the initial weights are set as follows: the weight of the gaze point data is 0.6, the weight of the head rotation data is 0.3, and the weight of the hand operation data is 0.1. The weights are then adjusted dynamically according to the complexity of the driving task. For example, during steering operations, the hand weight may increase to 0.25, and the gaze point weight may be reduced to 0.45 accordingly. The weight coefficients are combined with the Gaussian kernel function to generate the attention density distribution data. The standard deviation of the Gaussian kernel function is set to the pixel distance corresponding to a viewing angle of 5 degrees, for example, about 8.7 cm at 1 meter in front of the driver.
[0103] In the preset 5-second time window, the density feature vector is extracted from the attention density distribution data. The maximum density value, density center position, density distribution range and other features are calculated for the density distribution map at each time point to form a density feature vector. The exponential decay weight is used for time series fusion, and the decay coefficient is set to 0.8, so that the recent attention distribution has a higher weight. A dual-domain regularization function is constructed to eliminate outliers. The spatial domain threshold is set to 20% of the density mean, and the temporal domain threshold is set to 3 standard deviations. Points below the threshold are regarded as outliers and smoothed.
[0104] The initial hot zone boundary point set is determined based on the adaptive segmentation threshold, and the threshold is set to 40% of the maximum density value. For example, when the maximum density value is 0.85, the area with a density value greater than 0.34 is considered as the initial hot zone. The boundary is then guided to expand dynamically, with an expansion step of 2 pixels and 10 iterations, allowing the hot zone to expand appropriately along the gradient direction to form a more natural boundary. Finally, a heat map of the driver's attention distribution is generated, using a red-yellow-green scale to represent the intensity of attention, with red areas representing the most concentrated areas of attention, yellow areas representing the second highest attention areas, and green areas representing low attention areas.
[0105] This embodiment can accurately capture the driver's attention distribution during driving, providing an important basis for driving safety monitoring and assisted driving systems.
[0106] In an optional implementation, within a preset time window, the density feature vector is extracted from the attention density distribution data and time series fusion is performed through an exponential decay weight, a dual-domain regularization function is constructed to eliminate abnormal points, an initial hot zone boundary point set is determined based on an adaptive segmentation threshold, and the boundary is guided to dynamically expand, and the driver's attention distribution heat map is generated, including:
[0107] The statistical features and spatial moment features of the attention density distribution data are extracted within the preset time window to construct a density feature vector. The density feature vectors of the historical window and the current window are fused through the exponential decay weight to obtain a time series fusion feature vector.
[0108] Determine the coordinates of the attention center point according to the time series fusion feature vector, construct a dual-domain regularization function containing spatial distance constraints and density value constraints, and perform regularization processing to obtain the optimized data of attention density distribution;
[0109] Calculate the adaptive segmentation threshold based on the attention density distribution optimization data and determine the initial hot zone boundary point set;
[0110] Calculate the density gradient direction and density value of the initial hot zone boundary point set, and set the main threshold and secondary threshold in the corresponding eight neighborhoods along the density gradient direction;
[0111] When the density value of the eight neighborhoods is greater than the main threshold, the corresponding points are included in the initial hot zone boundary point set;
[0112] When the density value of the eight neighborhoods is between the primary threshold and the secondary threshold, the corresponding points are marked as candidate boundary point sets; the curvature value of the initial hot zone boundary point set is calculated to construct a dynamic scaling function, and the primary threshold and the secondary threshold are adjusted; the shortest path distance from the candidate boundary point set to the initial hot zone boundary point set is calculated, and the candidate boundary point set whose shortest path distance is less than the preset connectivity threshold is merged into the initial hot zone boundary point set;
[0113] When the area growth rate of the initial hot zone boundary point set is less than the preset growth threshold, and the number of candidate boundary point sets is less than the preset total number threshold, the boundary expansion is stopped to obtain the expanded hot zone boundary point set;
[0114] The kernel density estimation and normalization processing are performed on the boundary point set of the expansion hot zone to obtain the driver's attention distribution heat map.
[0115] In a specific implementation, for the attention density distribution data in each time window, its statistical features and spatial moment features are calculated. Statistical features include maximum density value, average density value, density standard deviation, etc.; spatial moment features include density centroid coordinates, main axis direction of density distribution, degree of discreteness of density distribution, etc. For example, for a 10-second time window, the collected attention density distribution data is a two-dimensional matrix D(x, y), where x and y represent the pixel coordinates in the horizontal and vertical directions, respectively, and D(x, y) represents the attention density value at the coordinate (x, y). The calculated maximum density value is 0.85, the average density value is 0.32, the density standard deviation is 0.18, the density centroid coordinates are (320, 240), the main axis direction is 30 degrees, and the degree of discreteness is 0.45. These features are combined to form a density feature vector F current .
[0116] Assume that the historical fusion feature vector is F history , the current eigenvector is F current , then the time series fusion feature vector F fusion Calculated as: F fusion = α × F history + (1-α) × F current , where α is the attenuation factor, and its value range is [0, 1]. In practical applications, α can be set to 0.7, which means that historical data accounts for 70% of the weight and current data accounts for 30% of the weight. For example, if F history is [0.80, 0.30, 0.20, (310, 230), 25, 0.40], F current is [0.85, 0.32, 0.18, (320, 240), 30, 0.45], then F fusion is [0.815, 0.306, 0.194, (313, 233), 26.5, 0.415].
[0117] The center of attention usually corresponds to the density centroid coordinates, which is (313, 233) in the above example. Based on the center of attention, a dual-domain regularization function is constructed for data optimization. The dual-domain regularization function contains two parts: spatial distance constraint and density value constraint. The spatial distance constraint makes the density value of the point farther from the center point decay more; the density value constraint suppresses abnormally high or abnormally low density values. In specific implementation, for each point (x, y), calculate its distance d to the center point and the original density value D(x, y), and then calculate the regularized density value D'(x, y). For example, when d=50 pixels and D(x, y)=0.6, if Gaussian space constraint and density threshold constraint are used, D'(x, y)=0.55 can be obtained.
[0118] Based on the optimized density distribution data, the adaptive segmentation threshold is calculated to determine the initial hot zone boundary point set. The adaptive segmentation threshold T init It can be determined by Otsu's method or density histogram-based method. For example, for the optimized density distribution, T init =0.4. Set the density value greater than T init The point set is defined as the initial hot zone boundary point set B init .
[0119] For B init For each point (x, y) in , calculate its density gradient in the x and y directions to obtain the gradient direction θ and gradient magnitude g. For example, for the boundary point (350, 260), θ = 45 degrees and g = 0.15 are calculated.
[0120] Then, the main threshold and secondary threshold are set in the corresponding eight neighborhoods along the density gradient direction. main Set to T init 80% of T main =0.32; subthreshold T secondary Set to T init 50% of T secondary =0.2. For each boundary point, check the density value of its eight neighboring points: if the density value is greater than T main , then the point is included in the initial hot zone boundary point set; if the density value is between T main and T secondary If the point is between , then the point is marked as the candidate boundary point set C.
[0121] Calculate the curvature value of the initial hot zone boundary point set and construct a dynamic scaling function to adjust the threshold. The curvature value k can be calculated from the local shape of the boundary point. For example, for the boundary point (350, 260), k=0.08 is calculated. Based on the curvature value, adjust the primary and secondary thresholds: T main ' = T main × (1-β×k), T secondary ' = T secondary × (1-β×k), where β is the adjustment coefficient, which can be set to 0.5. For k=0.08, the adjusted main threshold T main '=0.3072, sub-threshold T secondary '=0.192.
[0122] Calculate the shortest path distance from the candidate boundary point set to the initial hot zone boundary point set. For each point in the candidate boundary point set C, calculate its distance to B init The distance d between the nearest points min If d min Less than the preset connectivity threshold T connect(such as 10 pixels), then the candidate point is incorporated into the initial hot zone boundary point set. For example, the candidate point (360, 270) to B init The shortest distance is 8 pixels, which is less than T connect , so it is merged into B init , and get the expanded boundary point set B extended .
[0123] Repeat the boundary expansion process until the stopping condition is met. The stopping condition includes: the area growth rate of the boundary point set is less than the preset growth threshold T growth (e.g. 5%), and the number of candidate boundary point sets is less than the preset total threshold T count (For example, 10% of the total number of boundary points). For example, when the boundary is expanded to the fifth round, the area growth rate is 3%, and the number of candidate boundary points is 8% of the total number of boundary points, which meets the stopping condition and obtains the final expansion hot zone boundary point set B. final .
[0124] Most right B final For each point in , a kernel function (such as a Gaussian kernel) is applied for density estimation to obtain a smooth density distribution H. H is normalized so that its value range is mapped to the interval [0, 1] to obtain the final attention distribution heat map. For example, for an image with a resolution of 1920×1080, in the generated heat map, the value of the area with the most concentrated attention (such as the front of the road) is close to 1, represented by red; the value of the area with less attention (such as the sides of the vehicle) is close to 0, represented by blue; the intermediate transition area is represented by yellow-green.
[0125] like Figure 2 As shown, the performance of various technical solutions in driving scene recognition accuracy under different attenuation factor α values is demonstrated. This technical solution uses exponential decay weights to fuse the density feature vectors of the historical window and the current window. Its superiority can be clearly seen through experimental data. When the attenuation factor α varies between 0.1 and 0.8, the scene recognition accuracy of this technical solution shows a trend of first rising and then falling, and the best performance is achieved at α=0.6, with an accuracy of 89.2%. In contrast, the highest accuracy of the existing technology A (no time series fusion) is only 74.0%, the highest accuracy of the existing technology B (simple average fusion) is 77.3%, and the highest accuracy of the existing technology C (short-term memory fusion) is 82.6%. Under the optimal attenuation factor setting, this technical solution is 6.6 percentage points higher than the existing optimal technology. Experimental data also show that when the α value is too large (such as 0.8), over-reliance on historical data will cause the system to be insensitive to new situations, and the accuracy rate will drop to 86.7%; when the α value is too small (such as 0.1), the system is too sensitive to noise and interference, and the accuracy rate is only 70.5%. These data fully verify the effectiveness of the time series fusion strategy in this technical solution, and α = 0.6 ~ 0.7 is the optimal parameter range.
[0126] The existing driver attention distribution heat map generation technology mainly uses a simple accumulation method to directly accumulate and smooth the eye movement data points to generate a heat map without considering the time series characteristics. Or use a Gaussian mixture model to fit the attention distribution through multiple Gaussian kernel functions, but the parameter settings are fixed and lack adaptability. There is also a time window sliding average method that averages the attention data within a fixed time window without considering the attenuation effect of historical data. In addition, the threshold segmentation method uses a global fixed threshold to segment the attention density distribution, and the boundary extraction is not accurate. Convolution-based hot zone extraction uses a convolution kernel for attention density filtering, which is sensitive to density gradient changes, but has high computational complexity.
[0127] During driving, the driver's attention distribution has the dual characteristics of temporal continuity and mutation. Existing technologies either ignore historical information or use simple mean processing, which cannot accurately reflect the true distribution of attention. This technology introduces an exponential decay weight mechanism and integrates the density feature vectors of the historical window and the current window, making the representation of attention distribution more temporally coherent while retaining sensitivity to sudden changes.
[0128] Existing technologies are not effective in dealing with attention density noise and outliers, and cannot effectively balance the accuracy of the attention center and the integrity of the distribution boundary. This technology constructs a dual-domain regularization function that includes spatial distance constraints and density value constraints, comprehensively considering the information of the two dimensions of spatial position and density value, and realizes the optimization of attention density distribution.
[0129] The traditional fixed threshold segmentation method cannot adapt to the differences in attention distribution in different driving scenarios, and the boundary extraction is inaccurate. This technology is based on the density gradient direction and curvature characteristics, designs a dynamic adjustment mechanism for primary and secondary thresholds and a candidate boundary point connectivity analysis method to achieve adaptive and accurate expansion of the boundary.
[0130] Existing technologies are insufficient in processing the connectivity and boundary integrity of low-density areas, which easily leads to fragmented hotspots. This technology constructs a dynamic scaling function by calculating the boundary curvature, and combines the connectivity analysis of the shortest path distance to achieve effective connection and boundary smoothing of low-density areas.
[0131] Compared with the existing technology, the accuracy of attention hotspot recognition in different driving scenarios of this method is improved by an average of 18.6%, especially in complex road environments. By fusion of time-series feature vectors, the temporal continuity of attention distribution is guaranteed, reducing 75% of timing jump anomalies. The adaptive density gradient boundary expansion algorithm improves the accuracy of hotspot boundary extraction by 32.4% and the boundary smoothness by 41.7%. Compared with the convolution-based hotspot extraction method, this technology reduces the computational complexity by 43.2% and achieves real-time processing capabilities. In addition, it can automatically adapt to the attention distribution characteristics of different driving scenarios, different individual characteristics of drivers, and different driving tasks, and its versatility and robustness are significantly improved.
[0132] In summary, this technology effectively solves the problems of timing incoherence, imprecise boundaries, and insufficient adaptability in the existing driver attention distribution heat map generation technology through innovative methods such as timing adaptive fusion, dual-domain regularization optimization, adaptive density gradient boundary expansion, and multi-threshold dynamic scaling. It provides a more accurate and reliable attention distribution representation tool for driving behavior analysis, driving safety assessment, and intelligent driving assistance systems.
[0133] In this embodiment, exponential decay weights are used to fuse historical and current data to effectively suppress noise and outliers, thereby ensuring the stability and reliability of the attention density distribution data. Statistical features, spatial moment features, and density gradient analysis are used to accurately determine the attention center and the initial hot zone boundary, thereby achieving an accurate characterization of the driver's attention hotspots. The adaptive segmentation threshold and dynamic expansion strategy enable the boundary to be flexibly adjusted during the change process, thereby improving the real-time and adaptability of the detection. After kernel density estimation and normalization processing, the final generated heat map can clearly reflect the driver's attention distribution, providing reliable data support for subsequent monitoring and analysis.
[0134] In an optional implementation, vehicle operation status information and environmental perception information are collected, a dynamic scene recognition model is constructed, and the current driving scene type is identified in real time through a deep learning network. The usage probability prediction value of the output function module includes:
[0135] Acquire vehicle operation status information and environmental perception information, calculate the volatility of the vehicle operation status information and environmental perception information based on a sliding time window, and determine the volatility by the ratio of the change amplitude of the characteristic values at adjacent time points to the average change amplitude in the time window; divide the scene sequence into a steady-state interval and a mutation interval according to a preset fluctuation threshold; perform feature sampling of different frequencies on the steady-state interval and the mutation interval, respectively, to obtain an adaptive sampling feature sequence;
[0136] Matching the adaptive sampling feature sequence with historical scene samples, constructing and solving the scene evolution equation based on the element coupling relationship in the historical scene samples and the stability of the scene prototype, obtaining the state association matrix, and dynamically updating the dynamic coefficient using the prediction deviation;
[0137] In the adaptive sampling feature sequence, a sampling moment whose prediction deviation is greater than a preset deviation threshold is selected as a key scene frame; the features of the key scene frame are input into a pre-trained scene classifier to obtain a scene type recognition result; and based on the scene type recognition result and the prediction deviation, a usage probability prediction value of the corresponding functional module is calculated.
[0138] In a specific implementation, the vehicle operation status information includes vehicle speed, acceleration, steering wheel angle, brake pedal position, accelerator pedal position, etc.; the environmental perception information includes the distance to the vehicle ahead, relative speed, lane line type, road curvature, traffic sign recognition results, etc. This information is collected through the vehicle CAN bus and sensor network, with a sampling frequency of 10Hz.
[0139] Set the sliding time window size to 3 seconds, and the window contains 30 data points. For each feature, calculate the change amplitude of adjacent time points. For example, if the speed at the current time t is 60km / h and the speed at t+0.1s is 61km / h, the change amplitude is 1km / h. Calculate the average of all change amplitudes in the time window, such as the average change amplitude of the speed in the window is 0.5km / h. The volatility is equal to the ratio of the current change amplitude to the average change amplitude in the window, such as the current volatility is 1 / 0.5=2.
[0140] The preset fluctuation threshold is 3. When the fluctuation rate is greater than 3, it is determined to be a sudden change interval, otherwise it is a steady state interval. For example, during normal straight-line driving, the vehicle speed fluctuation rate remains below 1.5, which belongs to the steady state interval; while in emergency avoidance of pedestrians, the steering wheel angle fluctuation rate can reach 4.5, and the brake pedal position fluctuation rate can reach 5.2, which is determined to be a sudden change interval.
[0141] In the steady-state interval, the sampling frequency is reduced to 1Hz; in the mutation interval, the original sampling frequency of 10Hz is maintained. For example, in a 10-second driving process, the first 7 seconds are the steady-state interval and the last 3 seconds are the mutation interval. Then, 7 sample points are collected in the steady-state interval and 30 sample points are collected in the mutation interval, and a total of 37 sample points constitute the adaptive sampling feature sequence.
[0142] The historical scene sample library contains 100,000 annotated driving scene data, each of which contains a feature sequence and a corresponding scene type label. The KNN algorithm is used to calculate the similarity between the current feature sequence and the historical samples, and the 100 historical samples with the highest similarity are selected.
[0143] The scenario evolution equation is constructed based on the coupling relationship of elements in historical scenario samples and the stability of scenario prototypes. The coupling relationship of elements refers to the mutual influence between different features, such as the relationship between vehicle speed and the distance to the vehicle in front, the relationship between the steering wheel angle and the curvature of the road, etc. The stability of the scenario prototype refers to the regularity of feature changes in a specific scenario. By analyzing the temporal change rules of features in historical samples, the state transfer matrix is constructed to predict the feature values at the next moment.
[0144] For example, in the "highway following" scenario, the reduction in the distance to the vehicle ahead usually causes the vehicle to slow down, and the influence coefficient of the distance to the vehicle ahead on the speed in the state association matrix is -0.8; while in the "city congestion" scenario, the influence coefficient is -0.3. By comparing the predicted value with the actual value, the prediction deviation is calculated. When the prediction deviation exceeds 15%, the coefficient of the state association matrix is dynamically updated, such as adjusting -0.8 to -0.75.
[0145] In the adaptive sampling feature sequence, the sampling moment with a prediction deviation greater than the preset deviation threshold is selected as the key scene frame. The preset deviation threshold is 20%. When the prediction deviation at a certain moment exceeds 20%, the moment is marked as a key scene frame. For example, during a lane change, the predicted value of the steering wheel angle is 15 degrees, and the actual value is 22 degrees. The prediction deviation is (22-15) / 15=46.7%, which exceeds the threshold, and the moment is marked as a key scene frame.
[0146] The features of the key scene frames are input into the pre-trained scene classifier to obtain the scene type recognition results. The scene classifier adopts a deep neural network structure, which includes 4 convolutional layers and 2 fully connected layers. The input is the feature vector of the key scene frame, and the output is the probability distribution of 12 predefined scene types. The scene types include: high-speed cruising, high-speed following, city straight driving, city turning, city lane changing, congested driving, waiting for traffic lights, parking lot driving, reversing into the warehouse, emergency avoidance, driving in bad weather, mountain road driving, etc.
[0147] According to the scene type recognition results and prediction deviation, the usage probability prediction value of the corresponding functional module is calculated. Functional modules include adaptive cruise control, lane keeping assist, automatic emergency braking, automatic parking, etc. The usage probability prediction value is calculated by the weighted sum of the scene type probability and the historical usage frequency of the functional module in the scene.
[0148] For example, when the current scene is identified as "high-speed following" (probability 0.85) and "high-speed lane change" (probability 0.15), the historical usage frequencies of adaptive cruise control in these two scenes are 0.95 and 0.3 respectively, and the usage probability prediction value of adaptive cruise control is 0.85×0.95+0.15×0.3=0.8525. When the usage probability prediction value exceeds 0.7, the system will actively recommend to the driver to turn on the corresponding function module.
[0149] In an optional implementation, the adaptive sampling feature sequence is matched with the historical scene sample, the scene evolution equation is constructed and solved based on the element coupling relationship in the historical scene sample and the stability of the scene prototype, the state association matrix is obtained, and the dynamic coefficient is dynamically updated using the prediction deviation, including:
[0150] Extracting scene elements from the historical scene samples, wherein the scene elements include vehicles, pedestrians, and road facilities; constructing an element coupling matrix, wherein each element of the element coupling matrix represents an interaction influence intensity value between adjacent scene elements;
[0151] Clustering is performed based on the historical scene samples to obtain a plurality of scene prototypes; obtaining the eigenvalue of each scene prototype in the element coupling matrix, and calculating the stability index of each scene prototype in combination with the spatial distribution density of the scene elements;
[0152] Constructing a scenario evolution equation including a steady-state term and a disturbance term, wherein the coefficient of the steady-state term comes from the stability index, and the coefficient of the disturbance term comes from the element coupling matrix; solving the scenario evolution equation to obtain the state transition probability between scenario prototypes, and constructing a state association matrix;
[0153] The adaptive sampling feature sequence is input into the state association matrix for feature prediction to obtain predicted features; the predicted features are compared with the actual features to calculate the prediction deviation, and based on the prediction deviation, the interaction influence intensity value in the element coupling matrix and the disturbance term coefficient in the scenario evolution equation are updated respectively.
[0154] In a specific embodiment, key elements, including vehicles, pedestrians, and road facilities, are extracted from historical traffic scene samples. This process uses computer vision technology combined with a target detection algorithm to identify each element. A pre-trained deep learning target detection network is used to process historical scene images or videos to identify all vehicles, pedestrians, and road facilities. Each identified element is marked, and its feature information such as location coordinates, type, size, direction of movement, and speed is recorded. The interaction intensity value between adjacent elements is calculated, and an element coupling matrix is constructed. The interaction intensity value is calculated based on the spatial distance between elements, relative movement speed and direction, historical collision or conflict frequency, and traffic rule constraint relationship. The calculated interaction intensity value is filled into the coupling matrix, and each element in the matrix represents the interaction intensity between a specific element pair and another element pair.
[0155] The historical scenes are represented as multidimensional feature vectors, and the features include the number of elements, distribution density, type ratio, interaction mode, etc. The scene feature vectors are clustered by clustering algorithms to obtain multiple scene prototypes. The number of clusters can be determined by methods such as silhouette coefficient and elbow rule. For each scene prototype, its eigenvalue in the element coupling matrix is extracted. The specific method is to regard the coupling matrix of the prototype as a network structure and calculate the eigenvalue spectrum of the matrix. Combined with the spatial distribution density of scene elements, the stability index of each prototype is calculated. First, the spatial distribution of elements in the prototype is analyzed, the kernel density estimate is calculated, and the distribution of the eigenvalue of the element coupling matrix is analyzed. The dominant eigenvalue of the eigenvalue spectrum is combined with the spatial density distribution to obtain the stability index. The higher the stability index, the stronger the ability of the scene prototype to remain in its original state under external disturbances; the lower the index, the more likely the scene is to change or transform into other states.
[0156] Construct a scenario evolution equation containing steady-state terms and disturbance terms. The steady-state term describes the tendency of the system to maintain the current state, and the disturbance term describes the possibility of system change. The coefficient of the steady-state term directly adopts the stability index calculated previously, which indicates the tendency of the scenario to maintain the current state. The coefficient of the disturbance term comes from the interaction intensity value in the element coupling matrix, which indicates the state change that may be caused by the interaction between the elements in the scenario. Solve the scenario evolution equation and calculate the state transition probability between different scenario prototypes. The solution process adopts numerical methods such as finite difference method or Monte Carlo simulation. Organize the calculated transition probability into a state association matrix, in which the matrix elements represent the probability of transitioning from one scenario prototype to another.
[0157] Adaptively sample new traffic scene data to obtain feature sequences. Adaptive sampling dynamically adjusts the sampling frequency according to the rate of scene change. Input the feature sequence obtained by adaptive sampling into the state association matrix, perform forward prediction, and obtain prediction features. Compare the predicted features with the actual observed features to calculate the prediction deviation. Based on the prediction deviation, use methods such as gradient descent or Bayesian optimization to update the interaction intensity value in the factor coupling matrix. At the same time, adjust the coefficient of the disturbance term in the scene evolution equation so that the model can more accurately reflect the actual scene change law. Repeat the above process periodically to continuously optimize the model performance and improve the prediction accuracy.
[0158] For example, in a traffic surveillance video dataset of a city intersection, we first use the YOLOv5 target detection algorithm to extract scene elements. The system successfully identifies the following elements: vehicles include 30 private cars, 2 buses, and 5 trucks; pedestrians include 25 adult pedestrians and 3 child pedestrians; road facilities include 4 sets of traffic lights, 8 traffic signs, 2 crosswalks, and 4 lane dividers. For each element, the location information is recorded, such as "Car 1 is located at coordinates (120,350), driving east, and the speed is 30km / h". Then the interaction intensity between adjacent elements is calculated, for example, the interaction intensity between two vehicles traveling in the same direction is 0.3, while the interaction intensity between vehicles and crosswalks is 0.8. The size of the final constructed coupling matrix is 65×65, where the element value range is between 0-1, indicating the interaction intensity.
[0159] 100 historical intersection scene samples are represented as feature vectors, including dimensions such as vehicle density, pedestrian density, and signal light status. Five typical scene prototypes are obtained using the K-means clustering algorithm: dense traffic during the green light straight-through phase, stationary vehicles and pedestrians passing during the red light waiting phase, some vehicles turning during the left turn phase, slow vehicles moving in congestion, and low traffic at night. The coupling matrix eigenvalues are extracted for each prototype. For example, the main eigenvalue of the green light straight-through is 3.7, and the main eigenvalue of the red light waiting is 1.2. Combined with the spatial distribution density of the elements, the stability index is calculated: 0.65 for the green light straight-through phase and 0.82 for the red light waiting phase. This shows that the red light waiting state is more stable than the green light straight-through state and is less susceptible to external disturbances.
[0160] The evolution equation is constructed, and the coefficient of the steady-state term uses the stability index. The coefficient of the disturbance term uses the significant interaction intensity value in the coupling matrix, such as the interaction intensity of 0.85 between the vehicle and the traffic light as one of the disturbance coefficients. The state transition probability is calculated by the numerical solution method, and the state association matrix is obtained: the probability of transition from green light to red light is 0.73, the probability of transition from red light to left turn to green light is 0.68, the probability of completing the turn and returning to the straight state is 0.81, and the probability of turning to congestion increases to 0.35 during peak hours.
[0161] During the morning rush hour, an intersection was monitored and adaptive sampling was used to obtain feature sequences. The initial sampling interval was 10 seconds, which was reduced to 3 seconds when the traffic state changed rapidly. The sampled feature sequence was input into the state association matrix to predict the scene changes in the next 5 minutes. The prediction results showed that the intersection would change from going straight at a green light to waiting at a red light, followed by a left turn. After comparing with the actual observations, it was found that the prediction error was 12 seconds. Based on this deviation, the system updated the interaction intensity value between the vehicle and the traffic light in the coupling matrix from 0.85 to 0.82, and at the same time adjusted the corresponding disturbance term coefficient in the evolution equation from 0.75 to 0.71. After this adaptive update for a week, the average prediction error was reduced from the original 15 seconds to 7 seconds, significantly improving the prediction accuracy.
[0162] like Figure 3 As shown in the figure, five typical scene prototypes are formed after clustering 100 historical intersection scene samples. The figure uses t-SNE dimensionality reduction technology to map the high-dimensional feature space to a two-dimensional plane, and intuitively shows the distribution and boundaries of different scene prototypes in the feature space. The green light straight scene (square mark) is mainly concentrated in the upper right area of the figure, with the feature center coordinates of (0.78, 0.65). The cluster contains 28 samples, the feature cohesion is 0.83, and the stability index is 0.65. The red light waiting scene (triangle mark) is distributed in the upper left area, with the feature center point (-0.65, 0.72), containing 25 samples, the cohesion is as high as 0.91, and the stability index is 0.82, indicating that the scene structure is the most stable. The left turn stage scene (circular mark) is located in the right area, with the feature center point (0.82, 0.12), containing 17 samples, the cohesion is 0.85, and the stability index is 0.71. The congestion state scene (diamond mark) is distributed more dispersedly, located in the lower area of the figure, with the characteristic center point (0.05, -0.75), including 18 samples, the cohesion is only 0.67, and the stability index is 0.52, reflecting the diversity and instability of the congestion state. The night low-flow scene (star mark) is concentrated in the lower left area, with the characteristic center point (-0.72, -0.48), including 12 samples, the cohesion is 0.89, and the stability index is as high as 0.89. The figure clearly shows the boundaries between the prototypes of each scene, and the width of the boundary reflects the ambiguity of the scene transition. In particular, the boundary between the congestion state and other scenes is relatively fuzzy, with a boundary width index of 0.28, while the boundary between the red light waiting and the left turn stage is clear, with a boundary width index of only 0.09, indicating that the state transition is more certain. The element coupling matrix and steady-state-disturbance evolution equation constructed by this technical solution successfully capture these scene characteristics and their conversion characteristics.
[0163] The technical theoretical basis of the method in this embodiment mainly comes from complex system theory, coupling matrix analysis in graph theory, dynamic system stability theory, and clustering and state prediction methods in machine learning. Specifically, the construction of the factor coupling matrix draws on the node relationship strength quantification method in social network analysis; the stability analysis of the scene prototype is derived from the stability index theory in the dynamic system; and the construction of the scene evolution equation integrates the ideas of stochastic differential equations and Markov state transition models.
[0164] Existing traffic scene analysis technologies mainly rely on object detection algorithms in computer vision (such as FasterR-CNN, YOLO, etc.) to identify vehicles, pedestrians and road facilities in traffic scenes, but lack a systematic description of the relationship between elements. Traditional scene analysis usually uses feature-based clustering algorithms (such as K-means, DBSCAN, etc.) to classify scenes, but the clustering process mainly considers static features and ignores the dynamic interaction between scene elements. In addition, traditional prediction models mostly use time series analysis or simple state transition probability models, such as hidden Markov models (HMM) or conditional random fields (CRF), but these models are difficult to accurately capture the complex interactions and evolution laws between multiple elements in the scene. Most existing models use fixed-period offline updates or simple online learning methods. The feedback adjustment mechanism for prediction deviations is not sophisticated enough and it is difficult to adapt to the rapid changes in traffic scenes.
[0165] In view of the deficiencies of the above-mentioned prior art, the steps of this embodiment break through the limitation of the traditional method that only focuses on a single factor. By constructing an element coupling matrix, it quantitatively describes the intensity of the interaction between adjacent elements in the scene, making the scene analysis more systematic and structured. This method innovatively combines eigenvalue analysis and spatial distribution density to calculate the stability index of the scene prototype, making the scene classification more in line with physical laws. Different from the traditional single state transition model, this method constructs a scene evolution equation containing steady-state terms and disturbance terms, which more accurately describes the internal mechanism of scene evolution. This embodiment also designs a two-layer adaptive update mechanism, which updates the parameters in the element coupling matrix and the scene evolution equation according to the prediction deviation, realizes the two-layer adaptive adjustment of the model, and greatly improves the model's adaptability to scene changes.
[0166] The main starting point of this technical improvement is that the traffic scene is essentially a complex system. The analysis of a single element is difficult to accurately grasp the overall characteristics of the scene, and it is necessary to systematically describe the interaction between the elements. Traffic scenes are highly dynamic and uncertain, requiring models to be able to capture scene changes in real time and make adaptive adjustments. The generalization ability of traditional prediction models is limited, and it is difficult to cope with complex and changeable traffic scenes, requiring a more accurate prediction mechanism. At the same time, while ensuring the accuracy of the prediction, it is also necessary to consider the computational complexity of the model to achieve efficient scene prediction.
[0167] Through experimental verification, the scene prediction accuracy of this method is 15%-25% higher than that of traditional methods, especially in complex dynamic scenes. The two-layer adaptive update mechanism enables the model to respond quickly to scene changes, and the adaptability index is more than 40% higher than that of the traditional fixed update model. Although a more complex mathematical model is introduced, the calculation time of this method is only about 10% longer than that of the traditional method through the optimization algorithm, and it is still applicable in application scenarios with high real-time requirements. The introduction of the element coupling matrix and the stability index enables the model to deeply understand the internal structure and evolution law of the scene, providing a more reliable basis for subsequent decision support. Through the evolution equation of steady-state-disturbance coupling, the model shows stronger robustness in the face of abnormal scenes, and the error prediction rate is reduced by more than 30%. These improvements make this technology have broad application prospects in the fields of intelligent transportation systems, autonomous driving scene prediction, traffic safety warning, etc.
[0168] In this embodiment, by extracting key scene elements such as vehicles, pedestrians, and road facilities and constructing a coupling matrix, the intensity of the interaction between adjacent scene elements is effectively quantified, thereby achieving a comprehensive characterization of complex scene relationships; clustering is used to obtain multiple scene prototypes, and the stability index is calculated in combination with the eigenvalue and spatial density to provide an accurate steady-state term coefficient for the scene evolution equation, thereby accurately reflecting the stable state of each scene prototype; by constructing a scene evolution equation containing steady-state terms and disturbance terms, and solving the state transition probability, a state association matrix is constructed, so that the system can effectively predict the scene evolution process; based on the deviation between the predicted features and the actual features, the coupling matrix and the disturbance term coefficients are iteratively updated, thereby improving the model's adaptive adjustment capability to environmental changes and the overall prediction accuracy.
[0169] In an optional implementation, according to a preset index threshold, the corresponding functional module is adjusted to the hot spot area of the attention distribution heat map, and grouping and clustering are performed based on the correlation of the functional modules. The touch response area and display level of the functional module are dynamically adjusted, including:
[0170] Receive the functional module interaction experience optimization index and the attention distribution heat map, and select the functional modules to be adjusted from the functional module interaction experience optimization index based on a preset index threshold; divide the attention distribution heat map into multiple regional levels, and sort the functional modules by importance according to the functional module interaction experience optimization index; extract the interactive operation data of the functional module, including the operation time series, the operation interval duration, and the operation transfer probability;
[0171] Acquire the hierarchical relationship tree of the functional modules, calculate the shortest path distance between the functional modules, and construct the initial association strength based on the shortest path distance; calculate the interaction migration probability between the functional modules according to the interaction operation data, perform weighted combination of the interaction migration probability and the initial association strength, and generate a dynamic association matrix;
[0172] Performing eigendecomposition on the dynamic association matrix to obtain a main eigenvector; adaptively clustering the functional modules based on the main eigenvector to determine a functional module group; calculating the intra-group association density and inter-group association sparsity of each functional module group;
[0173] The functional module group is projected to the regional level of the attention distribution heat map, and the projection position is jointly determined by the intra-group association density and the functional module interaction experience optimization index; the relative layout spacing of the functional module group is determined based on the inter-group association sparsity; according to the position of the functional module in the attention distribution heat map, the touch response area and display level are dynamically adjusted.
[0174] In a specific implementation, the interactive experience optimization index is generated by summarizing multi-dimensional data such as user satisfaction scores when using various functions of the vehicle system, time required to complete tasks, and operation error rate. For example, the interactive experience optimization index of the navigation function module is 85 points, the music playback function is 78 points, the air conditioning control function is 92 points, and the telephone function is 75 points. Based on the preset index threshold of 80 points, the telephone function and the music playback function are selected as the function modules to be adjusted.
[0175] The attention distribution heat map is formed by capturing the driver's gaze position through the in-car camera and accumulating data. The system divides the heat map into five area levels: S-level (core attention area), A-level (high attention area), B-level (medium attention area), C-level (low attention area) and D-level (edge attention area). The S-level area is usually located in the center of the instrument panel and the central control screen, with the highest attention; while the D-level area is located at the edge of the screen, with the lowest attention. According to the interactive experience optimization index, the functional modules are ranked, with air conditioning control ranking first (92 points), navigation function ranking second (85 points), music playback ranking third (78 points), and telephone function ranking fourth (75 points).
[0176] Taking the navigation function as an example, the operation time series records are 7:30-7:45 in the morning, 12:00-12:10 in the afternoon, and 18:30-19:00 in the evening, and the operation interval is mainly concentrated in 5-6 hours. The operation transfer probability analysis of the music playback function shows that the probability of users switching to music playback after using navigation is 65%, while the probability of switching from music playback to the phone function is 45%.
[0177] In the in-vehicle interface, the primary functional modules include navigation, entertainment, communication, and vehicle settings, while the secondary functional modules include various sub-functions. It is calculated that the shortest path distance from navigation to music playback is 2 (needs to pass through the entertainment center), and the shortest path distance from navigation to phone function is 3 (needs to pass through the communication center). The initial correlation strength matrix is constructed based on the shortest path distance. The shorter the path distance, the higher the initial correlation strength. The initial correlation strength between navigation and music playback is 0.5, and the initial correlation strength between navigation and phone function is 0.33.
[0178] The probability of interactive migration from navigation to music playback is 0.65, and the probability of interactive migration from navigation to phone is 0.15. The interactive migration probability is weighted and combined with the initial association strength (the weight ratio is 7:3) to generate a dynamic association matrix. The dynamic association strength between navigation and music playback is 0.65×0.7+0.5×0.3=0.605, and the dynamic association strength between navigation and phone is 0.15×0.7+0.33×0.3=0.204.
[0179] Based on the main eigenvector, the functional modules are adaptively clustered to determine the functional module groups. Through cluster analysis, the navigation function and music playback function are divided into one functional module group (navigation-entertainment group), the telephone function and message function are divided into another functional module group (communication group), and the air conditioning control and seat adjustment are divided into the comfort control group. The intra-group association density and inter-group association sparsity of each functional module group are calculated. The intra-group association density of the navigation-entertainment group is 0.78, the intra-group association density of the communication group is 0.82, and the intra-group association density of the comfort control group is 0.85. The inter-group association sparsity between the navigation-entertainment group and the communication group is 0.35, and the inter-group association sparsity between the navigation-entertainment group and the comfort control group is 0.42.
[0180] The projection position is determined by the intra-group association density and the functional module interaction experience optimization index. The intra-group association density of the comfort control group is 0.85, the average interaction experience optimization index is 92 points, and it is projected to the S-level area; the intra-group association density of the navigation-entertainment group is 0.78, the average interaction experience optimization index is 81.5 points, and it is projected to the A-level area; the intra-group association density of the communication group is 0.82, the average interaction experience optimization index is 75 points, and it is projected to the B-level area.
[0181] The inter-group correlation sparsity between the navigation-entertainment group and the communication group is 0.35, and the layout spacing is set to 15% of the screen width; the inter-group correlation sparsity between the navigation-entertainment group and the comfort control group is 0.42, and the layout spacing is set to 20% of the screen width.
[0182] For the air-conditioning control function located in the S-class area, the touch response area is increased by 25%, and the display level is adjusted to the highest level; for the navigation function located in the A-class area, the touch response area is increased by 15%, and the display level is adjusted to the second highest level; for the communication function located in the B-class area, the touch response area remains unchanged, but the icon size is enlarged by 10% to improve visibility.
[0183] In an optional implementation, calculating the interaction migration probability between the functional modules according to the interaction operation data, and weighting and combining the interaction migration probability with the initial association strength to generate a dynamic association matrix includes:
[0184] Projecting the operation time series in the interactive operation data onto a two-dimensional plane to generate an operation trajectory point set; calculating the velocity vector and acceleration vector between adjacent trajectory points in the operation trajectory point set; and counting the residence time of the operation trajectory point set in the neighborhood of the functional module and the number of switching times between the functional modules;
[0185] Determining the type of operation intention according to the velocity vector and the acceleration vector;
[0186] Establishing an operation intention counter for each of the functional modules, and counting the counting results of each type of operation intention; and calculating the probability of interaction migration between the functional modules based on the counting results;
[0187] Setting a sliding time window, performing exponential smoothing on the interaction migration probability sequence within the sliding time window, and determining a combined weight of the interaction migration probability and the initial association strength, wherein the combined weight is positively correlated with the counting result within the sliding time window;
[0188] When new interactive operation data enters the sliding time window, the interactive migration probability and the combination weight are recalculated, and the weighted combination is re-performed to update the dynamic association matrix.
[0189] In a specific implementation, the interactive operation data of the user on the application interface is obtained, including the timestamp and coordinate information of operations such as clicking, sliding, and staying. These operation time series are projected onto a two-dimensional plane to generate an operation trajectory point set. For example, the operation coordinates of the user at time points t1 to t5 are (120, 350), (125, 355), (130, 360), (135, 365), and (140, 370), respectively, and these coordinate points constitute the user's operation trajectory point set.
[0190] The velocity vector represents the speed and direction of the user's operation, and the acceleration vector represents the speed and direction of the speed change. Taking the above trajectory points as an example, the velocity vector from t1 to t2 is (5, 5) / Δt, where Δt is the time difference between t2-t1; the velocity vector from t2 to t3 is (5, 5) / Δt. The acceleration vector from t1 to t2 and then to t3 is the change in the velocity vector divided by the time difference.
[0191] For example, the user stays in the "message list" function module for 15 seconds, then switches to the "contact" function module for 10 seconds, and then switches back to the "message list" function module. These dwell time and switching behaviors are recorded as basic data for calculating the probability of interactive migration.
[0192] According to the calculated velocity vector and acceleration vector, the user's operation intention type is determined. The operation intention type can be divided into: browsing type (moderate speed, small acceleration), search type (fast speed, large acceleration), precise operation type (slow speed, small acceleration), etc. The judgment standard can be set as follows: when the velocity vector modulus is less than 5 pixels / second and the acceleration vector modulus is less than 2 pixels / second², it is determined to be a precise operation type; when the velocity vector modulus is greater than 20 pixels / second and the acceleration vector modulus is greater than 10 pixels / second², it is determined to be a search type; other cases are determined to be browsing type.
[0193] For each functional module, an operation intention counter is established to count the counting results of each operation intention type. For example, the browsing operation count of the "message list" functional module is 25 times, the search operation count is 10 times, and the precise operation count is 15 times; the browsing operation count of the "contact" functional module is 20 times, the search operation count is 5 times, and the precise operation count is 10 times.
[0194] Based on the counting results, the probability of interactive migration between functional modules is calculated. The probability of interactive migration indicates the possibility of a user switching from one functional module to another. The calculation method is: the number of switches from functional module A to functional module B divided by the total number of switches from functional module A to any functional module. For example, the number of times a user switches from the "Message List" functional module to the "Contacts" functional module is 15 times, and the total number of switches from the "Message List" functional module to any functional module is 30 times, then the probability of interactive migration from "Message List" to "Contacts" is 15 / 30=0.5.
[0195] Set a sliding time window and perform exponential smoothing on the interaction migration probability sequence within the window. The sliding time window can be set to the operation data of the last 30 minutes. Exponential smoothing is a method of weighted averaging time series data, with recent data having a higher weight than long-term data. For example, if the smoothing factor α is 0.3, the current interaction migration probability is P_t, and the smoothed interaction migration probability at the previous moment is S_(t-1), then the current smoothed interaction migration probability S_t = α×P_t + (1-α)×S_(t-1).
[0196] Determine the combined weight of the probability of interactive migration and the initial association strength, and the combined weight is positively correlated with the count result in the sliding time window. The initial association strength is the degree of association between functional modules preset based on design and historical data. The combined weight calculation method is: when the total operation count in the sliding time window exceeds the threshold (such as 100 times), the combined weight β is set to 0.8; when the operation count is between 50 and 100 times, β is set to 0.6; when the operation count is less than 50 times, β is set to 0.4.
[0197] The interactive migration probability and the initial association strength are weighted together to generate a dynamic association matrix. The weighted combination is calculated as follows: dynamic association strength = β×interactive migration probability + (1-β)×initial association strength. For example, the interactive migration probability from "message list" to "contact" is 0.5, the initial association strength is 0.3, and the combination weight β is 0.6, then the dynamic association strength is 0.6×0.5 + 0.4×0.3 = 0.42.
[0198] When new interactive operation data enters the sliding time window, the interactive migration probability and combination weight are recalculated, and the weighted combination is re-calculated to update the dynamic association matrix. For example, the newly added operation data shows that the number of times the user switches from "message list" to "contact" has increased by 5 times, and the total number of times the user switches from "message list" to any functional module has increased by 8 times. The updated interactive migration probability is (15+5) / (30+8)=0.53. Assuming that the updated combination weight β is 0.7, the updated dynamic association strength is 0.7×0.53 + 0.3×0.3 = 0.461.
[0199] like Figure 4As shown in the figure, the comprehensive performance of different methods is intuitively demonstrated through the correspondence between the two key indicators of classification accuracy and calculation time. The concentric rounded rectangular thermal contours are used in the figure to represent the performance contours. The closer to the upper right corner area, the better the performance. It can be clearly observed from the scattered distribution that the three operation intention classifications of this technical solution (triangle mark) are all in the optimal performance area, among which the precise operation intention recognition has achieved the highest accuracy of 0.928 (calculation time 41 milliseconds), the browsing intention recognition is closely followed by 0.915 accuracy (calculation time 39 milliseconds), and the search intention recognition also has an excellent performance of 0.872 (calculation time 37 milliseconds). In contrast, although the traditional speed threshold method (circular mark) has the advantage of the shortest calculation time (only 13-16 milliseconds), it performs poorly in terms of accuracy. Its browsing, search and precise operation intention recognition accuracy rates are only 0.752, 0.658 and 0.674 respectively. The machine learning method (square mark) has a moderate performance in terms of accuracy (0.769-0.836), but the calculation time is significantly higher than other methods (55-60 milliseconds), causing its scatter points to be distributed in the upper left area of the chart. The rule engine method (diamond mark) is located in the middle performance area, with both accuracy (0.697-0.794) and calculation time (24-27 milliseconds) at a medium level. Through this heat map analysis, we can intuitively conclude that this technical solution achieves the best balance between accuracy and computing efficiency, forming a clear performance advantage in the range of 0.87-0.93 accuracy and 37-41 milliseconds calculation time, especially in the outstanding performance in precise operational intent recognition, providing a better choice for interactive intent recognition in complex operational scenarios.
[0200] Interaction behavior analysis is an important research direction in the field of human-computer interaction. Traditional methods mainly rely on preset rules and fixed thresholds to judge user intentions. Early implementations usually use static association matrices to represent the connection relationship between functional modules, and the matrix elements represent fixed transition probabilities or weight values. This method is artificially defined during the initial design of the system, lacks the ability to adapt to the actual usage habits of users, and is difficult to reflect the personalized needs of different users and the impact of changes in the usage environment.
[0201] In the existing technology, some improvement schemes try to introduce simple statistical methods, such as recording basic data such as user click frequency and dwell time, but these methods usually only consider single-dimensional interaction features and ignore the comprehensive analysis of multi-dimensional interaction behaviors. For example, some systems only judge the importance of functions by the number of clicks, or only judge user interests by the page dwell time. These single features cannot fully describe the user's operation intentions and usage habits.
[0202] The solution of this embodiment projects the operation time series onto a two-dimensional plane and generates a set of trajectory points, and analyzes the dynamic characteristics of the user's operation by calculating the velocity vector and acceleration vector. This method is derived from human behavioral research, which believes that users will show different operation speeds and acceleration characteristics under different operation intentions. For example, when the goal is clear, the operation speed is faster and the acceleration changes little, while exploratory operations are characterized by slower speeds and frequent acceleration changes. This analysis method based on the principles of physical kinematics can reveal the cognitive process behind the user's operation behavior from a more microscopic perspective.
[0203] Compared with the existing technology, the core improvement of this embodiment is to introduce sliding time window and exponential smoothing technology to dynamically adjust the probability of interaction migration. Traditional methods usually use cumulative statistics or simple averages to update interaction data, resulting in a lag in the system's response to changes in user behavior. The sliding time window mechanism enables the system to focus on recent interaction behaviors, while the exponential smoothing technology enables the system to capture the changing trends of user habits in a timely manner by giving higher weights to recent data, while avoiding drastic fluctuations caused by occasional behaviors.
[0204] Another important improvement is the dynamic adjustment mechanism of the combined weights. The system adaptively adjusts the combined ratio of the interaction migration probability and the initial association strength according to the amount of interaction data in the sliding time window. When there is sufficient user interaction data, the system tends to rely more on the actual interaction data; when the interaction data is sparse, it retains more of the influence of the initial association strength. This adaptive weight mechanism solves the limitations of pure data-driven methods in new features or cold start scenarios.
[0205] Through the above improvements, the solution of this embodiment can more accurately capture the user's operation intention, dynamically adjust the association strength between functional modules, and provide users with an interactive experience that is more in line with personal usage habits. Experimental results show that compared with the traditional static association matrix, this solution can improve the accuracy of function recommendations by more than 20%, shorten the user operation path by an average of 35%, and significantly improve the system's usability and user satisfaction. At the same time, the adaptive characteristics of the dynamic association matrix enable the system to maintain good adaptability under different user groups and different usage scenarios, providing strong support for personalized human-computer interaction.
[0206] In this embodiment, speed, acceleration, dwell time, switching times and other indicators are used to effectively identify operation intentions and achieve accurate analysis of interactive behaviors. By constructing an operation intention counter and calculating the probability of interactive migration between functional modules, combined with an exponential smoothing mechanism within a sliding time window, real-time updating of the dynamic association matrix is achieved, thereby improving the system's responsiveness to new data. Based on the combined weight of the interactive migration probability and the initial association strength, the positive correlation of the counting results is reflected, so that the model can adaptively adjust the prediction effect when faced with constantly changing interactive data.
[0207] The vehicle-mounted display terminal human-machine interface interactive experience optimization system of the embodiment of the present invention includes:
[0208] The first unit is used to obtain the interactive behavior data of the driver, wherein the interactive behavior data includes eye gaze point position data, hand operation trajectory data and head rotation angle data;
[0209] The second unit is used to fuse the eye gaze point position data with the head rotation angle data, and adjust the weights in combination with the hand operation trajectory data to generate a driver's attention distribution heat map; based on the heat values of each area in the attention distribution heat map, calculate the importance score of each functional module of the vehicle display terminal;
[0210] The third unit is used to collect vehicle operation status information and environmental perception information, build a dynamic scene recognition model, identify the current driving scene type in real time through a deep learning network, and output the usage probability prediction value of the functional module; the importance score and the usage probability prediction value are weighted and integrated to obtain the functional module interaction experience optimization index;
[0211] The fourth unit is used to adjust the corresponding functional modules to the hot spot area of the attention distribution heat map according to the preset index threshold, and group and cluster the functional modules based on the correlation degree, and dynamically adjust the touch response area and display level of the functional modules; when the change in the interactive behavior data exceeds the preset change threshold, it triggers the re-execution of the optimization step.
[0212] According to a third aspect of the embodiments of the present invention,
[0213] An electronic device is provided, comprising:
[0214] processor;
[0215] a memory for storing processor-executable instructions;
[0216] The processor is configured to call the instructions stored in the memory to execute the aforementioned method.
[0217] A fourth aspect of the embodiments of the present invention is:
[0218] A computer-readable storage medium is provided, on which computer program instructions are stored. When the computer program instructions are executed by a processor, the aforementioned method is implemented.
[0219] The present invention may be a method, an apparatus, a system and / or a computer program product. The computer program product may include a computer-readable storage medium carrying computer-readable program instructions for executing various aspects of the present invention.
[0220] Finally, it should be noted that the above embodiments are only used to illustrate the technical solutions of the present invention, rather than to limit it. Although the present invention has been described in detail with reference to the aforementioned embodiments, those skilled in the art should understand that they can still modify the technical solutions described in the aforementioned embodiments, or replace some or all of the technical features therein by equivalents. However, these modifications or replacements do not cause the essence of the corresponding technical solutions to deviate from the scope of the technical solutions of the embodiments of the present invention.
Claims
1. A method for optimizing the interactive experience of a human-machine interface of an in-vehicle display terminal, characterized in that: include: Acquiring interactive behavior data of the driver, wherein the interactive behavior data includes eye gaze point position data, hand operation trajectory data, and head rotation angle data; The eye gaze position data and the head rotation angle data are fused and calculated, and the weights are adjusted in combination with the hand operation trajectory data to generate a heat map of the driver's attention distribution; Based on the heat value of each area in the attention distribution heat map, calculate the importance score of each functional module of the vehicle display terminal; Collect vehicle operation status information and environmental perception information, build a dynamic scene recognition model, use the deep learning network to identify the current driving scene type in real time, and output the usage probability prediction value of the functional module; The importance score and the usage probability prediction value are weighted and integrated to obtain the functional module interaction experience optimization index; According to the preset index threshold, the corresponding functional modules are adjusted to the hot spot area of the attention distribution heat map, and grouped and clustered based on the correlation of the functional modules, and the touch response area and display level of the functional modules are dynamically adjusted, including: Receive the functional module interaction experience optimization index and the attention distribution heat map, and select the functional modules to be adjusted from the functional module interaction experience optimization index based on a preset index threshold; divide the attention distribution heat map into multiple regional levels, and sort the functional modules by importance according to the functional module interaction experience optimization index; extract the interactive operation data of the functional module, including the operation time series, the operation interval duration, and the operation transfer probability; Acquire the hierarchical relationship tree of the functional modules, calculate the shortest path distance between the functional modules, and construct the initial association strength based on the shortest path distance; calculate the interaction migration probability between the functional modules according to the interaction operation data, perform weighted combination of the interaction migration probability and the initial association strength, and generate a dynamic association matrix; Performing eigendecomposition on the dynamic association matrix to obtain a main eigenvector; adaptively clustering the functional modules based on the main eigenvector to determine a functional module group; calculating the intra-group association density and inter-group association sparsity of each functional module group; Projecting the functional module group to the regional level of the attention distribution heat map, the projection position is jointly determined by the intra-group association density and the functional module interaction experience optimization index; determining the relative layout spacing of the functional module group based on the inter-group association sparsity; dynamically adjusting the touch response area and display level according to the position of the functional module in the attention distribution heat map; When the change in the interactive behavior data exceeds a preset change threshold, the optimization step is triggered to be re-executed.
2. The method according to claim 1, characterized in that: The eye gaze point position data and the head rotation angle data are fused and calculated, and the weights are adjusted in combination with the hand operation trajectory data to generate a heat map of the driver's attention distribution, including: The pupil center coordinates are obtained based on the driver's eye gaze point position data, the high-frequency jitter data in the pupil center coordinates are eliminated by Kalman filtering, and the pupil center coordinates are smoothed by using a RANSAC algorithm to obtain continuous gaze point trajectory data; According to the yaw angle and translation vector in the head rotation angle data, the continuous gaze point trajectory data is compensated for spatial coordinates to obtain gaze point coordinate compensation data; Through the hand positioning points in the hand operation trajectory data, the three-dimensional space coordinate sequence is extracted, the hand movement speed, movement acceleration, movement direction angle, displacement distance and gesture type are calculated, and the hand trajectory feature vector is constructed; Using a dynamic time warping algorithm, the gaze point coordinate compensation data, the head rotation angle data, and the hand trajectory feature vector are time-series aligned to construct a multimodal feature tensor; Calculating the weight coefficient of each feature in the multimodal feature tensor, and combining the weight coefficient with a Gaussian kernel function to generate attention density distribution data; Within a preset time window, the density feature vector of the attention density distribution data is extracted and time series fusion is performed through exponential decay weights. A dual-domain regularization function is constructed to eliminate abnormal points. The initial hot zone boundary point set is determined based on an adaptive segmentation threshold, and the boundary is guided to dynamically expand to generate a driver's attention distribution heat map.
3. The method according to claim 2, characterized in that In a preset time window, the density feature vector is extracted from the attention density distribution data and time series fusion is performed through exponential decay weights, a dual-domain regularization function is constructed to eliminate abnormal points, an initial hot zone boundary point set is determined based on an adaptive segmentation threshold, and the boundary is guided to dynamically expand, and a driver's attention distribution heat map is generated, including: The statistical features and spatial moment features of the attention density distribution data are extracted within the preset time window to construct a density feature vector. The density feature vectors of the historical window and the current window are fused through the exponential decay weight to obtain a time series fusion feature vector. Determine the coordinates of the attention center point according to the time series fusion feature vector, construct a dual-domain regularization function containing spatial distance constraints and density value constraints, and perform regularization processing to obtain the optimized data of attention density distribution; Calculate the adaptive segmentation threshold based on the attention density distribution optimization data and determine the initial hot zone boundary point set; Calculate the density gradient direction and density value of the initial hot zone boundary point set, and set the main threshold and secondary threshold in the corresponding eight neighborhoods along the density gradient direction; When the density value of the eight neighborhoods is greater than the main threshold, the corresponding points are included in the initial hot zone boundary point set; When the density value of the eight neighborhoods is between the primary threshold and the secondary threshold, the corresponding points are marked as candidate boundary point sets; the curvature value of the initial hot zone boundary point set is calculated to construct a dynamic scaling function, and the primary threshold and the secondary threshold are adjusted; the shortest path distance from the candidate boundary point set to the initial hot zone boundary point set is calculated, and the candidate boundary point set whose shortest path distance is less than the preset connectivity threshold is merged into the initial hot zone boundary point set; When the area growth rate of the initial hot zone boundary point set is less than the preset growth threshold, and the number of candidate boundary point sets is less than the preset total number threshold, the boundary expansion is stopped to obtain the expanded hot zone boundary point set; The kernel density estimation and normalization processing are performed on the boundary point set of the expansion hot zone to obtain the driver's attention distribution heat map.
4. The method according to claim 1, characterized in that Collect vehicle operation status information and environmental perception information, build a dynamic scene recognition model, and use the deep learning network to identify the current driving scene type in real time. The output function module usage probability prediction values include: Acquire vehicle operation status information and environmental perception information, calculate the volatility of the vehicle operation status information and environmental perception information based on a sliding time window, and determine the volatility by the ratio of the change amplitude of the characteristic values at adjacent time points to the average change amplitude in the time window; divide the scene sequence into a steady-state interval and a mutation interval according to a preset fluctuation threshold; perform feature sampling of different frequencies on the steady-state interval and the mutation interval, respectively, to obtain an adaptive sampling feature sequence; Matching the adaptive sampling feature sequence with historical scene samples, constructing and solving the scene evolution equation based on the element coupling relationship in the historical scene samples and the stability of the scene prototype, obtaining the state association matrix, and dynamically updating the dynamic coefficient using the prediction deviation; In the adaptive sampling feature sequence, a sampling moment whose prediction deviation is greater than a preset deviation threshold is selected as a key scene frame; the features of the key scene frame are input into a pre-trained scene classifier to obtain a scene type recognition result; and based on the scene type recognition result and the prediction deviation, a usage probability prediction value of the corresponding functional module is calculated.
5. The method according to claim 4, characterized in that The adaptive sampling feature sequence is matched with the historical scene samples, the scene evolution equation is constructed and solved based on the element coupling relationship in the historical scene samples and the stability of the scene prototype, the state association matrix is obtained, and the dynamic coefficient is dynamically updated using the prediction deviation, including: Extracting scene elements from the historical scene samples, wherein the scene elements include vehicles, pedestrians, and road facilities; constructing an element coupling matrix, wherein each element of the element coupling matrix represents an interaction influence intensity value between adjacent scene elements; Clustering is performed based on the historical scene samples to obtain a plurality of scene prototypes; obtaining the eigenvalue of each scene prototype in the element coupling matrix, and calculating the stability index of each scene prototype in combination with the spatial distribution density of the scene elements; Constructing a scenario evolution equation including a steady-state term and a disturbance term, wherein the coefficient of the steady-state term comes from the stability index, and the coefficient of the disturbance term comes from the element coupling matrix; solving the scenario evolution equation to obtain the state transition probability between scenario prototypes, and constructing a state association matrix; The adaptive sampling feature sequence is input into the state association matrix for feature prediction to obtain predicted features; the predicted features are compared with the actual features to calculate the prediction deviation, and based on the prediction deviation, the interaction influence intensity value in the element coupling matrix and the disturbance term coefficient in the scene evolution equation are updated respectively.
6. The method according to claim 1, characterized in that Calculating the interaction migration probability between the functional modules according to the interaction operation data, and weighting and combining the interaction migration probability with the initial association strength to generate a dynamic association matrix includes: Projecting the operation time series in the interactive operation data onto a two-dimensional plane to generate an operation trajectory point set; calculating the velocity vector and acceleration vector between adjacent trajectory points in the operation trajectory point set; and counting the residence time of the operation trajectory point set in the neighborhood of the functional module and the number of switching times between the functional modules; Determining the type of operation intention according to the velocity vector and the acceleration vector; Establishing an operation intention counter for each of the functional modules, and counting the counting results of each type of operation intention; and calculating the probability of interaction migration between the functional modules based on the counting results; Setting a sliding time window, performing exponential smoothing on the interaction migration probability sequence within the sliding time window, and determining a combined weight of the interaction migration probability and the initial association strength, wherein the combined weight is positively correlated with the counting result within the sliding time window; When new interactive operation data enters the sliding time window, the interactive migration probability and the combination weight are recalculated, and the weighted combination is re-performed to update the dynamic association matrix.
7. A vehicle-mounted display terminal human-machine interface interactive experience optimization system, used to implement the method described in any one of claims 1 to 6, characterized in that: include: The first unit is used to obtain the interactive behavior data of the driver, wherein the interactive behavior data includes eye gaze point position data, hand operation trajectory data and head rotation angle data; The second unit is used to fuse the eye gaze point position data with the head rotation angle data, and adjust the weights in combination with the hand operation trajectory data to generate a driver's attention distribution heat map; based on the heat values of each area in the attention distribution heat map, calculate the importance score of each functional module of the vehicle display terminal; The third unit is used to collect vehicle operation status information and environmental perception information, build a dynamic scene recognition model, identify the current driving scene type in real time through a deep learning network, and output the predicted value of the use probability of the functional module; The importance score and the usage probability prediction value are weighted and integrated to obtain the functional module interaction experience optimization index; The fourth unit is used to adjust the corresponding functional modules to the hot spot area of the attention distribution heat map according to the preset index threshold, and group and cluster the functional modules based on the correlation degree, and dynamically adjust the touch response area and display level of the functional modules; when the change in the interactive behavior data exceeds the preset change threshold, it triggers the re-execution of the optimization step.
8. An electronic device, characterized in that: include: processor; a memory for storing processor-executable instructions; The processor is configured to call the instructions stored in the memory to execute the method according to any one of claims 1 to 6.
9. A computer-readable storage medium having computer program instructions stored thereon, characterized in that: When the computer program instructions are executed by a processor, the method according to any one of claims 1 to 6 is implemented.
Citation Information
Patent Citations
Driving behavior state recognition method and device, equipment and storage medium
CN114998870A
Controllable image description generation system and method fusing eye movement attention
CN116185182A