Method and apparatus for establishing eye movement dataset based on multiple traffic operation scenarios
By generating and analyzing eye-tracking video data and combining it with LSD-SLAM technology, an eye-tracking dataset for multiple traffic operation scenarios was constructed. This solved the problems of lack of systematic data collection and low annotation accuracy in the data system, achieving both rich data and accurate annotation.
Patent Information
- Application Number
- CN202510023963.1
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Filing Date
- 2025-01-07
- Publication Date
- 2025-11-25
- Estimated Expiration
- 2045-01-07
AI Technical Summary
In existing technologies, data systems that combine eye-tracking with operational scenarios lack systematic data collection and analysis, resulting in scarce data, a lack of diverse traffic operation scenarios, and low accuracy of labeled data.
By generating eye-tracking video data of target drivers in multiple traffic operation scenarios, using the gaze heatmap of each frame of eye-tracking data, and combining it with a pre-built LSD-SLAM, semantic segmentation information and 3D reconstruction information are obtained to construct an eye-tracking dataset that covers multiple traffic operation scenarios and performs data classification and key event annotation.
It effectively expands the diversity and accuracy of data collection, improves the annotation quality of eye-tracking data, solves the systematization problem of the data system, and realizes the construction of efficient and accurate eye-tracking datasets.
Smart Images

Figure CN119964108B_ABST
Abstract
Description
TECHNICAL FIELD
[0001] The present application relates to the technical field of vehicle driving, in particular to a method and device for establishing eye movement data set based on multiple traffic operation scenarios. BACKGROUND
[0002] Human vision can provide key insights for autonomous driving robots to improve visual perception algorithms and enhance algorithm performance. Experienced drivers can quickly identify and locate dangerous or important visual signals occurring in a blurred peripheral field of view and perform eye movements on these areas. By analyzing where drivers tend to focus their attention while driving, adjusting the sensor data processing of autonomous driving robots, and thus significantly improving the performance of autonomous driving robot models.
[0003] In related technologies, vehicle behavior features and eye movement speed in vehicle multi-dimensional data and driver eye movement data can be extracted to construct a dangerous behavior recognition model, recognize dangerous driving behavior, and generate a warning signal when dangerous driving behavior occurs. Eye trackers can also be used to collect driver pupil diameter data, and the pupil diameter in the traffic conflict period is used as a template for Fourier transform to generate template spectrum data, and then a target threshold is set to quickly identify traffic conflicts.
[0004] However, the data system of eye movement combined with operation scenarios in related technologies lacks systematic data collection and analysis, resulting in poor data, lack of diverse traffic operation scenarios, and low accuracy of annotated data, which needs to be improved. SUMMARY
[0005] The present application provides a method and device for establishing eye movement data set based on multiple traffic operation scenarios to solve the problems of poor data, lack of diverse traffic operation scenarios, and low accuracy of annotated data in related technologies due to the lack of systematic data collection and analysis of the data system of eye movement combined with operation scenarios.
[0006] The first aspect embodiment of the present application provides a method for establishing eye movement data set based on multiple traffic operation scenarios, comprising the following steps: generating eye movement video data of a target driver in multiple traffic operation scenarios according to scene video data of the multiple traffic operation scenarios; obtaining a fixation hotspot map corresponding to each frame of eye movement data in the eye movement video data based on the eye movement video data; obtaining semantic segmentation information and three-dimensional reconstruction information of the fixation hotspot map from a pre-constructed LSD-SLAM (Large-Scale Direct SLAM) based on the fixation hotspot map, to construct an eye movement data set of the target driver based on the semantic segmentation information and the three-dimensional reconstruction information.
[0007] Optionally, in an embodiment of the present application, the obtaining, based on the eye movement video data, a gaze heat map corresponding to each frame of eye movement data in the eye movement video data comprises: extracting pupil center information of the target driver in each frame of eye movement data in the each frame of eye movement data; obtaining line-of-sight estimation information of the target driver in the each frame of eye movement data according to the pupil center information; and obtaining the gaze heat map based on the line-of-sight estimation information.
[0008] Optionally, in an embodiment of the present application, the obtaining, based on the gaze heat map, semantic segmentation information and three-dimensional reconstruction information of the gaze heat map from a pre-constructed LSD-SLAM to construct an eye movement data set of the target driver based on the semantic segmentation information and the three-dimensional reconstruction information comprises: generating two-dimensional semantic segmentation information and a two-dimensional semantic label corresponding to each frame of eye movement data based on a SLAM in the pre-constructed LSD-SLAM; dividing the gaze heat map using the two-dimensional semantic label to obtain a first gaze heat map, first semantic segmentation information corresponding to the first gaze heat map, a second gaze heat map, and second semantic segmentation information corresponding to the second gaze heat map; obtaining an optimized first gaze heat map corresponding to the first gaze heat map based on the first gaze heat map, the first semantic segmentation information, the second gaze heat map, and the second semantic segmentation information; obtaining three-dimensional reconstruction information of the first gaze heat map based on the optimized first gaze heat map and the first semantic segmentation information in combination with a recursive Bayesian and conditional random field in the pre-constructed LSD-SLAM; and generating the eye movement data set based on the three-dimensional reconstruction information and the optimized first gaze heat map.
[0009] Optionally, in an embodiment of the present application, the generating, based on scene video data of a multi-traffic operation scene, eye movement video data of a target driver in the multi-traffic operation scene comprises: determining a key event in the multi-traffic operation scene; extracting scene key video frame data in the scene video data based on the key event; and determining the eye movement video data according to a scene timestamp corresponding to the scene key video frame data.
[0010] Optionally, in an embodiment of the present application, the determining the eye movement video data according to a scene timestamp corresponding to the scene key video frame data comprises: obtaining initial eye movement video data of the target driver in the multi-traffic operation scene according to the scene video data; and determining the eye movement video data based on the initial eye movement video data and the scene timestamp.
[0011] The second aspect embodiment of the application provides a device for establishing an eye movement data set based on a multi-traffic running scene, comprising: a generation module configured to generate eye movement video data of a target driver in the multi-traffic running scene according to scene video data of the multi-traffic running scene; an acquisition module configured to acquire a gaze hotspot map corresponding to each frame of eye movement data in the eye movement video data based on the eye movement video data; and a construction module configured to obtain semantic segmentation information and three-dimensional reconstruction information of the gaze hotspot map from a pre-constructed LSD-SLAM based on the gaze hotspot map, and to construct an eye movement data set of the target driver based on the semantic segmentation information and the three-dimensional reconstruction information.
[0012] Optionally, in an embodiment of the application, the acquisition module comprises: a first extraction unit configured to extract pupil center information of the target driver in each frame of eye movement data in the each frame of eye movement data; an acquisition unit configured to acquire line-of-sight estimation information of the target driver in the each frame of eye movement data according to the pupil center information; and a first generation unit configured to obtain the gaze hotspot map based on the line-of-sight estimation information.
[0013] Optionally, in an embodiment of the application, the construction module comprises: a second generation unit configured to generate two-dimensional semantic segmentation information and a two-dimensional semantic label corresponding to each frame of eye movement data based on SLAM in the pre-constructed LSD-SLAM; a division unit configured to divide the gaze hotspot map using the two-dimensional semantic label to obtain a first gaze hotspot map, first semantic segmentation information corresponding to the first gaze hotspot map, a second gaze hotspot map, and second semantic segmentation information corresponding to the second gaze hotspot map; a third generation unit configured to obtain an optimized first gaze hotspot map corresponding to the first gaze hotspot map based on the first gaze hotspot map, the first semantic segmentation information, the second gaze hotspot map, and the second semantic segmentation information; a fourth generation unit configured to obtain three-dimensional reconstruction information of the first gaze hotspot map by combining recursive Bayesian and conditional random fields in the pre-constructed LSD-SLAM based on the optimized first gaze hotspot map and the first semantic segmentation information; and a fifth generation unit configured to generate the eye movement data set based on the three-dimensional reconstruction information and the optimized first gaze hotspot map.
[0014] Optionally, in an embodiment of the application, the generation module comprises: a first determination unit configured to determine a key event in the multi-traffic running scene; a second extraction unit configured to extract scene key video frame data in the scene video data based on the key event; and a second determination unit configured to determine the eye movement video data according to a scene timestamp corresponding to the scene key video frame data.
[0015] Optionally, in an embodiment of the present application, the second determining unit comprises: an obtaining sub-unit, configured to obtain initial eye movement video data of the target driver in the multi-traffic running scene according to the scene video data; and a determining sub-unit, configured to determine the eye movement video data based on the initial eye movement video data and the scene timestamp.
[0016] A third aspect of the present application provides an electronic device, comprising a memory, a processor, and a computer program stored in the memory and executable on the processor, wherein the processor executes the program to implement the method for establishing an eye movement dataset based on a multi-traffic running scene as described in the above embodiments.
[0017] A fourth aspect of the present application provides a computer-readable storage medium, which stores a computer program executable by a processor to implement the method for establishing an eye movement dataset based on a multi-traffic running scene as described above.
[0018] A fifth aspect of the present application provides a computer program product, comprising a computer program executable to implement the method for establishing an eye movement dataset based on a multi-traffic running scene as described above.
[0019] The embodiments of the present application can generate eye movement video data according to scene video data of a multi-traffic running scene, and then utilize a fixation heat map corresponding to each frame of eye movement data to obtain semantic segmentation information and three-dimensional reconstruction information from a pre-constructed LSD-SLAM, so as to construct an eye movement dataset of a target driver, which not only covers conventional situations of robot lane-following and robot vehicle-tracking, but also extends to multi-traffic running scenes such as intersections and bicycle intrusions into motor vehicle lanes, and performs traffic scene generalization and data classification, and through smoothing of eye movement video data of the target driver, removal of sky noise by SLAM, labeling of key events, and efficient and accurate establishment of an eye movement dataset of the target driver in a multi-traffic running scene. Thus, the problems in the related art that an eye movement combined with a running scene lacks systematic data collection and analysis, resulting in poor data, lack of diverse traffic running scenes, and low accuracy of labeled data are solved.
[0020] Additional aspects and advantages of the present application will be in part apparent and in part pointed out hereinafter. BRIEF DESCRIPTION OF DRAWINGS
[0021] The above and / or additional aspects and advantages of the present application will become apparent and be readily appreciated from the following description, including the appended drawings.
[0022] Figure 1A flowchart of a method for establishing an eye movement dataset based on a multi-traffic running scenario according to an embodiment of the present application is provided.
[0023] Figure 2 A block schematic diagram of a driving simulation system according to an embodiment of the present application is provided.
[0024] Figure 3 A schematic diagram of an eye tracker according to an embodiment of the present application is provided.
[0025] Figure 4 A schematic diagram of a hotspot example according to an embodiment of the present application is provided.
[0026] Figure 5 A schematic diagram of another hotspot example according to an embodiment of the present application is provided.
[0027] Figure 6 A schematic diagram of a gaze trajectory according to an embodiment of the present application is provided.
[0028] Figure 7 A flowchart of constructing an eye movement dataset of a target driver by a pre-constructed LSD-SLAM according to an embodiment of the present application is provided.
[0029] Figure 8 A flowchart of semantic mapping according to an embodiment of the present application is provided.
[0030] Figure 9 A block schematic diagram of an apparatus for establishing an eye movement dataset based on a multi-traffic running scenario according to an embodiment of the present application is provided.
[0031] Figure 10 A structural schematic diagram of an electronic device according to an embodiment of the present application is provided. DETAILED DESCRIPTION
[0032] Embodiments of the present application are described in detail below with reference to the accompanying drawings, in which examples of the embodiments are shown, and the same or similar notations are used to denote the same or similar elements throughout. The embodiments described below by reference to the drawings are examples, and are intended to explain the present application, and cannot be understood as limiting the present application.
[0033] The method and device for establishing an eye movement data set based on multiple traffic operation scenarios according to the embodiments of the present application are described below with reference to the accompanying drawings. In view of the lack of systematic data collection and analysis in the data system combining eye movement and operation scenarios mentioned in the background art, which results in poor data, lack of diverse traffic operation scenarios, and low accuracy of annotated data, the present application provides a method for establishing an eye movement data set based on multiple traffic operation scenarios. In this method, eye movement video data can be generated according to scene video data of multiple traffic operation scenarios, and then the gaze heat map corresponding to each frame of eye movement data is used to obtain semantic segmentation information and three-dimensional reconstruction information from a pre-constructed LSD-SLAM, thereby constructing an eye movement data set of a target driver. In addition to covering conventional situations such as robot lane-following and robot vehicle-tracking, the method also expands multiple traffic operation scenarios such as intersections and bicycle intrusions into motor vehicle lanes, and performs traffic scenario generalization and data classification. By smoothing the target driver's eye movement video data, removing sky noise using SLAM, labeling and annotating key events, and efficiently and accurately establishing an eye movement data set of a target driver in multiple traffic operation scenarios. Thus, the problems of lack of systematic data collection and analysis in the data system combining eye movement and operation scenarios in related technologies, poor data, lack of diverse traffic operation scenarios, and low accuracy of annotated data are solved.
[0034] Specifically, Figure 1 A flowchart of a method for establishing an eye movement data set based on multiple traffic operation scenarios according to an embodiment of the present application is shown.
[0035] As Figure 1 shown, the method for establishing an eye movement data set based on multiple traffic operation scenarios includes the following steps:
[0036] In step S101, eye movement video data of a target driver in a multiple traffic operation scenario is generated according to scene video data of the multiple traffic operation scenario.
[0037] It can be understood that in the embodiments of the present application, the scene video data can be classified and annotated according to the location and event, and the multiple types of data of the same type of scene video data, eye movement video data, and various sensor data (including the target driver's operation behavior) are aligned and aggregated together in a time stamp method.
[0038] Further, in the embodiments of the present application, the eye movement video data can include, but is not limited to, three types of video data of the target driver's gaze data, saccade data, and pursuit movement data.
[0039] In addition, it should be noted that the target driver in the embodiments of the present application can be understood as a driver with more than 3 years of experience. The specific setting can be made by a person skilled in the art according to the actual situation, and the present application does not make specific limitations.
[0040] As a possible implementation manner, the embodiment of the application can obtain the eye movement video data generated by the target driver in the multi-traffic running scene.
[0041] Exemplarily, the embodiment of the application can generate the eye movement video data of the target driver in combination with Figure 2 the driving simulation system.
[0042] Specifically, as Figure 2 shown, the driving simulation system 20 comprises a display screen 201, a driving kit 202 and an eye tracker 203.
[0043] The display screen 201 can play the scene video data (multi-sensor actually collected data such as RGBD camera, or public driving video data set) of the multi-traffic running scene.
[0044] The driving kit 202 is equipped with a steering wheel to approach real driving, a seat similar to a car, and a foot pedal system for acceleration and deceleration, and records the acceleration and deceleration and steering behaviors of the target driver in real time during the whole process.
[0045] The eye tracker 203 is completed by a special eye tracking device and an eye movement analysis software for recording and monitoring the target driver, and an example diagram thereof is shown in Figure 3 .
[0046] Further, the embodiment of the application considers subsequent data verification for promoting actual traffic scene application, and the eye tracking device adopts a head-mounted eye tracker, which can but is not limited to setting a computing unit, an ET (eye camera) and an FT (scene camera), and the application does not make specific limitations.
[0047] In the embodiment of the application, the computing unit collects and processes FT and ET images through the FT and ET cameras, realizes accurate and real-time eyeball tracking, and transmits the processed FT images and eye movement data to the backend program through the network, and the backend program is responsible for the analysis and processing of the FT images and eye movement data. Data accuracy is very important, therefore, the embodiment of the application makes three calibrations before real data collection, and a short practical test is conducted after calibration to verify that the gaze tracking is accurate.
[0048] Further, after the embodiment of the application is verified to be accurate, the driving simulation system 20 and the eye tracker 203 are time-synchronized, the simulation driving and the recording start, and the driving simulation system 20 starts to collect the eye movement data. The scene video data and the eye movement video data are time-stamped in the later stage.
[0049] The eye movement analysis software can but is not limited to provide a hotspot map and a gaze trajectory map, and the application does not make specific limitations.
[0050] The hotspot map can be understood as being able to indicate the eye fixation points of the target driver within a few seconds, which is very useful for summarizing where the focus points are. The areas that are briefly or cautiously looked at are represented by blue shadows, and the areas that are highly focused on for a longer time are represented by red shadows, wherein Figure 4 and Figure 5 are examples of hotspot maps provided by embodiments of the present application, wherein Figure 4 It is described that a frame of the target driver's gaze is mainly concentrated on the pedestrian part; Figure 5 It is described that a frame is aggregated and smoothed by the gaze of multiple target drivers, covering a key traffic event.
[0051] The gaze trajectory map can be used to reveal the time sequence of the target driver's observation, the position of the observation, and the time of observing a certain position, which has explanatory and guiding significance for event-based driving behavior. The trajectory points show the spatiotemporal relationship of the target driver's attention information from a broad perspective (the scattered points need to be filtered), such as Figure 6 The gaze trajectory map shown in the figure.
[0052] Optionally, in an embodiment of the present application, the eye movement video data of the target driver in the multi-traffic running scene is generated according to the scene video data of the multi-traffic running scene, comprising: determining a key event in the multi-traffic running scene; extracting scene key video frame data in the scene video data based on the key event; determining the eye movement video data according to the scene time stamp corresponding to the scene key video frame data.
[0053] It can be understood that the key event in the embodiments of the present application can be understood as a brake event, a braking event, a turning event, a lane switching event, and an acceleration event, etc., which is not specifically limited in the present application.
[0054] In some embodiments, in order to effectively collect attention data in the eye movement video data, the embodiments of the present application can first determine the key event in the multi-traffic running scene, and then extract the scene key video frame data according to the key event, and determine the eye movement video data by using the scene time stamp.
[0055] It should be noted that in the embodiments of the present application, the visual attention of the target driver enables the target driver to quickly identify and locate potential dangers or important visual cues, such as a suddenly disappearing pedestrian, an intrusion of a nearby cyclist, or a change in a traffic light. That is, the scene video data of the embodiments of the present application not only includes driving running data under various weather and lighting conditions, but also includes crisis and rare traffic scenarios, such as various driving activities and environments including lane following, turning, lane switching, and parking in chaotic scenes.
[0056] Exemplarily, the embodiment of the present application can take a brake event, a braking event, a turning event, a lane switching event, an acceleration event and the like as a key event, and then video clips are made 6.5 seconds (which can also be other numerical values, and the present application does not make specific limitations) before and 3.5 seconds (which can also be other numerical values, and the present application does not make specific limitations) after the key event, and then more than 2000 scene video data are obtained, which contain a large number of different road users, and YOLO is used for detection. It should be noted that each scene video data of the embodiment of the present application usually contains a plurality of vehicles and / or pedestrians, and the present application does not make specific limitations.
[0057] Further, the embodiment of the present application extracts scene key video frame data in each scene video data according to the key event, and then obtains eye movement video data by using scene timestamps.
[0058] Optionally, in an embodiment of the present application, the eye movement video data is determined according to the scene timestamp corresponding to the scene key video frame data, comprising: obtaining initial eye movement video data of a target driver in a multi-traffic running scene according to scene video data; determining the eye movement video data based on the initial eye movement video data and the scene timestamp.
[0059] As a possible implementation manner, the target driver can first obtain initial eye movement video data through scene video data in a multi-traffic running scene, and then determine the eye movement video data of the target driver by using the scene timestamp.
[0060] It can be understood that each frame of video scene in the scene video data of the embodiment of the present application can be regarded as an image sequence, which has much richer content, stronger expression and larger information amount than an image, but there is usually a large amount of redundancy in each frame of video scene, and there is also the phenomenon of missing frames and redundancy in the extraction of video frames. Therefore, the embodiment of the present application can extract scene key video frame data in the scene video data through the key event, so as to effectively reduce the time spent in scene video data retrieval, and enhance the accuracy of scene video data retrieval.
[0061] Further, the embodiment of the present application finds the key frame of the initial eye movement video data according to the scene timestamp of the scene key video frame data, and then obtains the eye movement video data.
[0062] In step S102, a fixation heat map corresponding to each frame of eye movement data in the eye movement video data is obtained based on the eye movement video data.
[0063] It can be known from the above analysis that the embodiment of the present application can obtain a fixation heat map corresponding to each frame of eye movement data by using eye movement analysis software.
[0064] Optionally, in an embodiment of the present application, based on the eye movement video data, the gaze heat map corresponding to each frame of eye movement data in the eye movement video data is obtained, comprising: extracting the pupil center information of the target driver in each frame of eye movement data; obtaining the line of sight estimation information of the target driver in each frame of eye movement data according to the pupil center information; and obtaining the gaze heat map based on the line of sight estimation information.
[0065] As a possible implementation manner, the content of the gaze heat map generated by the eye movement analysis software in the embodiment of the present application can include two parts of pupil center detection and line of sight estimation.
[0066] In some embodiments, the pupil center detection can be performed based on image features in the embodiment of the present application, and then the pupil center information of the target driver in each frame of eye movement data is obtained.
[0067] For example, the embodiment of the present application first selects the eye area of the target driver in each frame of eye movement data, then detects the corneal reflection point and the pupil center of the target driver, and then obtains the pupil center information of the target driver.
[0068] It should be noted that the detection of the pupil center is affected by the presence of the corneal reflection point in the embodiment of the present application. In order to exclude the influence of the corneal reflection point on the calculation of the pupil center, the embodiment of the present application first performs threshold processing on each frame of eye movement data, then extracts the contour information of the pupil, then removes the pupil contour information within the light spot range, and finally performs ellipse fitting on the remaining contour. The center of the ellipse is the pupil center. Further, the embodiment of the present application extracts the position of the eyeball center in the camera coordinate system from the human eye image sequence using a human eye model, and performs Kalman filtering and other processing on the position result to obtain higher accuracy and robustness.
[0069] In some embodiments, the line of sight estimation can be performed based on machine learning in the embodiment of the present application, and then the line of sight estimation information of the target driver in each frame of eye movement data is obtained.
[0070] It can be understood that there are some differences in the physiological structure of each target driver's eye, and most of the current eye movement tracking systems need the target driver to concentrate on looking at the calibration point in the scene video data for calibration, so as to eliminate the inherent physiological deviation of the visual axis and the optical axis of the human eye, and obtain the real gaze point position. That is, the embodiment of the present application can use a human eye dataset to establish a human eye model with strong robustness through a machine learning method to perform line of sight estimation.
[0071] In step S103, based on the gaze heat map, the semantic segmentation information and three-dimensional reconstruction information of the gaze heat map are obtained from the pre-constructed LSD-SLAM, so as to construct the eye movement dataset of the target driver based on the semantic segmentation information and the three-dimensional reconstruction information.
[0072] It can be understood that the embodiments of the present application show through psychological research that when the target driver faces the need to pay attention to multiple visual clues at the same time, the order in which the target driver looks at these clues is very subjective, so the eye movement video data will encounter single focus and false positive gaze problems.
[0073] Further, in the embodiments of the present application, the single focus can be understood as being able to record only one position that the target driver is observing at each moment, while the target driver can pay attention to multiple important objects in the scene, and these objects are hidden, that is, a person's eyes are staring at one object while paying attention to another object; the false positive gaze can be understood as the target driver performing eye movement on the driving-irrelevant area such as the sky, trees and buildings, thereby generating irrelevant eye movement video data.
[0074] Further, in order to solve the above problems, the embodiments of the present application require more than 5 target drivers (which can also be other numerical values, and the present application does not make specific limitations) to participate in the collection of related data in the same scene video data. The heat map and gaze trajectory formed by the eye movement video data of these target drivers are aggregated and smoothed to make the gaze heat map of each frame of eye movement video data. The embodiments of the present application can record multiple important visual clues in a frame of video data by gathering the target driver's gaze; can effectively remove the noise sources such as the sky, trees and buildings by averaging the eye movement of the target driver and combining the pre-constructed LSD-SLAM to remove irrelevant areas; and realize multi-focus and false positive gaze filtering through eye movement aggregation smoothing and scene detection segmentation.
[0075] In the actual execution process, the embodiments of the present application can obtain semantic segmentation information and three-dimensional reconstruction information of the gaze heat map based on the gaze heat map from the pre-constructed LSD-SLAM, and further construct the eye movement data set of the target driver.
[0076] Optionally, in an embodiment of the present application, based on the gaze hotspot map, the semantic segmentation information and the three-dimensional reconstruction information of the gaze hotspot map are obtained from the pre-constructed LSD-SLAM, so as to construct the eye movement data set of the target driver based on the semantic segmentation information and the three-dimensional reconstruction information, comprising: generating two-dimensional semantic segmentation information and two-dimensional semantic labels corresponding to each frame of eye movement data based on the SLAM in the pre-constructed LSD-SLAM; dividing the gaze hotspot map by using the two-dimensional semantic labels to obtain a first gaze hotspot map of the gaze hotspot map, first semantic segmentation information corresponding to the first gaze hotspot map, a second gaze hotspot map, and second semantic segmentation information corresponding to the second gaze hotspot map; obtaining an optimized first gaze hotspot map corresponding to the first gaze hotspot map based on the first gaze hotspot map, the first semantic segmentation information, the second gaze hotspot map, and the second semantic segmentation information; obtaining three-dimensional reconstruction information of the first gaze hotspot map based on the optimized first gaze hotspot map and the first semantic segmentation information in combination with the recursive Bayesian and conditional random field in the pre-constructed LSD-SLAM; and generating the eye movement data set based on the three-dimensional reconstruction information and the optimized first gaze hotspot map.
[0077] As understood by those skilled in the art, the steps of constructing the eye movement data set of the target driver by the pre-constructed LSD-SLAM in the embodiments of the present application are as shown in Figure 7 , and the main content is as follows:
[0078] Step S701: generating two-dimensional semantic segmentation information and two-dimensional semantic labels of each frame of eye movement data by using SLAM.
[0079] It can be understood that in the SLAM, the motion of the camera can be estimated, and the position changes of each object in each frame of eye movement data can also be predicted, thereby generating a large amount of new data to provide more optimization conditions for semantic tasks and save the cost of manual calibration. Therefore, the semantic mapping uses dense or semi-dense feature points, so that the two-dimensional semantic segmentation information and the two-dimensional semantic labels (especially the camera pose) obtained by the SLAM can improve the performance of semantic segmentation. The flowchart of the semantic mapping is as shown in Figure 8 .
[0080] Step S702: dividing the gaze hotspot map by using the two-dimensional semantic labels to obtain a first gaze hotspot map of the gaze hotspot map, first semantic segmentation information corresponding to the first gaze hotspot map, a second gaze hotspot map, and second semantic segmentation information corresponding to the second gaze hotspot map.
[0081] In the embodiments of the present application, the two-dimensional semantic segmentation information of each frame of eye movement data, i.e., the pixels with two-dimensional semantic labels, can be mapped into three-dimensional point clouds, and then the first gaze hotspot map, the first semantic segmentation information corresponding to the first gaze hotspot map, the second gaze hotspot map, and the second semantic segmentation information corresponding to the second gaze hotspot map are obtained.
[0082] In addition, it should be noted that in the embodiments of the present application, the first gaze hotspot map can be understood as a gaze hotspot map that does not contain targets such as sky, trees and buildings; and the second gaze hotspot map can be understood as a gaze hotspot map that contains targets such as sky, trees and buildings.
[0083] Step S703: obtaining an optimized first gaze hotspot map based on the first gaze hotspot map, the first semantic segmentation information, the second gaze hotspot map and the second semantic segmentation information.
[0084] In the embodiments of the present application, the first gaze hotspot map is subjected to semantic segmentation, the second gaze hotspot map is subjected to depth estimation optimization of the first gaze hotspot map based on a small-baseline stereo vision comparison method, and then an optimized first gaze hotspot map is obtained.
[0085] Step S704: obtaining three-dimensional reconstruction information by using recursive Bayesian and conditional random field based on the optimized first gaze hotspot map and the first semantic segmentation information.
[0086] In the embodiments of the present application, recursive Bayesian can be used to enhance semantic segmentation, and conditional random field can be used to optimize three-dimensional reconstruction information, so as to realize semantic segmentation optimization and three-dimensional reconstruction.
[0087] It can be understood that in the embodiments of the present application, a semantic fusion recursive Bayesian method can be used, the semantic classification probability of a pixel in a current frame is multiplied by the classification probability at the old position in the previous frame as the final probability according to the estimation of the pixel motion by SLAM, that is, the probability of the pixel is multiplied along each frame, so as to enhance the result of semantic segmentation.
[0088] Step S705: generating eye movement data set based on the three-dimensional reconstruction information and the optimized first gaze hotspot map.
[0089] According to the method for establishing an eye movement data set based on a multi-traffic running scene provided in the embodiments of the present application, eye movement video data can be generated according to scene video data of the multi-traffic running scene, and then, by using a fixation heat map corresponding to each frame of eye movement data, semantic segmentation information and three-dimensional reconstruction information obtained from a pre-constructed LSD-SLAM, an eye movement data set of a target driver is constructed, which not only covers conventional situations such as robot lane-following and robot vehicle-tracking, but also extends to multi-traffic running scenes such as intersections and motorway invasion by cyclists, and performs traffic scene generalization and data classification, and through smoothing of the eye movement video data of the target driver and removal of sky noise by the SLAM, key events are labeled and annotated, and an eye movement data set of the target driver in the multi-traffic running scene is efficiently and accurately established. In this way, the problems in the related art that the data system combining eye movement and running scenes lacks systematic data collection and analysis, the data is poor, lacks diverse traffic running scenes, and the accuracy of the annotated data is low are solved.
[0090] Secondly, the device for establishing an eye movement data set based on a multi-traffic running scene provided in the embodiments of the present application is described with reference to the accompanying drawings.
[0091] Figure 9 A block schematic diagram of the device for establishing an eye movement data set based on a multi-traffic running scene provided in the embodiments of the present application is shown.
[0092] As shown in Figure 9 the device 90 for establishing an eye movement data set based on a multi-traffic running scene includes a generation module 100, an acquisition module 200 and a construction module 300.
[0093] The generation module 100 is configured to generate eye movement video data of a target driver in a multi-traffic running scene according to scene video data of the multi-traffic running scene.
[0094] The acquisition module 200 is configured to acquire a fixation heat map corresponding to each frame of eye movement data in the eye movement video data based on the eye movement video data.
[0095] The construction module 300 is configured to obtain semantic segmentation information and three-dimensional reconstruction information of the fixation heat map based on the fixation heat map by using a pre-constructed LSD-SLAM, and construct an eye movement data set of the target driver based on the semantic segmentation information and the three-dimensional reconstruction information.
[0096] Optionally, in an embodiment of the present application, the acquisition module 200 includes a first extraction unit, an acquisition unit and a first generation unit.
[0097] The first extraction unit is configured to extract pupil center information of the target driver in each frame of eye movement data in each frame of eye movement data.
[0098] The acquisition unit is configured to acquire the line-of-sight estimation information of the target driver in each frame of eye movement data according to the pupil center information.
[0099] The first generation unit is configured to obtain a gaze hotspot map based on the line-of-sight estimation information.
[0100] Optionally, in an embodiment of the present application, the construction module 300 comprises a second generation unit, a division unit, a third generation unit, a fourth generation unit and a fifth generation unit.
[0101] The second generation unit is configured to generate two-dimensional semantic segmentation information and a two-dimensional semantic label corresponding to each frame of eye movement data based on the SLAM in the pre-constructed LSD-SLAM.
[0102] The division unit is configured to divide the gaze hotspot map by using the two-dimensional semantic label to obtain a first gaze hotspot map of the gaze hotspot map, first semantic segmentation information corresponding to the first gaze hotspot map, a second gaze hotspot map and second semantic segmentation information corresponding to the second gaze hotspot map.
[0103] The third generation unit is configured to obtain an optimized first gaze hotspot map corresponding to the first gaze hotspot map based on the first gaze hotspot map, the first semantic segmentation information, the second gaze hotspot map and the second semantic segmentation information.
[0104] The fourth generation unit is configured to obtain three-dimensional reconstruction information of the first gaze hotspot map based on the optimized first gaze hotspot map and the first semantic segmentation information in combination with the recursive Bayesian and conditional random field in the pre-constructed LSD-SLAM.
[0105] The fifth generation unit is configured to generate an eye movement data set based on the three-dimensional reconstruction information and the optimized first gaze hotspot map.
[0106] Optionally, in an embodiment of the present application, the generation module 100 comprises a first determination unit, a second extraction unit and a second determination unit.
[0107] The first determination unit is configured to determine a key event in a multi-traffic running scene.
[0108] The second extraction unit is configured to extract scene key video frame data in scene video data based on the key event.
[0109] The second determination unit is configured to determine eye movement video data according to a scene timestamp corresponding to the scene key video frame data.
[0110] Optionally, in an embodiment of the present application, the second determination unit comprises an acquisition subunit and a determination subunit.
[0111] The acquisition subunit is configured to acquire initial eye movement video data of the target driver in the multi-traffic running scene according to the scene video data.
[0112] The determination subunit is configured to determine the eye movement video data based on the initial eye movement video data and the scene timestamp.
[0113] It should be noted that the foregoing explanation and description of the method for establishing the eye movement data set based on the multi-traffic running scene also applies to the device for establishing the eye movement data set based on the multi-traffic running scene, which will not be described here again.
[0114] The device for establishing the eye movement data set based on the multi-traffic running scene according to the embodiments of the present application can generate the eye movement video data according to the scene video data of the multi-traffic running scene, and then utilize the gaze heat map corresponding to each frame of eye movement data to obtain the semantic segmentation information and the three-dimensional reconstruction information from the pre-constructed LSD-SLAM, so as to construct the eye movement data set of the target driver. In addition to covering the conventional cases of robot lane-following and robot vehicle-tracking, the device also expands the multi-traffic running scene such as intersection and bicycle invasion of motor lane, and performs traffic scene generalization and data classification. Through smoothing of the eye movement video data of the target driver, the device removes the sky noise of SLAM, labels and marks the key events, and efficiently and accurately establishes the eye movement data set of the target driver in the multi-traffic running scene. Thus, the device solves the problems in the related art that the data system of eye movement combined with the running scene lacks systematic data collection and analysis, the data is poor, lacks diverse traffic running scenes, and the accuracy of the marked data is low.
[0115] Figure 10 A structural schematic diagram of an electronic device according to an embodiment of the present application is provided. The electronic device can include:
[0116] The memory 1001, the processor 1002, and the computer program stored in the memory 1001 and executable on the processor 1002.
[0117] The processor 1002 implements the method for establishing the eye movement data set based on the multi-traffic running scene provided in the above embodiments when executing the program.
[0118] Further, the electronic device further includes:
[0119] The communication interface 1003 is configured to communicate between the memory 1001 and the processor 1002.
[0120] The memory 1001 is configured to store the computer program executable on the processor 1002.
[0121] The memory 1001 can include a high-speed RAM memory and can also include a non-volatile memory, such as at least one disk memory.
[0122] If the memory 1001, the processor 1002 and the communication interface 1003 are implemented independently, the communication interface 1003, the memory 1001 and the processor 1002 can be connected to each other through a bus and complete communication between each other. The bus can be an Industry Standard Architecture (ISA) bus, a Peripheral Component Interconnect (PCI) bus or an Extended Industry Standard Architecture (EISA) bus, etc. The bus can be divided into an address bus, a data bus, a control bus, etc. For convenience of representation, Figure 10 In the figure, only one thick line is used to represent the bus, but it does not mean that there is only one bus or only one type of bus.
[0123] Optionally, in a specific implementation, if the memory 1001, the processor 1002 and the communication interface 1003 are integrated on a chip, the memory 1001, the processor 1002 and the communication interface 1003 can complete communication between each other through an internal interface.
[0124] The processor 1002 can be a Central Processing Unit (CPU), or an Application Specific Integrated Circuit (ASIC), or one or more integrated circuits configured to implement one or more embodiments of the present application.
[0125] The embodiment of the present application also provides a computer readable storage medium, which stores a computer program, and the program is executed by a processor to implement the method for establishing an eye movement data set based on a multi-traffic running scenario as described above.
[0126] The embodiment of the present application also provides a computer program product, which includes a computer program, and the program is executed to implement the method for establishing an eye movement data set based on a multi-traffic running scenario as described above.
[0127] In the description of the application, reference to "one embodiment", "some embodiments", "an example", "a specific example", or "some examples" means that a particular feature, structure, material, or characteristic being described is included in at least one embodiment or example of the application. The appearances of the phrase in various places in the specification are not necessarily all referring to the same embodiment or example. Furthermore, the described specific features, structures, materials, or characteristics can be combined in any suitable manner in one or more embodiments or examples. In addition, the usage of "N" means at least two, for example, two, three or the like, unless explicitly stated otherwise.
[0128] Furthermore, the terms "first", "second", or the like, are used merely as a designation of certain elements or features of the application, and do not imply or connote relative importance or a specific order of precedence. Thus, features defined with "first", "second", etc. can include at least one of the features, either explicitly or implicitly.
[0129] Any process or method descriptions or blocks in flow charts or otherwise described herein represent embodiments of modules, segments, or portions of code which include one or more executable instructions for implementing specific logic functions or steps, and alternate implementations are possible. In some embodiments, the processes or methods described in flow charts or otherwise described herein are not necessarily performed in the order shown or discussed, including, for example, performing or depending from other operations or stages, in parallel, in reverse order, or in some other suitable manner. Blocks can also be omitted.
[0130] The logic and / or steps represented in the flowcharts and / or described herein, for example, can be considered as a sequence of executable instructions stored in a computer readable medium, which can be executed by an instruction execution system, apparatus or device, such as a computer-based system, a processor-based system, or other system that can fetch the instructions from the instruction execution system, apparatus or device and execute the instructions, or a combination of the above. For the purposes of this specification, a "computer readable medium" can be any apparatus that can contain, store, communicate, propagate, or transport the program for use by or in connection with the instruction execution system, apparatus or device. The computer readable medium can be a computer readable storage medium or a computer readable signal medium. The computer readable storage medium can include, but is not limited to, an electronic, magnetic, optical, electromagnetic, infrared, or semiconductor system, apparatus, or device, or a propagation medium. The computer readable signal medium can include, but is not limited to, a computer readable medium that facilitates transfer of the program from one place to another. A specific example of a computer readable medium is a non-transitory computer-readable storage medium. A specific example of a computer readable signal medium is a source or destination of the computer readable medium. Another specific example of a computer readable signal medium is a computer readable signal travelling through space. Thus, a computer readable medium can take many forms of hardware to carry out the program for use by or in connection with the instruction execution system, apparatus or device.
[0131] It should be understood that aspects of the application can be implemented in hardware, software, firmware or combinations thereof. In the above embodiments, the N steps or methods can be implemented in software or firmware stored in a memory and executed by a suitable instruction execution system. If implemented in hardware and in another embodiment, the hardware can be implemented using any or a combination of the following technologies, which are each well known in the art: a discrete logic circuit(s) having logic gates for implementing logic functions upon an application of data signals, an application specific integrated circuit having appropriate combinational logic gates, a programmable gate array(s) (PGA), a field programmable gate array (FPGA), etc.
[0132] Those of skill in the art would understand that the steps of the methods carried out above can be carried out wholly or partly by a program instructing relevant hardware, and the program can be stored in a computer readable storage medium, and when executed, includes one or a combination of the steps of the method embodiments.
[0133] In addition, each of the functional units in the various embodiments of the present application can be integrated in one processing module, or each of the units can be physically present separately, or two or more units can be integrated in one module. The integrated module can be implemented in the form of hardware or in the form of a software functional module. When the integrated module is implemented in the form of a software functional module and sold or used as an independent product, it can also be stored in a computer readable storage medium.
[0134] The storage medium mentioned above can be a read-only memory, a magnetic disk or an optical disk, etc. Although the embodiments of the present application have been shown and described above, it should be understood that the above embodiments are exemplary and should not be construed as limiting the present application, and those skilled in the art can make changes, modifications, replacements and variations to the above embodiments within the scope of the present application.
Claims
1. A method for establishing an eye movement dataset based on a multi-traffic scenario, characterized in that, The method comprises the following steps: generating eye movement video data of a target driver in a multi-traffic operation scene according to scene video data of the multi-traffic operation scene; obtaining a gaze hotspot map corresponding to each frame of eye movement data in the eye movement video data based on the eye movement video data; obtaining semantic segmentation information and three-dimensional reconstruction information of the gaze hotspot map based on the gaze hotspot map by using a large-scale direct method simultaneous localization and mapping (LSD-SLAM) pre-constructed, and constructing an eye movement data set of the target driver based on the semantic segmentation information and the three-dimensional reconstruction information.
2. The method of claim 1, wherein, The method comprises the following steps: extracting pupil center information of the target driver in each frame of eye movement data in the eye movement data; obtaining line-of-sight estimation information of the target driver in each frame of eye movement data according to the pupil center information; obtaining the gaze hotspot map based on the line-of-sight estimation information.
3. The method of claim 1, wherein, The method comprises the following steps: generating two-dimensional semantic segmentation information and a two-dimensional semantic label corresponding to each frame of eye movement data based on SLAM in the pre-constructed LSD-SLAM; dividing the gaze hotspot map by using the two-dimensional semantic label to obtain a first gaze hotspot map, first semantic segmentation information corresponding to the first gaze hotspot map, a second gaze hotspot map, and second semantic segmentation information corresponding to the second gaze hotspot map; obtaining an optimized first gaze hotspot map corresponding to the first gaze hotspot map based on the first gaze hotspot map, the first semantic segmentation information, the second gaze hotspot map, and the second semantic segmentation information; obtaining three-dimensional reconstruction information of the first gaze hotspot map based on the optimized first gaze hotspot map and the first semantic segmentation information, and combining recursive Bayesian and conditional random fields in the pre-constructed LSD-SLAM; generating the eye movement data set based on the three-dimensional reconstruction information and the optimized first gaze hotspot map.
4. The method of claim 1, wherein, The method comprises the following steps: determining a key event in the multi-traffic operation scene; extracting scene key video frame data in the scene video data based on the key event; determining the eye movement video data according to a scene timestamp corresponding to the scene key video frame data.
5. The method of claim 4, wherein, The method comprises the following steps: obtaining initial eye movement video data of the target driver in the multi-traffic operation scene according to the scene video data; determining the eye movement video data based on the initial eye movement video data and the scene timestamp.
6. An apparatus for establishing an eye movement dataset based on a multi-traffic operating scenario, the apparatus comprising: The method comprises the following steps: The generating module is configured to generate eye movement video data of a target driver in a multi-traffic operation scene according to scene video data of the multi-traffic operation scene; The acquiring module is configured to acquire a gaze hotspot map corresponding to each frame of eye movement data in the eye movement video data based on the eye movement video data; The constructing module is configured to obtain semantic segmentation information and three-dimensional reconstruction information of the gaze hotspot map from a pre-constructed LSD-SLAM based on the gaze hotspot map, and to construct an eye movement data set of the target driver based on the semantic segmentation information and the three-dimensional reconstruction information.
7. The apparatus of claim 6, wherein, The acquiring module comprises: A first extracting unit configured to extract pupil center information of the target driver in each frame of eye movement data in the each frame of eye movement data; An acquiring unit configured to acquire line-of-sight estimation information of the target driver in the each frame of eye movement data according to the pupil center information; A first generating unit configured to obtain the gaze hotspot map based on the line-of-sight estimation information.
8. The apparatus of claim 6, wherein, The constructing module comprises: A second generating unit configured to generate two-dimensional semantic segmentation information and a two-dimensional semantic label corresponding to each frame of eye movement data based on SLAM in the pre-constructed LSD-SLAM; A dividing unit configured to divide the gaze hotspot map by using the two-dimensional semantic label to obtain a first gaze hotspot map, first semantic segmentation information corresponding to the first gaze hotspot map, a second gaze hotspot map, and second semantic segmentation information corresponding to the second gaze hotspot map; A third generating unit configured to obtain an optimized first gaze hotspot map corresponding to the first gaze hotspot map based on the first gaze hotspot map, the first semantic segmentation information, the second gaze hotspot map, and the second semantic segmentation information; A fourth generating unit configured to obtain three-dimensional reconstruction information of the first gaze hotspot map by combining recursive Bayesian and conditional random fields in the pre-constructed LSD-SLAM based on the optimized first gaze hotspot map and the first semantic segmentation information; A fifth generating unit configured to generate the eye movement data set based on the three-dimensional reconstruction information and the optimized first gaze hotspot map.
9. The apparatus of claim 6, wherein, The generating module comprises: A first determining unit configured to determine a key event in the multi-traffic operation scene; A second extracting unit configured to extract scene key video frame data in the scene video data based on the key event; A second determining unit configured to determine the eye movement video data according to a scene timestamp corresponding to the scene key video frame data.
10. The apparatus of claim 9, wherein, The second determining unit comprises: An acquiring sub-unit configured to acquire initial eye movement video data of the target driver in the multi-traffic operation scene according to the scene video data; A determining sub-unit configured to determine the eye movement video data based on the initial eye movement video data and the scene timestamp.
11. An electronic device, comprising: The computer program is stored in the memory and executable on the processor, and the processor executes the program to implement the method for establishing an eye movement data set based on a multi-traffic operation scene according to any one of claims 1-5. 12. A computer readable storage medium having stored thereon a computer program, characterized in that, The program is executed by a processor for implementing the method of establishing an eye movement dataset based on a multi-traffic running scenario according to any one of claims 1-5.
13. A computer program product, characterised in that, The computer program is executed for implementing the method of establishing an eye movement dataset based on a multi-traffic running scenario according to any one of claims 1-5.
Citation Information
Patent Citations
Driving behavior collaborative recognition method and device fusing eye movement and scene information
CN118898830A