Computer-implemented method for generating an area of attention
By generating a geometric attention region through a single image capture unit and data processing unit, the problem of complex gaze direction localization in existing technologies is solved, enabling rapid, flexible division and accurate localization of gaze direction in three-dimensional space.
Patent Information
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- EMOTION3D GMBH
- Filing Date
- 2022-04-26
- Publication Date
- 2026-05-05
AI Technical Summary
In existing technologies, analyzing a person's gaze direction requires the use of multiple image capture units and additional sensors, which makes the system complex and makes it difficult to accurately locate the gaze direction of different people in different positions, and makes it impossible to effectively divide the attention area in three-dimensional space.
By using a single image capture unit combined with a data processing unit and a database, the intersection points are calculated to generate a geometric attention region by taking multiple images and analyzing the human gaze direction vector, thus avoiding the use of additional sensors.
It enables the rapid and flexible generation of attention regions in three-dimensional space without the need for additional hardware installation, simplifying the system structure and improving the accuracy and flexibility of gaze direction.
Smart Images

Figure CN117121061B_ABST
Abstract
Description
Technical Field
[0001] This invention relates to a computer-implemented method for generating geometric attention regions in three-dimensional space. Background Technology
[0002] As is known in existing technology, image-based methods are used to determine a person's attention by analyzing their gaze direction. This can be achieved, for example, by extracting the head direction or eye position from a person's photograph or video. Attention levels can be calculated by classifying gaze directions into predetermined regions. This attention modeling is primarily used in the automotive field, but also in other human-machine applications. In vehicles, this attention modeling can be used to determine whether people are paying attention to traffic or are distracted. In stores, this attention modeling can be used to identify which products attract more or less attention. In robotics, a person's gaze direction can be used to control machines.
[0003] To classify a person's gaze direction, their visual field needs to be divided into attention regions. Since these attention regions must be valid in three-dimensional space, they are typically defined by polygons in three-dimensional space; that is, these attention regions comprise three or more points bounded by x, y, and z coordinates spanning the region. After dividing a person's visual field into such attention regions, each gaze direction can be assigned to a defined attention region.
[0004] However, one problem is that the image capture unit, which is expected to analyze the person's gaze direction, must be pointed at the person, and therefore cannot capture the person's field of vision simultaneously. In order to capture the person's field of vision as well, two or more image capture units are needed.
[0005] Extracting a person's gaze direction from an image captured by an image capture unit also presents a problem: the same gaze direction from different people being analyzed does not necessarily mean that they are looking at the same area in space. For example, the driver and passengers in a vehicle may have the same gaze direction, but due to their different seating positions, the areas they are focusing on are completely different.
[0006] Therefore, in existing technologies, additional sensors are frequently used. These sensors have an attentional region within the field of view and are connected to the image capture unit that captures the person, enabling the establishment of a direct geometric relationship between the person's gaze direction and the attentional region. For example, these additional sensors are mounted next to or behind the person, allowing the attentional region to be defined for each individual. Summary of the Invention
[0007] One object of the present invention is to improve upon these methods known in the prior art, and in particular, to enable the creation of geometric attention regions without the use of additional sensors.
[0008] These and other objects of the present invention are achieved by computer-implemented methods and apparatus according to the independent claims.
[0009] The computer-implemented method according to the invention for generating geometric attention regions for at least one person in three-dimensional space uses a single image capture unit, a data processing unit, a database, and a display unit arranged in space.
[0010] The three-dimensional space can be the interior of a vehicle; however, the method of the present invention is not limited to vehicles.
[0011] The image capture unit, data processing unit, and database can preferably be entirely housed within the vehicle. However, it is also possible to specify that the data processing unit and database are housed within the vehicle and communicate with an external server (e.g., an internet server) via an interface (e.g., a wireless connection), where a database containing previously stored and / or continuously supplemented reference data may exist for processing the captured data.
[0012] The image capture unit can be a photo camera or video camera designed to capture two-dimensional or three-dimensional photographs or videos. In particular, the image capture unit can be a TOF (Time-of-Flight) camera, etc. The use of a TOF camera helps to robustly detect people and extract gaze direction. However, conventional 2D cameras combined with previously stored 3D models of people or people's faces can also robustly extract gaze direction from individual 2D images. The image capture unit can preferably be arranged in such a way that, if possible, the entire interior, at least all passenger seats, but at least the driver's seat and passenger seats are visible in the image area.
[0013] The data processing unit can be designed as a microcontroller or microcomputer and includes a central processing unit (CPU), volatile semiconductor memory (RAM), non-volatile semiconductor memory (ROM, SSD), magnetic storage (hard disk) and / or optical storage (CD-ROM), and interface units (Ethernet, USB), etc. The database can be provided as a software module within the data processing unit, in a separate computer, or on an external server. The components of such a data processing unit are generally known to those skilled in the art.
[0014] The method according to the present invention includes at least the following steps:
[0015] Initially, the display unit requires the person to focus on the apex of a specific area of attention through head and / or eye movement. This area of attention can be, for example, a polygon corresponding to the position of a vehicle's windshield, dashboard, side window, or side mirror. However, it can also be a polygon defining a shop's display window in a shopping street or a display panel of a technical device in a manufacturing plant.
[0016] Then, the image capture unit captures a series of N consecutive first photographs of the space contained within a predetermined time period and transmits these photographs to the data processing unit. The number N represents the number of vertices of the polygon defining the region of attention. If the shape and extent of the region of attention are predetermined, the number N can be equal to 1. However, the number N can also be greater than or equal to 3. Preferably, in the case of a rectangular or trapezoidal region of attention, N equals 4. The image capture unit can also capture short videos of the space instead of capturing N consecutive photographs. The data processing unit analyzes the first photograph or video, detects people within it, and determines a first gaze direction vector BV1-BV for the number of people in the vector. n N first gaze direction vectors BV1-BV n The attention is guided to the vertices of the relevant area of focus and is defined, for example, as vectors in three-dimensional Cartesian or polar coordinates.
[0017] Determining the gaze direction vector of at least one person from a photograph or video can be performed by extracting and analyzing the person's head orientation and / or eye position. This method of determining a person's gaze direction vector from a photograph or video is known in principle; known image processing algorithms can be used for this purpose. In particular, image analysis libraries in databases and / or detectors (e.g., neural networks) trained with training examples can be referenced.
[0018] In addition, the data processing unit can extract the gaze direction vectors of several people (e.g., driver and passenger) from a single photo series or a single video.
[0019] As a result, the display unit prompts the person to change their position in space. For example, the person can move their seat backward, forward, up, or down, or tilt it forward or backward. Afterward, the display unit prompts the person to refocus on the apex of the attention area by moving their head and / or eyes.
[0020] In the next step, the image capture unit takes a second set of photos of the space accommodating the person, numbered N, and transmits these photos to the data processing unit. Similarly, video can be recorded for a short period instead of taking photos. The number N can be equal to or greater than 1, thus allowing the system to specify that the person has changed position after each time they gaze at a vertex and are required to gaze at another vertex.
[0021] The data processing unit then extracts the second three-dimensional gaze direction vector BV1′-BV for the number of people in N. N Due to the different positions of the person, the second gaze direction vector is BV1′-BV. N ′ and the first gaze direction vector BV1-BV N different.
[0022] Next, the data processing unit will view the gaze direction vector BV1-BV N Superimposed on the gaze direction vector BV1′-BV N On top, N three-dimensional intersection points P1-P are determined by triangulation. N These N intersection points correspond to the vertices of the region of attention in the common three-dimensional coordinate system of space. The N intersection points P1-P N The vertices that serve as the attention region are stored in the database.
[0023] This method can also be performed at two or more different locations of the person to increase the accuracy of triangulation.
[0024] For example, it can be stipulated that the person is first required to move their seat downwards, then upwards, then forwards, and finally backwards. After each change in position, N gaze direction vectors are determined, and these N gaze direction vectors are then used to calculate the precise intersection point P1-P. N .
[0025] The advantage of this invention is that it eliminates the need for additional sensors or image capture units to create the configuration of the attention region. In any case, using an image capture unit pointed at the person being analyzed is sufficient. There is no need to photograph the person from behind.
[0026] This allows for rapid and flexible configuration of the attention area, ensuring that the entire system can be converted or installed into existing image capture and image analysis systems more quickly. No new hardware needs to be installed.
[0027] According to the present invention, the method can be repeated to create multiple attentional regions for a single person. For example, it can be specified that after performing the method according to the invention, the person is now asked to gaze at the apex of another attentional region.
[0028] According to the present invention, the method can be repeated to create multiple attention regions for different individuals. For example, it can be specified that after performing the method according to the present invention, all detected individuals are now required to gaze at the apex of another attention region.
[0029] According to the present invention, after generating and storing one or more attention regions of at least one person, the image capture unit continuously detects the person's gaze direction vector and transmits it to the data processing unit, and the data processing unit continuously checks whether the detected gaze direction vector falls into one of the stored attention regions. This allows for monitoring of a person's attention during operation.
[0030] According to the present invention, the data processing unit can be configured to calculate the person's level of attention by determining the frequency at which the gaze direction vector of at least one person detected per unit time falls into a region of a stored attention region. It can be configured that the display unit issues a warning message when the level of attention drops below a predetermined threshold. For example, a threshold can be defined according to which the person looks at the attention region defined by the windshield at least once per second.
[0031] The present invention also relates to a computer-readable storage medium including instructions for causing a data processing unit to perform the method according to the present invention.
[0032] The present invention also relates to an apparatus for generating a geometric attention region for at least one person in three-dimensional space, comprising a single image capture unit, a data processing unit, a database, and a display unit, wherein the apparatus is adapted to perform the method according to the invention.
[0033] According to the present invention, the space can be specifically defined as the interior of the vehicle, and the image capturing unit is arranged in front of the driver's seat or the passenger seat along the vehicle's driving direction, and preferably in the center position above the vehicle's windshield or above the dashboard.
[0034] Other features of the invention are derived from the claims, exemplary embodiments, and drawings.
[0035] Brief description of the attached diagram
[0036] The following sections provide a detailed explanation of non-exclusive exemplary embodiments of the present invention.
[0037] Figure 1 A schematic diagram of a vehicle having a data processing unit for performing the method according to the invention is shown;
[0038] Figure 2 Illustrative examples of various attention areas in a vehicle according to the invention are shown;
[0039] Figures 3a-3d A schematic diagram of the interior of a vehicle during the execution of the method according to the invention is shown;
[0040] Figure 3e A schematic diagram is shown to determine the intersection of two gaze direction vectors of a person. Detailed Implementation
[0041] Figure 1A schematic diagram of a vehicle 8 is shown, in which an electronic data processing unit 4 is integrated for performing the method according to the invention. Inside the interior 2 of the vehicle 8, an image capture unit 3 in the form of a camera is arranged on the ceiling. The camera is configured and arranged such that it can capture the interior 2 of the vehicle 8, i.e., at least two front seats and the people seated therein. Several people sit in two rows inside. In this embodiment, the vehicle is configured as a passenger car. The data processing unit 4 is connected to a database 5 and has an interface (not shown) for communicating with external electronic components. Furthermore, a display unit 6 is provided, which is also connected to the data processing unit 4. The display unit 6 is arranged in the dashboard of the vehicle 8.
[0042] Figure 2 A schematic example of various attention areas 1, 1′, 1”, 1”′ in the interior 2 of a vehicle 8 according to the present invention is shown.
[0043] In this example, the windshield forms the first attention area 1, the two side mirrors form the second attention area 1', the two side windows form the third attention area 1'", and the dashboard forms the fourth attention area 1"'. Furthermore, the image capture unit 3 is shown as a video camera positioned above the center of the windshield. The display unit 6 is shown as a centrally located touchscreen.
[0044] Figures 3a-3d A schematic diagram of the interior of a vehicle is shown when the method according to the invention is performed.
[0045] Figure 3a Photograph 7 shows an interior 2 with three people, where the gaze directions of the driver and passenger have been extracted as gaze direction vectors BV1. These individuals were instructed to focus their attention on vertex P1 of attention region 1. The extraction of the gaze direction vectors for the two individuals can be performed substantially simultaneously or separately.
[0046] Figure 3b Another photograph 7 shows the interior 2 behind another vertex P2 of attention region 1. Again, the gaze direction vectors of the driver and passenger are extracted as gaze direction vector BV2. Here, extraction can be performed simultaneously or separately.
[0047] Figure 3c Another photograph 7 shows the interior 2 after people are asked to change their position and refocus their attention on the vertex P1 of attention area 1. Again, the gaze direction vectors of the driver and passenger are extracted as gaze direction vector BV1'.
[0048] Figure 3d Another photograph 7 shows the interior 2 after people are asked to refocus their attention on the vertex P2 of attention area 1. Again, the gaze direction vectors of the driver and passenger are extracted as gaze direction vector BV2'.
[0049] Repeat these steps for all vertices of the desired polygonal region of attention until sufficient (but at least two distinct) gaze direction vectors have been extracted for each person in each region of attention to perform triangulation to determine point P1-P2. N The location.
[0050] Figure 3e A schematic diagram of triangulation for determining the intersection of two gaze direction vectors of a person is shown. Two different sitting postures are proposed, with gaze direction vectors BV1 and BV2 corresponding to the first sitting posture, and gaze direction vectors BV1′ and BV2′ corresponding to the second sitting posture.
[0051] By superimposing the gaze direction vectors BV1 and BV1' or BV2 and BV2', the intersection points P1 and P2 can be determined. N In this case, the intersection points P1 and P N Pay attention to the vertices of region 1.
[0052] The present invention is not limited to the exemplary embodiments described, but also includes other embodiments of the invention within the scope of the following patent claims.
[0053] List of reference numerals
[0054] 1, 1', 1”, 1”' Note the area
[0055] 2. Three-dimensional space
[0056] 3 Image capture unit
[0057] 4 Data Processing Unit
[0058] 5. Database
[0059] 6 display units
[0060] 7 photos
[0061] 8 vehicles
Claims
1. A computer-implemented method for generating multiple geometric attention regions (1) for at least one person in a three-dimensional space (2), the method using a single image capture unit (3), a data processing unit (4), a database (5), and a display unit (6), comprising the following steps: a. The display unit (6) outputs a request to the person to gaze at at least one point in the space (2) corresponding to a vertex of the attention region (1) to be created. b. The image capture unit (3) captures a first photograph (7) of N number accommodating the space (2) of the person, and the data processing unit (4) extracts a first three-dimensional gaze direction vector BV1-BV of N number accommodating the person. N Where N is greater than or equal to 3, c. The display unit (6) outputs a request to the person to change their position in the space (2) and refocus on the vertex of the attention area (1). d. The image capture unit (3) captures a second photograph (7′) of N number accommodating the space (2) of the person, and the data processing unit (4) extracts a second three-dimensional gaze direction vector BV1′-BV of N number accommodating the person. N ′, e. The data processing unit (4) processes the gaze direction vector BV1-BV N The line and the gaze direction vector BV1′-BV N The lines are superimposed to determine N three-dimensional intersection points P1-P'. N , f. The space (2) containing vertices P1-P N The polygons are stored as attention regions (1) in the database (5).
2. The method according to claim 1, characterized in that, The number of vertices, N, is 4.
3. The method according to claim 1 or 2, characterized in that, The space (2) is the interior of the vehicle (8).
4. The method according to claim 1, characterized in that, Repeat the above method to create multiple attentional regions (1, 1′, 1'') for a single person, N=4.
5. The method according to claim 1 or 2, characterized in that, Repeat the above method to create multiple attentional regions (1, 1′, 1'') for different individuals.
6. The method according to claim 1 or 2, characterized in that, a. After generating and storing one or more attention regions (1, 1′, 1'') of at least one person, the image capture unit (3) continuously detects the person's gaze direction vector and transmits the gaze direction vector to the data processing unit (4), and b. The data processing unit (4) checks whether the detected gaze direction vector falls into one of the stored attention regions (1, 1′, 1'').
7. The method according to claim 6, characterized in that, a. The data processing unit (4) calculates the attention level of at least one person by determining the frequency at which the detected gaze direction vector falls into one of the stored attention regions (1, 1′, 1'') per unit time, and b. When the level of attention drops below a predetermined threshold, the display unit (6) outputs a warning message.
8. A computer-readable storage medium comprising a data processing unit (4) for performing the method according to any one of claims 1 to 7.
9. An apparatus for generating multiple geometric attention regions (1) for at least one person in a three-dimensional space (2), comprising a single image capture unit (3), a data processing unit (4), a database (5), and a display unit (6), characterized in that, The apparatus is designed to perform the method according to any one of claims 1 to 7.
10. The apparatus according to claim 9, characterized in that, The space (2) is the interior of the vehicle (8), and the image capture unit (3) is arranged in front of the driver's seat or the passenger seat along the vehicle's driving direction.
11. The apparatus according to claim 10, characterized in that, The image capture unit (3) is arranged in the center above the windshield of the vehicle.
Citation Information
Patent Citations
Methods and systems for processing attention data from a vehicle
CN104851242A
Display control apparatus and method
CN107278187A