COMPUTER-IMPLEMENTED METHOD FOR CREATING AN ATTENTION ZONE
Patent Information
- Application Number
- DE502022004927
- Authority / Receiving Office
- DE · DE
- Patent Type
- Patents
- Current Assignee / Owner
- Priority Date
- 2021-05-05
- Filing Date
- 2022-04-26
- Publication Date
- 2025-08-28
- Estimated Expiration
- 2042-04-26
AI Technical Summary
Existing gaze direction detection methods require additional sensors to define geometric attention zones in three-dimensional space, necessitating multiple image acquisition units and complicating the setup.
A computer-implemented method using a single image recording unit, data processing unit, and database to create geometric attention zones by triangulating gaze direction vectors from multiple positions, eliminating the need for additional sensors.
Enables quick and flexible configuration of attention zones without additional hardware, allowing continuous monitoring of gaze directions and attention levels with improved accuracy and reduced setup complexity.
Description
[0001] The invention relates to a computer-implemented method for creating a geometric attention zone in a three-dimensional space.
[0002] It is known from the state of the art to use image-based methods to determine a person's attention by analyzing their gaze direction. For example, head orientation or eye position can be extracted from photos or videos of the person. By classifying the gaze direction into predefined areas, a level of attention can be calculated. This type of attention modeling is used primarily in the automotive sector, but also in other human-machine applications. In a vehicle, it can be used to determine whether people are paying attention to the traffic or are distracted. In stores, it can be determined which products attract more or less attention. In robotics, a person's gaze direction can be used to control machines.
[0003] To classify a person's gaze direction, the person's field of vision must be divided into attention zones. Since these attention zones must be valid in three-dimensional space, they are usually defined using polygons in three-dimensional space. They comprise three or more defined points with x, y, and z coordinates that span a surface. After dividing the person's field of vision into such attention zones, each of the person's gaze directions can be assigned to a defined attention zone.
[0004] Methods for detecting a person's gaze direction are known from the state of the art. For example, the publications Pichitwong Widthipong et al., "An Eye-Tracker-Based 3D Point-of-Gaze Estimation Method Using Head Movement," IEEE ACCESS, Vol. 7, August 7, 2019, pages 99086-99098, and Mardanbegi Diako, "Resolving Target Ambiguity in 3D Gaze Interaction through VOR Depth Estimation", CCS '18: Proceedings of the 2018 ACM SIGSAG Conference on Computer and Communications Security, ACM Press, NY, USA, May 2, 2019, pages 1 - 12 such known methods for gaze direction detection.
[0005] One problem, however, is that the image acquisition unit intended to analyze the person's gaze direction must necessarily be directed at the person themselves, and thus cannot record the person's field of vision itself. If the person's field of vision is also to be recorded, two or more image acquisition units are required.
[0006] When extracting a person's gaze direction from the images captured by the image acquisition unit, the problem arises that identical gaze directions of different analyzed individuals do not necessarily mean that these individuals are looking at the same areas in space. For example, the driver and front passenger in a vehicle may have identical gaze directions, but focus on completely different areas due to their different seating positions.
[0007] Therefore, the state of the art regularly uses additional sensors that focus on the attention zones and are linked to the image capture unit that captures the person, thus establishing a direct geometric relationship between the person's gaze direction and the attention zones. These additional sensors are installed, for example, next to or behind the person, so that attention zones can be defined for each individual.
[0008] An object of the invention is to improve these methods known from the prior art and in particular to enable the creation of geometric attention zones without the use of additional sensors.
[0009] These and other objects of the invention are achieved by a computer-implemented method and a device according to the independent claims.
[0010] A computer-implemented method according to the invention for creating a plurality of geometric attention zones for at least one person in a three-dimensional space uses a single image recording unit arranged in the space, a data processing unit, a database and a display unit.
[0011] The three-dimensional space may be the interior of a vehicle; however, the method according to the invention is not limited to vehicles.
[0012] The image acquisition unit, data processing unit, and database can preferably be arranged entirely within a vehicle. However, it can also be provided that the data processing unit and the database are arranged in the vehicle and communicate via an interface, for example a wireless connection, with an external server, for example a server on the Internet, which may contain a database with pre-stored and / or continuously updated reference data that is used to process the recorded data.
[0013] The image recording unit can be a photo or video camera designed to capture two- or three-dimensional photos or videos. In particular, it can be a ToF (Time-of-Flight) camera or the like. The use of a ToF camera facilitates the robust detection of persons and the extraction of the viewing direction. However, a conventional 2D camera, in combination with a pre-stored 3D model of the person or the person's face, can also enable robust extraction of the viewing direction from individual 2D images. The image recording unit can preferably be arranged in a vehicle in such a way that the entire interior, in any case all passenger seats, but at least the driver and front passenger seats, are visible in the image area.
[0014] The data processing unit can be embodied as a microcontroller or microcomputer and comprise a central processing unit (CPU), a volatile semiconductor memory (RAM), a non-volatile semiconductor memory (ROM, SSD hard drive), a magnetic memory (hard drive) and / or an optical memory (CD-ROM), as well as interface units (Ethernet, USB), and the like. The database can be provided as a software module in the data processing unit, in a computer separate from the data processing unit, or in an external server. The components of such data processing units are generally known to those skilled in the art.
[0015] A method according to the invention comprises at least the following steps: Initially, the person is prompted by the display unit to fixate the corner points of a specific attention zone using head and / or eye movements. The attention zone can, for example, be polygons that correspond to the position of the windshield, dashboard, side windows, or side mirrors of a vehicle. However, it can also be polygons that define the shop window of a store on a high street or the display panel of a technical device in a production facility.
[0016] The image acquisition unit then takes N consecutive first photos of the room with the person over a predetermined period of time and transmits them to the data processing unit. The number N denotes the number of vertices of the polygon defining the attention zone. The number N is greater than or equal to three. Preferably, N is four in the case of rectangular or trapezoidal attention zones. Instead of taking N consecutive photos, the image acquisition unit can also record a short video of the room. The data processing unit analyzes the first photos or video, detects the person contained therein, and determines N first gaze direction vectors BV 1 - BV N for that person.The first N gaze direction vectors BV 1 - BV N are directed from the person to the corner points of the relevant attention zone and are defined, for example, as vectors in three-dimensional Cartesian coordinates or in polar coordinates.
[0017] The gaze direction vectors of at least one person can be determined from the photos or video by extracting and analyzing the person's head orientation and / or eye position. Such methods for determining a person's gaze direction vectors from a photo or video of the person are generally known; known image processing algorithms can be used for this purpose. In particular, image analysis libraries in the database and / or a detector trained with training examples, for example, a neural network, can be used.
[0018] Furthermore, it can be provided that the data processing unit extracts the viewing direction vectors of several persons (for example the driver and the passenger) from a single photo series or a single video.
[0019] The display unit then prompts the person to change their position in the room. For example, the person can move their seat back, forward, up, or down, or lean forward or backward. The display unit then prompts the person to fixate the corner points of the attention zone again using head and / or eye movements.
[0020] In the next step, the image acquisition unit takes N second photos of the room with the person and transmits them to the data processing unit. Again, instead of photos, a video can be recorded over a short period of time.
[0021] The data processing unit again extracts a number N of second three-dimensional gaze direction vectors BV 1 ' - BV N ' of the person. Due to the different position of the person, the second gaze direction vectors BV 1 ' - BV N ' differ from the first gaze direction vectors BV 1 - BV N .
[0022] The data processing unit then superimposes the gaze direction vectors BV 1 - BV N with the gaze direction vectors BV 1 ' - BV N ' to determine N three-dimensional intersection points P 1 - PN by triangulation. These N intersection points correspond to the corner points of the attention zone in a common three-dimensional spatial coordinate system. The N intersection points P 1 - PN are stored in the database as corner points of the attention zone.
[0023] The procedure can also be performed with more than two different positions of the person to increase the accuracy of triangulation.
[0024] For example, it may be provided that the person is first asked to move their seat all the way down, then to move their seat all the way up, then to move their seat all the way forward, and finally to move their seat all the way back, whereby after each change of position N viewing direction vectors are determined, which are then used to exactly calculate the intersection points P 1 - PN.
[0025] The present invention offers the advantage that no additional sensors or image acquisition units are required to create the attention zone configuration. It is sufficient to use the image acquisition unit, which is already directed at the person to be analyzed. A picture of the person from behind is not necessary.
[0026] This allows for quick and flexible configuration of the attention zones, ensuring faster conversion or integration of the entire system into an existing image acquisition and analysis system. No new hardware installation is required.
[0027] According to the invention, the method is repeated to create multiple attention zones for a single person. For example, after completing the method according to the invention, the person is asked to fixate on the corner points of another attention zone.
[0028] According to the invention, the method is repeated to create multiple attention zones for different people. For example, it can be provided that all detected people are asked, after completing the method according to the invention, to fixate on the corner points of a different attention zone.
[0029] According to the invention, after the creation and storage of several attention zones of at least one person, the image recording unit continuously detects the gaze direction vectors of this person and transmits them to the data processing unit, and the data processing unit continuously checks whether the detected gaze direction vectors fall within one of the stored attention zones. This allows monitoring of the person's attention during operation.
[0030] According to the invention, the data processing unit can calculate the attention level of at least one person by determining the frequency with which the detected gaze direction vectors of this person fall into one of the stored attention zones per unit of time. The display unit can output a warning message if the attention level falls below a predetermined threshold. For example, a threshold can be defined according to which the person looks into the attention zone defined by the windshield at least once per second.
[0031] The invention further relates to a computer-readable storage medium comprising instructions that cause a data processing unit to execute a method according to the invention.
[0032] The invention further relates to a device for creating a geometric attention zone for at least one person in a three-dimensional space, comprising a single image recording unit, a data processing unit, a database and a display unit, wherein the device is designed to carry out a method according to the invention.
[0033] According to the invention, it can be provided in particular that the space is the interior of a vehicle and the image recording unit is arranged in the direction of travel of the vehicle in front of a driver's seat or in front of a passenger seat and preferably centrally above a windshield or above the dashboard of the vehicle.
[0034] Further features of the invention emerge from the claims, the embodiments and the figures.
[0035] The invention is explained below using an exemplary, non-exclusive embodiment. Fig. 1 shows a schematic representation of a vehicle with a data processing unit for carrying out a method according to the invention; Fig. 2 shows schematic examples of various attention zones according to the invention in a vehicle; Figs. 3a - 3d show schematic representations of the interior of a vehicle during the execution of a method according to the invention; Fig. 3e shows a schematic representation of the determination of the intersection points of two gaze direction vectors of a person.
[0036] Fig. 1 shows a schematic representation of a vehicle 8 with an electronic data processing unit 4 integrated therein for carrying out a method according to the invention. In the interior 2 of the vehicle 8, an image recording unit 3 in the form of a camera is arranged on the ceiling. The camera is designed and arranged such that it can record the interior 2 of the vehicle 8, namely at least the two front seats and the people sitting therein. Several people are sitting in two rows in the interior. In this exemplary embodiment, the vehicle is designed as a passenger car. The data processing unit 4 is connected to a database 5 and has interfaces (not shown) for communication with external electronic components. Furthermore, a display unit 6 is provided, which is also connected to the data processing unit 4. The display unit 6 is arranged in the dashboard of the vehicle 8.
[0037] Fig. 2 shows schematic examples of various attention zones 1, 1', 1", 1‴ according to the invention in the interior 2 of a vehicle 8.
[0038] In this example, the windshield forms the first attention zone 1, the two side mirrors the second attention zones 1', the two side windows the third attention zones 1" and the dashboard the fourth attention zone 1‴. Furthermore, the image recording unit 3 is shown in the form of a video camera arranged centrally above the windshield. The display unit 6 is shown as a centrally arranged touch screen.
[0039] Figs. 3a - 3d show schematic representations of the interior of a vehicle during the execution of a method according to the invention.
[0040] Fig. 3a shows a photo 7 of the interior 2 with three people, where the gaze directions of the driver and the passenger were extracted as gaze direction vectors BV 1. The people were asked to focus on a corner point P 1 of an attention zone 1. The extraction of the gaze direction vectors of the two people can essentially take place simultaneously or separately.
[0041] Fig. 3b shows another photo 7 of the interior 2 after the subjects were asked to focus on another corner point P 2 of the attention zone 1. Again, the gaze direction vectors of the driver and the passenger are extracted as gaze direction vectors BV 2. Here, too, the extraction can be performed simultaneously or separately.
[0042] Fig. 3c shows another photo 7 of the interior 2 after the people were asked to change their position and again focus on the corner point P 1 of the attention zone 1. Again, the gaze direction vectors of the driver and the passenger are extracted as gaze direction vectors BV 1 '.
[0043] Fig. 3d shows another photo 7 of the interior 2 after the people were asked to focus again on the corner point P 2 of the attention zone 1. Again, the gaze direction vectors of the driver and the passenger are extracted as gaze direction vectors BV 2 '.
[0044] These steps are repeated for all vertices of the desired polygonal attention zones until sufficient (but at least two different) gaze direction vectors of each person have been extracted for each attention zone in order to perform triangulation to determine the position of the points P 1 - PN.
[0045] Fig. 3e shows a schematic representation of triangulation for determining the intersection points of two gaze vectors of a person. Two different sitting positions of the person are indicated, with the gaze vectors BV 1 and BV 2 corresponding to the first sitting position, and the gaze vectors BV 1 ' and BV 2 ' corresponding to the second sitting position.
[0046] By superimposing the gaze direction vectors BV 1 and BV 1 ' or BV 2 and BV 2 ', the intersection points P 1 and P 2 can be determined, which are corner points of the attention zone 1.
[0047] The invention is not limited to the described embodiments, but also includes further embodiments of the present invention within the scope of the following patent claims. Bezugszeichenliste
[0048] 1, 1', 1", 1‴Attention zone 2Three-dimensional space 3Image acquisition unit 4Data processing unit 5Database 6Display unit 7Photo 8Vehicle
Claims
1. A computer-implemented method for generating multiple geometric attention zones (1, 1', 1") for at least one person in a three-dimensional space (2) with a single image capturing unit (3), a data processing unit (4), a database (5) and a display unit (6), comprising the following steps: a. outputting, by the display unit (6), a request to the person to fixate points in the space (2) corresponding to the vertices of the attention zone (1, 1', 1") to be created, b. capturing, by the image capturing unit (3), a number N of first photographs (7) of the space (2) with the person and extracting, by the data processing unit (4), a number N of first three-dimensional gaze direction vectors BV1 - BVN of the person, where N is greater than or equal to one, c. outputting, by the display unit (6), a request to the person to change their position in the space (2) and to re-fixate the vertices of the attention zone (1, 1', 1"), d. capturing, by the image capturing unit (3), a number N of second photographs (7') of the space (2) with the person and extracting, by the data processing unit (4), a number N of second three-dimensional gaze direction vectors BV1' - BVN' of the person, e. superimposing, by the data processing unit (4), the gaze direction vectors BV1 - BVN with the gaze direction vectors BV1' - BVN' in order to determine N three-dimensional intersection points P1 - PN, f. storing, in the database (5), an attention zone (1, 1', 1") as a polygon in the space (2) with the vertices P1 - PN, wherein g. the method is repeated to generate multiple attention zones (1, 1', 1") of a single person or of several persons, and h. the number of vertices N is greater than or equal to three.
2. The method according to claim 1, characterised in that the number of vertices N is equal to four.
3. The method according to claim 1 or 2, characterised in that the space (2) is the interior of a vehicle (8).
4. The method according to one of claims 1 to 3, characterised in that a. after generating and storing one or multiple attention zones (1, 1', 1") of at least one person, the image capturing unit (3) continuously detects the gaze direction vectors of this person and transmits them to the data processing unit (4), and b. the data processing unit (4) verifies whether the detected gaze direction vectors fall within one of the stored attention zones (1, 1', 1").
5. The method according to claim 4, characterised in that a. the data processing unit (4) calculates a degree of attention of at least one person by determining the frequency with which the detected gaze direction vectors fall within one of the stored attention zones (1, 1', 1") per unit of time, and b. the display unit (6) outputs a warning message when the level of attention drops below a predefined threshold.
6. A computer-readable storage medium, comprising commands which cause a data processing unit (4) to execute a method according to one of claims 1 to 5.
7. A device for generating multiple geometric attention zones (1, 1', 1") for at least one person in a three-dimensional space (2), comprising a single image capturing unit (3), a data processing unit (4), a database (5) and a display unit (6), characterised in that the device is designed for executing a method according to one of claims 1 to 5.
8. The device according to claim 7, characterised in that the space (2) is the interior of a vehicle (8) and the image capturing unit (3) is arranged in front of a driver's seat or in front of a passenger's seat in the direction of travel of the vehicle and preferably centrally above a windscreen of the vehicle.