Processing image data of grandstands to identify an area of ​​interest

An automated method for detecting areas of collective attention in seated crowds using head orientation analysis addresses the inefficiencies of manual monitoring, enabling rapid incident detection and response in large gatherings.

FR3163195A1Pending Publication Date: 2025-12-12ORANGE SA
View PDF 1 Cites 0 Cited by

Patent Information

Application Number
FR2024006192
Authority / Receiving Office
FR · FR
Patent Type
Applications
Current Assignee / Owner
Filing Date
2024-06-11
Publication Date
2025-12-12

AI Technical Summary

Technical Problem

Existing methods for detecting incidents in seated crowds, such as in stadiums, are inefficient and rely heavily on manual monitoring, lacking an automated process to quickly identify potential threats from abnormal crowd movements.

Method used

An automated method to estimate head orientations of spectators and detect areas of collective attention by analyzing image data, generating a signal for potential incidents based on significant head orientations towards a specific area, which can trigger alerts or camera zooms.

Benefits of technology

Enhances the reliability of incident detection by leveraging collective spectator intelligence, allowing for rapid response to potential threats through automated means, reducing reaction time and improving crowd safety in large gatherings.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure 00000000_0000_ABST
    Figure 00000000_0000_ABST
Patent Text Reader

Abstract

A method for processing image data of a space containing spectators, for example in a stadium grandstand, is proposed. The processing involves: - estimating (P4), within a typical neighborhood of spectators in the image, the respective head orientations of spectators in said neighborhood, - detecting (P5), at least based on said estimated head orientations, whether said heads are oriented towards a specific area of ​​the space, in order to generate (P6), if applicable, a signal containing data from said area as the spectators' attention zone within said space. Abstract figure: Figure 1
Need to check novelty before this filing date? Find Prior Art

Description

Title of the invention: Processing image data of grandstands to identify an area of ​​attention technical field

[0001] This disclosure relates to the processing of image data, in particular of spectator stands attending an event, for example a sporting event.

[0002] It finds a possible but not limiting application in the monitoring of stands, particularly for security reasons. Previous technique

[0003] Typically, very large gatherings of people present an increasingly acute security challenge. The world of sport regularly experiences precisely this concern. France is meeting this challenge by hosting the 2024 Summer Olympic Games.

[0004] Stadiums offer numerous advantages for managing spectator safety, as the arrangement of seated spectators presents far fewer risks than a moving crowd, which carries the risk of a stampede. In particular, it is important to detect any incident as early as possible, including in the stands of a seated stadium, because if an incident is not detected and dealt with very quickly, it can escalate into an uncontrolled crowd phenomenon.

[0005] There are methods for managing crowd safety. These include video analysis, for example, which focuses on capturing human density, speed (of a procession), and local movements of groups of people. Thus, for crowd monitoring, particularly in motion, a key indicator of an incident is precisely the abnormal movement of groups of people.

[0006] Conversely, in a stadium or performance hall, equipped with seats, the detection of the beginnings of incidents on the basis of "movements" is not applicable.

[0007] Ensuring the safety of seated spectators (who are stationary or relatively immobile) is a complex task in a large stadium. No technique other than manual monitoring is known: an operator uses a motorized camera to scan the crowd and, if in doubt, can point the camera axis in a chosen direction and zoom in on an area of ​​interest.

[0008] The reaction time to an incident is crucial. Therefore, the detection time is important and an automated process is sought.

[0009] Computer vision and artificial intelligence algorithms have made significant progress recently. Pre-trained models already make it possible to detect more of a thousand types of common objects in an image, including human faces for example.

[0010] However, alerting to an incident by detecting an abnormal situation is more complex to formulate for a machine learning algorithm, which precisely requires being fed by a substantial set of examples.

[0011] Thus, the prior formulation of many threatening situations is very complex to achieve. Summary

[0012] This disclosure improves the situation.

[0013] A method for processing image data of a space containing spectators is proposed, the method comprising: - to estimate, within a typical neighborhood of spectators in the image, the respective head orientations of spectators in said neighborhood, - detect, at least according to said estimated head orientations, whether said heads are oriented towards an area of ​​space, to generate, where appropriate, a signal containing data from said area as an area of ​​spectator attention in said space.

[0014] Thus, according to the proposed approach, the aim is to "capture," in real time, the collective intelligence of a group of people who can react to any type of abnormal situation. In this approach, the spectators themselves possess the relevant information to be identified. For example, if the attention of a significant number of people is drawn to the same area of ​​a space where spectators gather, then this area very likely presents a noteworthy situation, which benefits from closer analysis, particularly if it is an incident requiring a response.

[0015] In addition to the increased reliability due to collective intelligence (a group of spectators in a neighborhood), this approach also has the advantage of focusing the effort on a simple detection of human faces (or "heads" hereafter) of the spectators and the orientations of these heads.

[0016] The aforementioned space "grouping spectators" can typically be a stand, for example in a stadium, performance hall, theatre, or other venues.

[0017] Thus, if a significant number of eyes converge on the same area of ​​this space then this area presents a particular interest or, for example, a potential incident.

[0018] In one embodiment, the aforementioned signal may include geographic coordinate data for the detected area of ​​attention. Typically, if the image is acquired by a mobile camera, the geographic coordinates can be deduced from the current settings of the mobile camera.

[0019] In this case, law enforcement can be alerted by this signal and go to the location, using the geographical coordinates, to verify if a danger is indeed imminent.

[0020] Alternatively, the aforementioned signal can be transmitted to a surveillance camera to zoom in on this area of ​​attention (this camera being for example connected to a monitoring center, for rapid intervention for example).

[0021] Alternatively, a digital zoom can be performed in the same image acquired by the first camera, towards the detected area of ​​attention, and the zoomed image can be transmitted to a monitoring center or to a control room for insertion of this zoomed image into a television stream, or other.

[0022] In one embodiment, for each spectator in the neighborhood associated with a candidate zone as an attention zone, an angular deviation is estimated between: - a straight line passing through the spectator's head and the candidate zone, and - a head orientation of the spectator. This angular deviation is then all the smaller in absolute value the more the spectator's head is oriented towards the aforementioned candidate area.

[0023] In this embodiment, a representative average of the angular deviations can be estimated over all the spectators in the neighborhood, for the aforementioned candidate area, and the estimated average can be compared to a threshold to determine the candidate area as an attention area (i.e. whether the candidate area is an attention area or not).

[0024] Such an embodiment may further include a determination of a natural orientation of the heads of the neighborhood towards a game action zone located in front of said space, and the aforementioned estimated average may then be weighted by an angular difference between the head orientation of the spectator and his natural orientation towards the game action.

[0025] It is also possible to use a wider image than that of the neighborhood to determine this natural orientation, as typically illustrated by [Fig.2] of a wide field image, described later.

[0026] In an embodiment where the space grouping spectators is typically a multi-row, multi-column grandstand, the aforementioned typical neighborhood may consist of spectators located: - at the same rank as a candidate area as an area of ​​attention, and at least one rank above and at least one rank below said same rank, and - in the same column as the candidate area, and at least one column to the left and at least one column to the right of said same column.

[0027] In this case, for the estimation of the average, spectators of the rank below may be given a greater weight than spectators of the same rank mentioned above, and spectators of the rank above may be given a lesser weight than spectators of the same rank mentioned above.

[0028] Such an implementation thus respects the principle of taking into account the natural orientation of a spectator's head towards the action of the game. Indeed, a spectator in a lower row has a natural tendency to look at the action of the game generally in front of them and therefore has their head naturally oriented generally downwards. If, on the contrary, their head is detected as being oriented upwards (for example, towards the upper stand), then this situation is unusual and more weight is given to such a determination in the estimation of the average.

[0029] Similarly, more weight can be assigned to spectators who are several columns away from the candidate area than to those who are only one column away, for example from the candidate area, because, even if spectators are far from the candidate area, the latter still attracts their attention.

[0030] In one embodiment, the estimation and detection of the process are repeated for a plurality of successive neighborhoods in one or more successively acquired images.

[0031] For example, the aforementioned threshold can be determined based on the estimated averages for these successive neighborhoods.

[0032] Typically, before estimating the respective head orientations of nearby spectators, head detection of the spectators (as objects recognized, for example, by artificial intelligence) can be implemented, and from this, it is then possible to estimate the orientation of the heads thus detected. It should be understood in particular that this head detection is not necessarily facial recognition of a specific individual but simply the detection of an object corresponding to a human head.

[0033] This detection of spectator heads can typically be followed by a determination of the respective positions of these heads, which makes it possible to determine, for each spectator, the aforementioned line passing through the spectator's head and the candidate area as the attention zone.

[0034] According to another aspect, a computer program is proposed comprising instructions for implementing all or part of a process as defined herein when this program is executed by a processor. According to another aspect, a non-transient, computer-readable recording medium is proposed on which such a program is recorded.

[0035] According to another aspect, a device for processing image data of a grandstand including spectators is proposed, comprising a processing circuit for implementing the above process. Brief description of the drawings

[0036] Other features, details and advantages will become apparent from the detailed description below and from the analysis of the accompanying drawings, in which: Fig. 1

[0037] [Fig-1] illustrates an example of a sequence of general steps, of a process of the type presented above. Fig. 2

[0038] [Fig.2] shows a wide-field image of a spectator stand, showing in particular a general orientation of the spectators' heads towards a game action taking place in front of the stands. Fig. 3

[0039] [Fig.3] illustrates a possible embodiment for detecting head orientations of spectators in a stand. Fig. 4

[0040] [Fig.4] illustrates the head orientations of a neighborhood towards an attention zone represented by a grey disc, while other spectators, outside this area (typically at the bottom of [Fig.4]), have their heads naturally oriented more towards a game action on the field. Fig. 5

[0041] [Fig.5] illustrates an example of a possible neighborhood to be associated with a candidate zone in as an area of ​​attention. Fig. 6

[0042] [Fig.6] illustrates an angle normally expected between a natural orientation of the head of a spectator towards the game action, and a real current orientation of that spectator's head. Fig. 7

[0043] [Fig.7] illustrates a succession of steps more specific than those of [Fig.1], in an example of a specific implementation of the process. Fig. 8

[0044] [Fig.8] schematically illustrates one possible embodiment of a device of the type presented above. Description of the implementation methods

[0045] Reference is now made to [Fig. 1] showing an example of a sequence of general steps of a method according to an embodiment. This method proposes an automatic determination of an area of ​​attention by spectators in a grandstand, using computer image data processing to analyze the orientation of the spectators' attention themselves.

[0046] During a first PI step, the process is triggered by manual intervention, or automatically via a request from another system (for example, a surveillance system or a television channel control room), or regularly (for example, every five minutes), or otherwise.

[0047] The second step P2 comprises head detection, where heads are simply identified as heads (or faces). It is therefore not facial recognition. These heads are assigned ID#i identifiers for set calculations described later.

[0048] In step P3, the respective positions POS_i of these heads are determined to define lines (or directions) towards a region of interest, a candidate in the aforementioned set calculations. The positions can be determined in the 2D image acquired by a camera filming the stands, for example.

[0049] However, it is also possible to determine positions in 3D space by determining the distance between each detected object and the camera, which typically allows the conversion from a cylindrical coordinate system to an orthonormal system. This distance estimation can be implemented by artificial intelligence, simply from a front view. Indeed, (relative) proximity estimation approaches can be robust using only a single monocular view. Thus, using a single image, such artificial intelligence can provide a table of values ​​corresponding to an inference result for each pixel. This approach has the advantage of using a standard, low-cost camera, and it is not necessary to use a stereoscopic camera or one equipped with a sensor such as LiDAR.However, it is advisable to avoid using cameras with optics that induce strong distortions, such as wide-angle lenses (e.g., fisheye lenses).

[0050] An advantage of applying such an embodiment to the context of crowd tracking in a stadium is that the arrangement of people mainly seated and facing the main area of ​​interest (which normally takes place in the center of the stadium) makes the detection of heads all the more efficient: there are indeed few occlusions, as shown in [Fig.2].

[0051] For example, the result of multiple detections of these objects can be a list of information, each item including a rectangle encompassing the detected occurrence, as illustrated in [Fig. 2]. This position data, in pixels, in the acquired image, can be converted into an absolute location due to knowledge of the geometry of the stands, the location of the fixed camera, and the orientation of the camera's optics if it is motorized.

[0052] The fourth step P4 aims to determine the direction of attention of each spectator whose head has been detected. This step can only be implemented in a restricted neighborhood as described later, and not in a whole global wide-field image, especially for computational efficiency.

[0053] Generally speaking, "attention orientation" means a determination according to various possible modalities: orientation of the gaze, orientation of the face or more generally orientation of the spectator's head.

[0054] For large audiences (a situation in which the proposed processing is particularly useful), it can be advantageous to implement head orientation "capture." The established term in the English-language literature is "head pose estimation." Indeed, learning and inference models exist for predicting head orientation from a single image. Most of them perform face and facial marker detection (eyes, nose, mouth) to facilitate learning. These approaches are then valid for the typical interval [-90°, +90°]. Other approaches do away with facial markers to rely directly on head detection (without facial markers alone) and thus offer estimation over a much wider range. A difficulty then arises, related to the discontinuity of the rotation angle from +180° to -180°.Machine learning models struggle to handle discontinuities, but recent research proposes a radical change in the representation method. Euler angles offer the dual advantage of conciseness and ease of interpretation. Conversely, a general rotation matrix lacks these advantages but provides the quality of continuity, useful for machine learning algorithms. Without loss of generality, orthonormal matrices can be used: the last column is derived from the first two, so there are (only) six parameters to determine. Detection is then possible, even for people facing away from you.

[0055] However, in a particular embodiment, capturing attentional orientation by detecting facial markers can also be a simple and suitable solution for the application envisaged here. Figure 3 illustrates such detection.

[0056] The fifth step P5 aims to determine whether a particular area of ​​attention emerges. For this purpose, it is necessary to perform a set-theoretic processing of the information carried by each spectator (location and axis of attention).

[0057] In particular, we seek to determine whether a small area of ​​the crowd (or near the crowd) is the focus of multiple attentions from nearby locations, typically within a given neighborhood. For example, a threshold of one hundred people (all with their heads facing the area), among the two hundred nearest neighbors of that area, must be reached for the aforementioned area to be considered an area of ​​interest.

[0058] A specific method is described below, by way of illustration, for determining that a small area is the center of attention of people located in the vicinity.

[0059] It is assumed a priori that all persons are positioned according to a grid corresponding to the arrangement of seats in a grandstand.

[0060] Spatial processing can be used for 2D matrices, particularly in computer vision, using for example a convolution and cross-correlations algorithm.

[0061] A 2D matrix is ​​denoted hereafter as “I”, in which each group of pixels has a value corresponding to an angle alpha in radians of the orientation of the attention of the person located in the location. Furthermore, the entire area to be monitored can be scanned successively by small areas. The diagram in [Fig. 5] illustrates the use of a scanning disk; this can also be a rectangle if the area to be monitored lends itself to a grid. In all cases, the scanning shape should be of moderate size, as it is desirable to capture the attention of people located near the area being sought.

[0062] An elementary scanning area is then denoted as "O" and a neighborhood of people around this elementary area as "V(O)" (for example, the set of pixels located in a rectangle or a disk centered at O). In the diagram in [Fig. 5], these are the people appearing in gray.

[0063] For any point M in the neighborhood V(O), we define a unit cross-correlation coefficient c(0,M) = -cos(OM,alpha).

[0064] Thus, for any zone O to be swept, we define a mean set coefficient of neighborhood: c(O) = S0M(M in V(O)) [ - cos(OM, alpha(M) ] / card(V(O).

[0065] The interpretation is as follows: for a given point O, such a coefficient with a value close to 1 means that almost all neighbors are looking at it.

[0066] Admittedly, such a treatment, as such, is costly in terms of the number of calculations, including angle correlation operations, but the various techniques for quantifying and parallelizing the treatments nevertheless allow a real-time approach.

[0067] A very moderate precision is sufficient, and a true cosine calculation can be avoided in 64-bit floating-point values. Ultimately, a simple table of values ​​with, for example, only 100 intervals is more than adequate. Finally, a precision of 10⁶ is also sufficient, so that one can easily manipulate natural numbers.

[0068] Finally, this 2D matrix of natural integers (or image) is suitable for a massively parallel processing architecture thanks to "computer vision" technologies.

[0069] In the case of bleachers, the seats are typically oriented towards the playing field. A user's attention is naturally directed in this direction. Conversely, the more a user "looks backwards," the more valuable this information is for the purpose of detecting atypical situations.

[0070] Thus, the unit cross-correlation coefficient can be weighted by a deviation coefficient from the nominal axis for that user. This is illustrated in [Fig. 6]. With an interval [0,1], using, for example: (l-cos(alpha,axis)) / 2. Here again, quantization with only 100 values ​​allows us to limit ourselves to integer operations.

[0071] The nominal axis for a given spectator can be simple to determine. It can naturally be the orientation axis of their seat. In a simple implementation, it can be decided to assign more weight to spectators in the vicinity who are seated in seats that are below a presumed area of ​​interest (spectators below the circle of [Fig. 5]) than to those who are seated in seats located above this area, and to assign an intermediate weight to spectators seated on the "same row" as the area of ​​interest.

[0072] However, the area of ​​interest may be outside the stands (for example, in one of the access doors to the stands). Indeed, the area of ​​interest may not necessarily coincide exactly with the area where the spectators are seated and may extend beyond it (particularly at the level of the access doors, for example), while noting that it is important that there be spectators not too far from the area of ​​interest since they are precisely the carriers of information.

[0073] Alternatively, the nominal axis can be determined more generally by the area of ​​attention of the spectacle. This area can change over time: for example, it could be the position of the ball in a football match. Particularly in the world of sports, there are numerous video analysis methods that allow the area of ​​interest of the activity to be determined in real time. Typically, an average of head orientations in a wide-field image, as illustrated in [Fig. 2], can also be used to determine a nominal axis of normal head orientation towards the action of the game. Furthermore, at the precise moment of calculating the weighting coefficient for a given spectator, it is possible to determine the angular difference between the spectator's axis of attention and the nominal axis of the action on the field.

[0074] In the next general step P6 of [Fig. 1], a signal containing the coordinates of a detected area of ​​interest can be transmitted to a third-party system, for example, to generate an alert for an incident detection surveillance system. This signal can also be transmitted to a control room (for example, a mobile control room truck located near the stadium) to generate an audiovisual feed for broadcast by a television channel, from feeds acquired by different cameras. In such an embodiment, the signal generated in step P6 can be interpreted by a control room device capable of controlling a camera to zoom in on the area of ​​interest, and thus film a scene that a significant number of spectators are watching (for example, a person well-known to the general public, or a spectator dancing, or others).

[0075] Figure 7 illustrates a possible embodiment with an alternative implementation of step P5 above, for determining whether a candidate area O is an area of ​​attention (or at least a possible area of ​​interest). In step S1, a current image IM is considered, for example, a wide-field image. In step S2, respective neighborhoods of possible candidate areas O are defined. Each candidate area O can correspond to a group of pixels in a few rows and columns of the image IM, and the dimensions of this area can correspond to the apparent dimensions of a seat in the image IM, for example. Thus, the image IM is gridded, and each element of the grid can have the dimensions of a seat and will be considered, in what follows, as a candidate area O as a possible area of ​​attention.

[0076] In step S2, a neighborhood M(i) is assigned to each candidate zone O(i). A neighborhood can extend, for example, to two rows (of seats) above and below the candidate zone O and to three columns (of seats) to the left and right of the zone O. The number of neighbors M of a neighborhood can thus be about thirty points M (5x7-1).

[0077] In step S3, for a candidate zone O as the attention zone, each angle ANG between is first calculated: - the 0RI_M orientation of the head of a neighbor M, and - the right passing through this neighbor M and the candidate zone O.

[0078] An angle measurement metric (for example, its sine or tangent, or directly the value of the angle in radians) is used to calculate the sum of these angles (in absolute values ​​of these angles) over all the neighbors M of the candidate zone O, and, from this, an angular mean MOY(O) can be calculated as a possible set-theoretic metric over the neighborhood of a candidate zone O. The angular mean MOY(O) can be weighted according to the relative positions of each neighbor M with respect to the candidate zone O. For example, more weight can be assigned to the angles formed by neighbors seated in the lower tiers, with an even greater weight for the second row below the candidate zone O. Similarly, the weight Wp can increase according to the number of columns between the candidate zone O and the seat of a neighbor M.

[0079] Once the average (thus weighted) has been calculated for a candidate area O, this average is compared to a threshold in step S4, and if this average is lower than the THR threshold, then the area Oj with this average can be a target area. The aforementioned threshold can be parameterized. It can be, for example, the smallest average determined for all successive candidate areas O in the given wide-angle IM image. Thus, for example, in step S5, the averages over their neighborhood are successively calculated for each candidate area O, and selects the one with the smallest average to, for example, order a camera zoom on that area as a possible area of ​​attention.

[0080] However, this step S5 is optional (and illustrated by dashed lines for this purpose). Alternatively, it is still possible to determine a THR threshold value based on feedback from previous experiments, independently of determining a minimum mean in the image. Thus, if a candidate area is identified whose mean is below this THR threshold, then step S6 is triggered (OK arrow at the exit of test S4). Otherwise (KO arrow at the exit of S4), the processing is repeated on another IM image.

[0081] Step S6 corresponds to step P6 of [Fig.1], namely the transmission of a signal including coordinates of the attention zone (the average of which has been determined to be less than a THR threshold).

[0082] The steps SI to S5 (and possibly S6) can then be reproduced on a new general image acquired at step S7 (NEXT IM), either with another camera position to monitor other stands of the stadium for example, or with the same camera position and at a later time typically.

[0083] All or part of these steps can be implemented by a device as illustrated in [Fig. 8], the device comprising a CT processing circuit equipped with: - an IN input interface to receive image data acquired by a CAM camera, and possibly current camera settings data (if it is, for example, a mobile camera), - a MEM memory storing, in particular, instruction data from a computer program for the above implementation (and possibly, for example, data on estimated averages for different candidate areas as areas of attention for the implementation of step S5 typically), - a PROC processor capable of accessing the MEM memory to read and execute the instructions of the aforementioned computer program, for the implementation of the method, the processor also receiving image data acquired via the IN input interface, on the basis of which one or more attention zones can be detected as described above with reference to one and / or the other of Figures 1 and 7, - and an OUT output interface capable of delivering, in particular, the SIG signal containing data from the area of ​​attention detected by the implementation of the process, for example, geographic coordinate data of this area of ​​attention (determined according to the camera settings, for example).

[0084] This signal can be shaped to feed a surveillance system and / or an additional camera on the stadium capable of zooming the image in an area of ​​attention whose coordinates have been transmitted and interpreted by the additional camera. Industrial application

[0085] These technical solutions may be applicable in particular to ensure the safety of spectators, especially for events which have worldwide media coverage such as football cup matches or the Olympic Games.

[0086] This disclosure is not limited to the examples described above, which are only examples, but encompasses all the variations that a person skilled in the art may consider in the context of the protection sought.

[0087] Knowledge of the geometry of the stands and the current settings of the acquisition camera can be used to determine the absolute location of the detected object. In a particular embodiment, a stereoscopic camera can be used to provide location information more directly.

[0088] Instead of using one or more fixed cameras, the method can also be applied to mobile cameras, whether aerial drones or devices moving along a cable. In this case, the capture system should be equipped with a positioning device, for example GPS or simply a linear device along the carrier cable, which allows calculations of coordinate system changes to be performed to return to a situation of the type described above.

[0089] The method can be applied to an audience with seating. However, it involves tracking people with relatively limited mobility. More generally, the method can then be applied to any type of static audience or crowd, at least for the duration of calculating a candidate area as a zone of interest.

Claims

Demands

1. A method for processing image data of a space grouping spectators, the method comprising: - estimating (P4), in a current neighborhood of spectators in the image, the respective head orientations of spectators in said neighborhood, - detecting (P5), at least as a function of said estimated head orientations, whether said heads are oriented towards an area of ​​the space, to generate (P6), where appropriate, a signal comprising data of said area as the area of ​​attention of spectators in said space.

2. A method according to claim 1, wherein the signal includes geographic coordinate data of said attention area.

3. A method according to claim 2, wherein the image is acquired by a mobile camera and the geographic coordinates are deduced from current settings of the mobile camera.

4. A method according to any one of the preceding claims, wherein the signal is transmitted to a surveillance camera to perform a zoom in said area of ​​attention.

5. A method according to any one of the preceding claims, wherein, for each spectator in the neighborhood associated with a candidate zone as an attention zone, an angular deviation (ANG(MO, ORI_M)) is estimated between: - a straight line (MO) passing through the spectator's head and the candidate zone, and - a head orientation of the spectator (ORI_M), said angular deviation being smaller in absolute value the more the spectator's head is oriented towards said candidate zone (0).

6. A method according to claim 5, wherein a representative average of the angular deviations is estimated over all the spectators in the neighborhood (S3), for said candidate area, and the estimated average is compared to a threshold (S4) to determine the candidate area as an attention area.

7. A method according to claim 6, further comprising determining a natural orientation of the heads of the neighborhood towards a playing area located in front of said area, the estimated average being weighted by an angular difference between the viewer's head orientation and said natural orientation.

8. A method according to any one of the preceding claims, wherein said spectator area is a multi-row, multi-column grandstand, and wherein the current neighborhood consists of spectators located: - in the same row as a candidate zone as an attention zone, and at least one row above and at least one row below said same row, and - in the same column as the candidate zone, and at least one column to the left and at least one column to the right of said same column.

9. A method according to claim 8 taken in combination with any one of claims 6 and 7, wherein, for the purpose of estimating the average, spectators in the row below are assigned a greater weight than spectators in the same row, and spectators in the row above are assigned a lesser weight than spectators in the same row.

10. A method according to any one of the preceding claims, wherein estimation and detection are repeated for a plurality of successive neighborhoods in one or more successively acquired images (S7).

11. A method according to claim 10, taken in combination with any one of claims 6 to 9, wherein said threshold is determined (S5) as a function of the estimated averages for said successive neighborhoods.

12. A method according to any one of the preceding claims, wherein the estimation of the respective head orientations of the spectators in the vicinity is preceded by a spectator head detection (P2).

13. A method according to claim 12, wherein the detection of spectator heads is followed by a determination of the respective positions of said heads (P3), to determine, for each spectator, a straight line (MO) passing through the spectator's head and a candidate area (0) as an attention zone.

14. A computer program comprising instructions for carrying out the method according to any one of the preceding claims, when this program is executed by a processor.

15. Image data processing device for a space grouping spectators, comprising a processing circuit (PROC, MEM, IN, OUT) for implementing the method according to any one of claims 1 to 13.

Citation Information

Patent Citations

  • Control apparatus, control system, and control program

    EP3713211A1