METHOD FOR DETECTING THE MOVEMENT OF AT LEAST ONE OBJECT, ELECTRONIC DEVICE, SYSTEM, COMPUTER PROGRAM PRODUCT AND MEDIUM FOR IT
Patent Information
- Authority / Receiving Office
- DE · DE
- Patent Type
- Patents
- Current Assignee / Owner
- Filing Date
- 2022-12-14
- Publication Date
- 2026-03-25
AI Technical Summary
Existing object detection solutions require heavy physical infrastructure and telecommunications resources, are dependent on object connectivity, and are costly, making them impractical for widespread application.
A method that analyzes a video stream to detect moving objects by generating a motion map from differences between successive images, segmenting this map into geometric shapes, and tracking these shapes across frames, allowing for object identification and counting without requiring extensive infrastructure or connectivity.
Enables efficient and economical detection of moving objects by analyzing video streams, reducing the need for costly infrastructure and telecommunications, and providing real-time object tracking and counting capabilities.
Description
1. Domaine technique
[0001] This application relates to the field of object detection, such as the detection of moving objects. Depending on the implementation of the solution described in this application, this may involve various objects, connected or not, such as vehicles, people, animals, robots, and / or machines.
[0002] This application relates in particular to a method for detecting the movement of at least one object, as well as a corresponding electronic device, system, computer program product and storage medium. 2. État de la technique
[0003] For several decades, we have witnessed a surge in automation in industry, transportation, and agriculture, and more recently, the development of connected environments in the private and public spheres (smart homes, smart offices, smart cities). These automations and connected environments often involve objects, whether connected or not, and frequently require (or would benefit from) automatic detection of these objects.
[0004] Some existing solutions rely on information transmitted by the objects themselves. However, these solutions are only applicable to connected objects and generally require the use of telecommunications resources and / or specific protocols to receive this information. Furthermore, they are often dependent on external constraints, such as the willingness of object owners to enable this data transmission, the load on the objects, network transmission quality, etc. Other existing solutions can be more easily applied to non-connected objects. For example, in the transportation sector, there are various vehicle counting techniques. Some road traffic analysis techniques record the passage of vehicles using inductive loops placed under the road surface. A major drawback of such technologies is the cost of the fixed infrastructure required for these loops.Thus, if it is not planned from the time of the road construction, the installation of inductive loops requires heavy development work to integrate the inductive loops into the road and the necessary equipment at the roadside.
[0005] Therefore, there is a need for an object detection solution that does not present all of the above drawbacks. US patent 2003 / 044045 A1 (Schoepflin et al. – “Video object tracking by estimating and subtracting background”) describes tracking an object across a plurality of image frames. In an initial frame, an operator selects an object. The object is distinguished from the remaining background portion of the image to define a background and a foreground. A background model is used and updated in subsequent frames. A foreground model is used and updated in subsequent frames. Pixels in subsequent frames are classified as belonging to the background or the foreground.In subsequent images, decisions are made, including: which pixels are not in the background; which foreground pixels need updating; which background pixels were incorrectly observed in the current image; and which background pixels are being observed for the first time. Additionally, mask filtering is performed to correct errors, eliminate small islands, and maintain the spatial and temporal consistency of a foreground mask.
[0006] US patent 6,240,197 B1 (Christian et al. – “Technique for disambiguating proximate objects within an image”) discloses a technique for resolving the ambiguity of nearby objects in an image. In one embodiment, the technique is performed by obtaining an image that is a representation of a plurality of pixels, in which at least one group of substantially adjacent pixels has been identified. Discontinuities are identified in each of the identified groups of substantially adjacent pixels. Each of the identified groups of substantially adjacent pixels is divided according to the identified discontinuities. It is then determined whether each of the divided identified groups of substantially adjacent pixels corresponds to an object to be classified.
[0007] US patent 2021 / 0133483 A1 (Prabhu et al. – “Object detection based on pixel differences”) concerns machine learning-based object recognition using pixel difference information. A difference image generated by subtracting one or more previous images from a current image can be provided as input to a machine learning engine. The machine learning engine can produce a detected object or action based, at least in part, on the difference image. In this way, temporal information about the object can be provided and used by a structured machine learning model to accept the image input.
[0008] The US document 2015 / 0279051 A1 (Kovesi et al. - "Image detection and processing for building control") discloses an occupancy detection for the environmental control of a building.An apparatus includes at least one image sensor and a controller, the controller being operational to: obtain a plurality of images from the image sensor, determine a color difference image corresponding to a color difference between consecutive images, determine areas of the color difference image in which the color difference is greater than a threshold, calculate a total change area as an aggregation area of the determined areas, create an occupancy candidate list by filtering the determined areas if the total change area is less than a light change threshold, in which each occupancy candidate is represented by a connected component, track the occupancy candidate(s) over a period of time, and detect occupancy if it is detected that one or more of the occupancy candidates have moved beyond a movement threshold over the period of time.
[0009] However, there is a need for a simple implementation solution that does not require the establishment of heavy physical infrastructure (mechanical or telecommunications) and is more economical in terms of digital resource consumption. The purpose of this application is to propose improvements to at least some of the drawbacks of the current state of the art. 3. Exposé de l'invention
[0010] The present application aims to improve the situation using a process implemented at least partially in an electronic device.
[0011] A motion detection method according to claim 1 is disclosed.
[0012] In some embodiments, the value of a portion of the motion map obtained takes into account a variation of at least one R, G, B component, between at least two of said images obtained, for an area of said images obtained corresponding to said portion of the motion map.
[0013] In some embodiments, the value of a portion of the motion map obtained takes into account a variation of each component R, G, B between at least two of said images obtained, for an area of said images obtained corresponding to said portion of the motion map.
[0014] In some embodiments, the value of a portion of said first motion map is a binary value taking into account the differences between the values of each video channel for said area of said images.
[0015] In certain embodiments, said value of a portion of said first motion map takes into account the maximum of the differences of each video channel for said area of said images.
[0016] In some embodiments, the method includes obtaining a motion map for each video channel of said images, the motion map of a video channel of said images being cut into portions representing the differences for an area of at least two of said images, for the video channel associated with said motion map, and said first motion map is the motion map, among the motion maps associated with each of the video channels of said images, which has the most portions whose binary value is representative of a difference for an area of said images, for the associated video channel.
[0017] In some embodiments, said geometric shape is a rectangular shape.
[0018] In some embodiments, the process includes, prior to cutting the first motion map, rotating the first motion map. In some embodiments, the motion map is a binary map whose dimensions correspond to those of the images in the video stream.
[0019] In some embodiments, said process is implemented iteratively on successive images of the video stream and comprises: obtaining a first movement map and a second movement map temporally preceding said first movement map, a conditional association of at least one first encompassing geometric shape obtained for the first movement map with at least one second encompassing geometric shape obtained for said second movement map, said association taking into account the relative positions of said first and second encompassing shapes on said first and second movement map.
[0020] In some embodiments, said association takes into account a distance between the centroids of the first and second encompassing geometric shapes obtained.
[0021] In some embodiments, the second geometric shape is the encompassing geometric shape of said second movement map that is spatially closest to said first encompassing geometric shape.
[0022] In some embodiments, said association is implemented when said distance between the centroids of the first and second encompassing geometric shapes obtained is less than a first distance.
[0023] In some embodiments, the process includes assigning the same object identifier to said first and second associated geometric shapes.
[0024] In some embodiments, the process includes assigning an object type to at least one encompassing geometric shape, said assignment taking into account a size, a width and / or a length of said encompassing geometric shape (for example, a ratio between the width and / or the length of the encompassing geometric shape).
[0025] In at least one embodiment, the method includes counting the encompassing geometric shapes of at least one movement map, said counting taking into account proximity to at least one so-called part of interest of said movement map.
[0026] In some embodiments, said counting of encompassing geometric shapes takes into account a direction of movement of an object corresponding to several encompassing geometric shapes associated with each other, of several movement maps, between said movement maps.
[0027] In at least one embodiment, the process comprises: obtaining a third histogram counting the number of moving portions per column in at least one third of said bounding shapes obtained; obtaining a fourth histogram counting the number of moving portions per row in said third bounding shape obtained; obtaining at least one fourth bounding geometric shape corresponding to the intersection in the movement map of the columns including at least one non-zero element of the third histogram and the rows including at least one non-zero element of the fourth histogram.
[0028] In some embodiments, the said obtaining of the said third and fourth histograms are implemented successively on at least one encompassing geometric shape previously obtained from the movement map.
[0029] The characteristics, presented individually in this application in connection with certain embodiments of the process of this application, may be combined with each other according to other embodiments of this process.
[0030] In another respect, the present application also relates to an electronic device according to claim 2.
[0031] The present application also relates to a computer program according to claim 13.
[0032] The present application also relates to a recording medium readable by a processor of an electronic device according to claim 14.
[0033] The programs mentioned above may use any programming language, and be in the form of source code, object code, or code intermediate between source code and object code, such as in a partially compiled form, or in any other desirable form.
[0034] The information storage media mentioned above can be any entity or device capable of storing the program. For example, a medium can include a storage means, such as a ROM, for example a CD-ROM or a microelectronic circuit ROM, or a magnetic recording means.
[0035] Such a storage device could be, for example, a hard drive, a flash memory, etc.
[0036] On the other hand, an information medium can be a transmissible medium such as an electrical or optical signal, which can be transmitted via an electrical or optical cable, by radio, or by other means. A program according to the invention can, in particular, be downloaded from a network such as the Internet.
[0037] Alternatively, an information carrier may be an integrated circuit in which a program is incorporated, the circuit being adapted to execute or to be used in the execution of any of the embodiments of the method which is the subject of this patent application. 4. Brève description des dessins
[0038] Other features and advantages of the invention will become clearer upon reading the following description of particular embodiments, given by way of simple illustrative and non-limiting examples, and the accompanying drawings, among which: There [ Fig 1 ] presents a simplified view of a system, cited by way of example, in which at least some embodiments of the process of the present application can be implemented, The [ Fig 2 ] presents a simplified view of a device adapted to implement at least certain embodiments of the process described in this application, The [ Fig 3 ] presents an overview of the process of this application, in some of its implementations. The [ Fig 4 ] presents an example of obtaining a movement chart, in certain embodiments of the process described in this application. The [ Fig 5 ] presents an example of obtaining histograms from a portion of a motion map, in certain embodiments of the method of this application. The [ Fig 6 ] presents an example of obtaining encompassing geometric shapes from the histograms of the figure 5 , in certain embodiments of the process described in this application. The [ Fig 7 ] presents an example of obtaining a new bounding geometric shape from histograms relating to one of the bounding geometric shapes of the figure 5 , in certain embodiments of the process described in this application. The [ Fig 8 ] presents an example of encompassing geometric shapes obtained from the movement map portion of the figure 5 , at the end of segmentation, according to certain embodiments of the process described in this application. The [ Fig 9 ] presents an example of obtaining a bounding geometric shape from the movement map of the figure 4 , in certain embodiments of the process described in this application. The [ Fig 10 ] presents an example of a line of interest from a movement chart, according to certain embodiments of the process in this application. The [ Fig 11 ] presents an example of application to pedestrian counting of the method of this application, in some of its embodiments. 5. Description des modes de réalisation
[0039] This application proposes a solution based on analyzing a video stream representing a scene (or area) to be monitored, in order to detect moving objects within that scene. Movement within the monitored scene will be reflected by differences between successive images of the video stream.
[0040] The differences between several successive images of the stream allow us to obtain a map representing the movements that occurred in the monitored scene between these images. This map is also referred to in this application as a "movement map".
[0041] The motion map is divided into segments indicating the presence or absence of movement between images in specific areas of the scene. In some embodiments, the method described in this application may involve grouping segments associated with moving areas (also called motion segments for simplicity) and associating them with objects, as well as, optionally, tracking these objects across the motion maps. It may be possible, for example, to count these objects, identify them, or generate statistics about them.
[0042] In the detailed examples below, the scene to be monitored is a highway scene with moving vehicles (cars, trucks, motorcycles, etc.). It is clear, however, that certain embodiments of the method described in this application can be implemented in other environments (particularly for other types of scenes to be monitored, such as a factory workshop, a shopping center, a living room, a shopping street, an animal crossing point, etc.) for the detection of various objects (robots, vehicles, pedestrians, animals, etc.).
[0043] We now describe, in connection with the figure 1 , in more detail the present request.
[0044] There figure 1 represents a telecommunications system 100 in which the method of the present application can be implemented, in at least some of its embodiments. The system 100 comprises one or more electronic devices, at least some of which can communicate with each other via one or more communication networks 120, possibly interconnected, such as a local area network (LAN) and / or a wide area network (WAN). For example, the network may include a corporate or home LAN and / or a WAN of the internet, or cellular, GSM (Global System for Mobile Communications), UMTS (Universal Mobile Telecommunications System), Wi-Fi (Wireless), etc. type.
[0045] As illustrated in figure 1 The system 100 may also include several electronic devices, such as a communication terminal (such as a computer 110, for example a laptop, a smartphone 120, a tablet 130), and / or a server 140, for example a server adapted to process (analyze for example) data obtained via at least one of the electronic devices of the system 1, a storage device 150 for example a storage device adapted to process (analyze for example) data obtained via at least one of the electronic devices of the system 1. The system may also include network management and / or interconnection elements (not shown).
[0046] There figure 2 illustrates a simplified structure of an electronic device 200 of system 100, for example device 110, 130 or 140 of the figure 1 adapted to implement the principles of this application. Depending on the embodiment, this may be a server and / or a terminal.
[0047] Device 200 includes, in particular, at least one memory M 210. Device 200 may include, in particular, a buffer memory, volatile memory, for example of the RAM (Random Access Memory) type, and / or non-volatile memory (for example of the ROM (Read Only Memory) type). Device 200 may also include a processing unit UT 220, equipped, for example, with at least one P 222 processor, and driven by a computer program PG 212 stored in memory M 210. At initialization, the code instructions of the computer program PG are, for example, loaded into RAM before being executed by the P processor. Said at least one P 222 processor of the processing unit UT 220 may, in particular, implement, individually or collectively, any one of the embodiments of the method of this application (described in particular in relation to the figure 3 ), according to the instructions of the PG computer program.
[0048] The device may also include, or be coupled to, at least one I / O input / output module 230, such as a communication module, enabling, for example, the device 200 to communicate with other devices in the system 100, via wired or wireless communication interfaces, and / or such as an acquisition module enabling the device to obtain (for example, acquire) data representative of an environment to be monitored, and / or such as an interface module with a user of the device (also referred to more simply in this application as a "user interface").
[0049] The device user interface refers, for example, to an interface integrated into the device 200, or to a part of a third-party device connected to this device by wired or wireless means. For example, it could be a secondary screen of the device or a set of speakers connected wirelessly to the device. A user interface can, in particular, be an "output" user interface adapted for rendering (or controlling the rendering) of an output element of a computer application used by the device 200, for example, an application running at least partially on the device 200 or an "online" application running at least partially remotely, for example, on the server 140 of the system 100.Examples of the device's output user interface include one or more screens, including at least one graphic display (e.g., touchscreen), one or more speakers, and connected headphones. The interface of device 200, for example, can be adapted to render the images illustrated by one of the... figures 4 , 9 à 11 . By rendering, we mean here a display (or “output” according to English terminology) on at least one user interface, in any form, for example including text, audio and / or video components, or a combination of such components.
[0050] Furthermore, a user interface can be an "input" user interface adapted for acquiring a command from a user of device 200. This may include an action to be performed in connection with a returned item, and / or a command to be transmitted to a computer application used by device 200, for example an application running at least partially on device 200 or an "online" application running at least partially remotely, for example on server 140 of system 100. Examples of input user interfaces for device 200 include one or more sensors, an audio and / or video acquisition means (microphone, camera (webcam) for example), a keyboard, a mouse.
[0051] As indicated above, the device may include or be coupled (via its communication means) to at least one acquisition module, such as a sensor (hardware and / or software) coupled to said device, such as a video camera, enabling the acquisition of a video stream representative of the scene to be monitored. Depending on the embodiments (and applications) of the method of this application, different video cameras may be used for capturing the scene: for example, it may be a wide-angle or narrow-angle video camera, a USB or IP camera (connected, for example, to the I / O module 230 of the device 200).
[0052] The device may also include or be coupled with other types of sensors, capable for example of reporting on a physical environment to be monitored and / or the situation of a capture module (for example a position and / or orientation of a camera monitoring the scene).
[0053] Information acquired via the acquisition and input / output modules can, for example, be transmitted to the processing module 210.
[0054] For example, said at least one processor of device 200 can be adapted, in particular, to: obtaining at least two time-spaced images capturing the same scene from a video stream; obtaining at least one motion map representing the differences between the components, according to at least one video channel, of said images; obtaining at least one geometric shape encompassing a set of moving portions of said motion map.
[0055] For example, said at least one processor of device 200 can be adapted, in particular, to: obtaining at least two time-spaced images capturing the same scene from a video stream; obtaining at least one first motion map cut into portions representative of the differences between the values of at least two video channels, for an area of at least two of said images; obtaining at least one geometric shape encompassing a set of moving portions of said first motion map.
[0056] Some of the above input / output modules are optional and may therefore be absent from device 200 in certain embodiments. In particular, while this application is sometimes detailed in relation to a device communicating with at least one other device of system 100, the method can also be implemented locally by device 200, without requiring communication with another device, in certain embodiments. For example, in some embodiments, the device can use information acquired by a camera internal to the device and operating continuously, and display in real time the resulting data from the processing of this information on a screen local to the device (for example, an alert about traffic congestion at a crossing point, constituting the scene to be monitored, on a road where the device is installed).
[0057] Thus, at least some of the embodiments of the process in this application offer a solution that can help to analyze a flow of vehicles locally, without requiring the sending of a video stream (capturing images of the flow of vehicles) to a remote server or to the cloud.
[0058] In some of its embodiments, on the contrary, the process can be implemented in a distributed manner between at least two devices 110, 120, 130, and / or 150 of the system 100. For example, device 200 can be installed in a centralized control station and obtain, via a communication module of the device, information acquired by one or more cameras installed on crossing points to be monitored on a road axis and belonging to the same local network as device 200.
[0059] The terms "module," "component," or "element" of the device refer to a hardware element, particularly a wired one, a software element, or a combination of at least one hardware element and at least one software element. The method according to the invention can therefore be implemented in various ways, including in wired and / or software form.
[0060] There figure 3 illustrates certain embodiments of process 300 of this application. Process 300 can, for example, be implemented by the electronic device 200 illustrated in figure 2 .
[0061] As illustrated in figure 3 , the process 300 can include obtaining 310 a plurality N of successive images of a video stream capturing a scene to be monitored (i.e. capturing images of the scene) (with N: an integer greater than or equal to 2, for example with a value of 2, 3 or 5).
[0062] For example, this could refer to the N most recent images of a video stream being acquired. In this application, obtaining an element means, for example, receiving that element from a communication network, acquiring that element (via, for example, user interface elements or sensors), creating that element through various processing methods such as copying, encoding, decoding, transformation, etc., and / or accessing that element from a local or remote storage medium accessible to the device implementing this acquisition.
[0063] Obtaining images (item 310) may, for example, include receiving at least a portion of the video stream on a device's communication interface and / or reading at least a portion of the video stream from a local or remote storage area (e.g., a database as illustrated by item 150 of the [document / method / etc.]). figure 1 ), or read access to a buffer accessed for writing by at least one camera monitoring the scene.
[0064] In some embodiments, the video stream can be acquired by a single camera, in an unchanged position, orientation and / or shooting angle throughout the acquisition.
[0065] Each of the N images can have several components, for example three components R, G, and B, corresponding to the video channels (red, green, blue) (or red, green, blue (RGB) according to English terminology).
[0066] In the remainder of this document, we note respectively imgt, img t- 1 and img t- 2 the image at a time t, for example the current time of the video stream and the two previous images.
[0067] Furthermore, we designate respectively: by imgt(R), img t ( V ) And imgt(B) the red, green and blue components of the image img t , by img t- 1 ( R ), img t- 1 ( V ) And img t- 1 ( B ) the red, green and blue components of the image img t -1 and by img t -2 (R), img t- 2 (V) And img t- 2 (B) the red, green and blue components of the image img t -2.
[0068] Optionally, image acquisition 310 may also include the acquisition of additional data (or metadata) associated with that image, such as a date and / or time of capture, an identifier of the camera that acquired the image, an identifier of the image in the acquired video stream, information relating to the positioning and / or capture of a camera that acquired at least a portion of the video stream, etc.
[0069] As illustrated in figure 3 as well as in figure 4 The process may also include obtaining (for example, generating) 320 a map (called a motion map) representative of the movements that occurred in the monitored scene between the N images obtained. We note mvt t the motion map obtained from the image imgt, and the previous N-1 images. The figure 4 This illustrates an example of a 420 motion map obtained from 3 successive images. img t -2,410 img t -1,412 and img t 414, (the three most recent images of the video stream, for example). In the illustrated example, a portion of the motion map 420 corresponds to one pixel of the obtained images.
[0070] In some embodiments, the motion map can be divided into disjoint portions (forming a partition of the map), each portion being associated with a distinct area of the images of the monitored scene, so as to partition the obtained images (and therefore partition the scene to be monitored). Depending on the embodiment, these portions may be of identical size and / or shape throughout the map or, conversely, of variable size and / or shape. For example, in some embodiments, each portion (pixel, for example) of the motion map may represent P pixels in the images obtained, with P a constant natural number (i.e., it may be a 1-P correspondence (P>= 1).
[0071] In some embodiments, portions of the motion map can correspond to areas of different sizes in the resulting images. In particular, in some embodiments, two identically sized portions of the same motion map can correspond to image areas of different sizes, the size of the image area assigned to a portion of the motion map depending, for example, on the probability of motion occurring within that image area. Thus, when the scene to be monitored includes a crossing point on a road, the size of an image area representing a portion of the road (i.e., the roadway), where cars frequently travel, can be smaller than the size of an area representing the roadside, which generally remains unoccupied.
[0072] In some embodiments, the motion map can be a binary map, where all portions of the map corresponding to an image area where motion has been detected are assigned a first and same constant value (for example 1), all portions of the motion map corresponding to an image area where no motion has been detected are assigned a second and same constant value (for example 0).
[0073] In some embodiments, obtaining the binary map can take into account the differences between the same components (for example, according to at least one of the same R, G, or B video channels) of at least some of the N images obtained. For example, in some embodiments, it can take into account the differences between the components of each R, G, and B video channel of at least some of the N images obtained, for example, of the set of all N images obtained. In other embodiments, it can take into account the differences between the components of a single R, G, or B video channel of the set of all N images obtained.
[0074] In some embodiments, the differences between each video component of the images can be evaluated for each portion, so as to account for the existence of a significant difference in at least one of the video components for each portion. Such an embodiment can detect more visible differences in a first video component of the images for a first portion of the motion map and more visible differences in a second video component of the images for a second portion of the motion map (depending, for example, on the colors of the moving objects in the obtained images). For example, in an embodiment where the 3 RGB video components of 3 successive images (N=3) are taken into account and where the motion map is a binary map where a pixel mvt t ( i,j ) of the movement map corresponds to a pixel with abscissa i and ordinate jin a component of a obtained image ( img t ( R,i,j) representing, for example, the value of the R component of the pixel with coordinates (i, j) of img t ), the pixel value mut t ( i,j The movement map can be obtained by applying the following formula. In this formula, T is a strictly positive natural number that represents an amplitude of variation (for example, a minimum variation). mvt t i j = 1 si: max img t R i j − 2 img t − 1 R i j + img t − 2 R i j ; img t V i j − 2 img t − 1 Vi j + img t − 2 V i j ; img t B i j − 2 img t − 1 B i j + img t − 2 B i j > T mvt t i j = 0 sinon
[0075] In one variant, the value of a portion of the motion map can be obtained by calculating the difference between the R, G, or B values of each component for the pixels of the images corresponding to the considered portion of the motion map. For example, in some embodiments where the 3 RGB video components of 3 successive images (N=3) are taken into account and where the motion map is a binary map, the value of a portion of a motion map can be obtained by calculating the difference between the average R, G, or B values of each component for the pixels of the images corresponding to the considered portion.
[0076] In yet another variant, the motion map can be chosen from among the motion maps of the 3 video components, as the motion map (among these 3 maps) with the most differences.
[0077] The value of T can differ depending on the embodiment. For example, the value of T can be a function of a desired sensitivity (which can be defined via a user interface of the device and / or via read access to a configuration file) and / or of scene capture characteristics. For example, it can be a function of the amount of light received by the camera capturing the scene.
[0078] In some embodiments, T can have a constant value (on the order of a few tens, for example 25). In the example illustrated in figure 3 , the detection process further includes a 330 segmentation of the previously generated mvt t movement map, taking into account at least one moving object in the map.
[0079] This segmentation aims, in certain embodiments, to locate on the motion map geometric shapes (for example, the smallest possible) encompassing at least one set of moving portions (i.e., a set of white pixels in the example of the figure 4 ) and potentially associated with the same object.
[0080] The motion map is segmented along vertical and horizontal axes, forming rows and columns of portions (squares or rectangles), and a set of portions associated with the same object is a set of squares or rectangles. Segmentation 330 aims to generate 336 rectangular shapes, each encompassing a set of portions. Of course, in other embodiments where the motion map is segmented differently (for example, into arcs of circles), the geometric shapes resulting from the segmentation can be shapes other than rectangles. In the examples illustrated in figures 5 à 8 which represent a part of the movement map, it can be obtained 336 of the encompassing rectangular geometric shapes 540, 550, 560, 570 corresponding to the intersections of the columns and rows each comprising at least one moving portion.
[0081] This 336 result is based on histograms (as described in more detail below).
[0082] Thus, as illustrated by the figures 3 And 5 à 8 The 330 segmentation includes: a 332 acquisition of a first histogram 510, along the vertical axis of the map mvt t , counting the number of portions to which a movement is associated (for example, the white pixels in the example of the figure 4 ) for each column of the movement chart. This histogram 510 can, for example, take the form, in the illustrated example, of a first table (or list) whose length (i.e., the number of elements) is identical to the number of columns in the movement chart. This results in a second histogram 520, obtained along the horizontal axis of the chart. mut t , counting the number of portions to which a movement is associated (for example, the white pixels in the example of the figure 4 ) for each row of the movement chart. This histogram 520 can, for example, take the form, in the illustrated example, of a second table (or list) whose length is identical to the number of rows of the movement chart.
[0083] In some embodiments, the two histograms can be obtained simultaneously. For example, as illustrated in figure 5 The process may include initializing the first 510 and second 520 arrays to 0, then traversing the map mvt t 420, to increment for each portion p i 530 in movement of the map (i.e., each white pixel (with a value of 1, for example) according to the example of the figure 4 ) the value of cells 512, 522 of the first and second tables corresponding respectively to the column and row number of the portion p i 530 in the movement map. obtaining 336 encompassing geometric shapes (rectangular in the illustrated example) 540, 570, 560, 570, on the movement map, in the intersections of columns and rows where the histogram values are non-zero.
[0084] There figure 5 This illustrates the encompassing geometric shapes 540, 550, 560, 570 obtained by considering the first 510 and second 520 histograms. The periphery (boundary) of at least one encompassing geometric shape can be obtained by identifying the row and column numbers corresponding to the cells in histograms 510, 520 where a transition is observed between a zero value and a non-zero value, or vice versa.
[0085] For example, by traversing the first histogram from left to right and the second histogram from top to bottom, a transition from 0 to a non-0 value in both histograms may correspond to the upper left corner 542, 552, 562, 572 of an encompassing geometric shape 540, 550, 560, 570, and a transition in both histograms from a non-0 value to a value equal to 0 may correspond to the lower right corner 544, 564, 574 of an encompassing geometric shape 540, 550, 560, 570.
[0086] In some embodiments, as illustrated in figure 6 The process may also include the removal of 338 of the encompassing geometric shapes 550 (called "empty") that do not include any moving portions (i.e., containing only black pixels in the example of the figure 4 ), so as to eliminate the "ghost" encompassing geometric shapes that appeared due to the "cumulative" histogram approach and to retain only the encompassing geometric shapes that include at least one moving portion. As illustrated in figure 6 Some of the encompassing geometric shapes 540, 560, retained after the removal of empty encompassing geometric shapes, thus including at least one moving portion (and therefore corresponding to a "real" object moving in the captured scene), may still "extend" beyond the real objects (i.e., contain portions not corresponding to the object's movement). Also, in some embodiments, as illustrated in figure 7 The process may include at least one iteration of obtaining histograms and bounding geometric shapes, not from the motion map as described above, but from at least one bounding geometric shape already obtained in a previous iteration (for example, a non-"empty" bounding geometric shape). This at least one iteration may possibly allow (as illustrated in figure 7 ) to obtain at least one other encompassing geometric shape on which, for example, a new iteration can be performed. During an iteration, the first table 550 and second table 560, representing the histograms of an encompassing geometric shape 540, then have sizes (i.e., number of elements) corresponding to one of the dimensions (width, length) of the encompassing geometric shape on which the iteration is performed, as illustrated in figure 7 Using smaller histogram sizes can help to obtain at least one encompassing geometric shape smaller than the one from which it was derived, and optionally to identify and remove, among these newly obtained shapes, new "empty" shapes, as illustrated in figure 8 Therefore, in some embodiments, iterations can help to more finely delineate the moving parts of an object represented by an encompassing geometric shape.
[0087] A criterion for stopping the iterations could be, for example, the absence of generation of a new encompassing geometric shape during the last M iterations, where M is a strictly positive integer (for example, equal to 1), as illustrated in figure 8 In such an embodiment, the final encompassing geometric shapes obtained no longer contain any empty rows or columns, as illustrated in figure 8 ).
[0088] In the example illustrated in figures 3 à 8 The motion map is cut along horizontal and vertical axes. However, in some embodiments, other cutting axes may be used. For example, orthogonal axes inclined at a first angle to the vertical may be used. Such a variant may be useful, for instance, when at least some objects on the motion map are not separable along the horizontal and vertical axes. In such an embodiment, process 300 may include an optional step (not illustrated in figure 3 ) transformation of the movement map (e.g., a rotation around the first angle) before the implementation of the segmentation 330 described above.
[0089] As illustrated in figure 3 In some embodiments, process 300 may include tracking 340 of at least one moving object in the monitored scene. This tracking may be optional in some embodiments.
[0090] At any given moment, an object is only partially visible to the camera capturing the monitored scene (for example, some parts of the object may be obscured by the object itself from the camera's view). Furthermore, as the object moves, different parts of the object become visible to the camera (and therefore in the resulting images) over time. Consequently, the set of portions associated with the same object in motion maps can vary in terms of shape, size, and / or position across images (and therefore between different motion maps). Using bounding geometric shapes for tracking, which are less variable than the objects themselves, can help to smooth out the tracking. In some embodiments, the tracking can also take into account the ratios between the dimensions of the bounding shapes (for example, a ratio between the width and length of a shape).Indeed, using such a dimensionality ratio can help, for example, to achieve tracking that is less sensitive to apparent size variations (in images) of an object following a change in its distance from the camera capturing the scene. This helps maintain the same object identifier across multiple bounding geometric shapes from different motion maps corresponding to the same object whose distance from the camera varies. 340° tracking can, for instance, be performed on at least one first bounding geometric shape resulting from the 330° segmentation of a first motion map, in conjunction with at least one second motion map that precedes the first. 340° tracking can also be performed on all bounding geometric shapes resulting from the 330° segmentation between at least two motion maps generated at different times.The aim is twofold: first, to assign a bounding geometric shape a different identifier than other bounding geometric shapes moving within the movement map to which it belongs; and second, to attempt to track the objects bounded by these geometric shapes across different movement maps (for example, consecutive movement maps). To achieve this, the same identifier can be assigned to the bounding geometric shapes in successive movement maps that appear to correspond to the same object.
[0091] We will detail below a more specific example of tracking between two movement charts mvt t- 1 and mvt t successive in time.
[0092] In some embodiments, the tracking 340 may include obtaining 342 relative positioning information for the encompassing geometric shapes of the movement map mvt t and encompassing geometric shapes of the movement map mvt t -1 . For example, the process may include obtaining 342 distances (e.g., Euclidean) between the centroids of the encompassing geometric shapes of the motion map mvt t and the movement map mvt t -1 .
[0093] The term "centroid" here refers to the center of a bounding geometric shape. The Euclidean distance d between two bounding geometric shapes whose centroids have respective coordinates x 1, y 1 and x 2 ,y 2 can be obtained using the equation: d = x 1 − x 2 2 + y 1 − y 2 2 .
[0094] As illustrated in figure 3 , the process may include, for at least a first encompassing geometric shape of the movement map mvt t , a conditional association 344 of this first encompassing geometric shape with a second encompassing geometric shape of the movement map mvt t -1, taking into account the relative positioning information obtained and an assignment of an object identifier to the first encompassing geometric shape, taking into account this conditional association. The second encompassing geometric shape is, for example, the encompassing geometric shape of the movement map. mvt t- 1. The closest (according to the relative positioning information obtained) to the first encompassing geometric shape of the movement map mvt t .
[0095] In some embodiments where the relative positioning information of two encompassing geometric shapes is a distance between these two encompassing geometric shapes, the association of the first encompassing geometric shape with the second encompassing geometric shape can be effective when the distance between the first encompassing geometric shape of the movement map mvt t and the second encompassing geometric shape of the movement map mvt t- 1 is less than a first distance (used as a threshold). In some embodiments, the process may include assigning to the first encompassing geometric shape the object identifier already assigned (previously) to the second encompassing geometric shape. Indeed, the first and second encompassing geometric shapes are considered to correspond to the same object.
[0096] When the distance between the first encompassing geometric shape of the movement map mvt t and the second encompassing geometric form of movements mvt t -1 is greater than the first distance, the first encompassing geometric shape is considered to correspond to a different object than the second encompassing geometric shape. When the distance between the first encompassing geometric shape of the movement map mvt t and the set of encompassing geometric shapes of the movement map mvt t-1 is greater than the first distance, the first encompassing geometric shape is considered to correspond to a newly appeared or newly moving object in the scene and the process includes an assignment 346 of a new object identifier (not assigned at time t to another encompassing geometric shape of the same or another movement map) to the first encompassing geometric shape.
[0097] New object identifiers can, for example, be generated in such a way as to guarantee their uniqueness with respect to other identifiers already assigned at a given time. For example, these could be object identifiers each comprising at least one numeric portion obtained by successively incrementing the same numeric variable.
[0098] In some embodiments, the process may further include a storage 348 (in a storage area accessible to the device) of the newly assigned object identifier (for example in a field of a data structure representing the first encompassing geometric shape).
[0099] Depending on the embodiment, the first distance used as a threshold can be a constant distance or a variable distance, for example adjustable by configuration (so as to depend, for example, on a type of object to be tracked). Indeed, the distance that the same object is likely to travel between two successive images obtained (and therefore between two successive motion maps) can vary greatly depending on the object (for example, depending on whether it is a car, a pedestrian, etc.).
[0100] In some embodiments, the method may include filtering the tracked bounding geometric shapes. This filtering may, for example, be based on the type of object targeted by the tracking, so as to exclude from the tracking bounding geometric shapes that do not a priori correspond to a targeted object type, and / or conversely, to limit the tracking to shapes that a priori correspond to a targeted object type. For example, in the illustrated embodiment, the filtering may limit the tracking to bounding geometric shapes identified as vehicles.
[0101] The identification of encompassing geometric shapes to exclude and / or to which tracking can be limited can, for example, be carried out as the process unfolds across all the encompassing geometric shapes tracked 340 in the movement maps.
[0102] The identification of encompassing geometric shapes to be excluded and / or to which tracking should be limited can, for example, be implemented by taking into account the size of the encompassing geometric shapes in relation to an expected (probable) size of the objects in question, or a ratio between the width and / or length of the encompassing geometric shapes.
[0103] In the illustrated embodiments, where vehicles are targeted (truck, car, motorcycle, ...), encompassing geometric shapes that are too small to correspond to a vehicle (and considered as extraneous objects - like a pedestrian or a bird) can, for example, be excluded from tracking.
[0104] Furthermore, in some embodiments, an indication of the type of object targeted (truck, car, motorcycle, etc.) to which a bounding geometric shape corresponds can be deduced from the width, length, and / or width-to-length ratio of the bounding geometric shape. This could be achieved, for example, by comparing it to a lookup table or a library of bounding geometric shapes classified by object type.
[0105] Alternatively, in some embodiments, the process can implement "frugal" artificial intelligence models to identify a target object type (e.g., a vehicle type) based on the width and length of the bounding geometries. These models could, for example, be neural networks with a low number of neurons per layer and whose parameters have been quantized to maintain frugality in terms of computing resource consumption (memory and / or processing complexity). The inference of such a neural network could, for example, be implemented locally on the device.
[0106] Depending on the embodiment, various 360° processing techniques can be applied to the tracked (and optionally filtered) bounding geometric shapes. For example, in some embodiments, the process may include a 360° count of the filtered bounding geometric shapes. In the illustrated embodiments, bounding geometric shapes corresponding to vehicles can, for example, be counted. The processing techniques applied may take into account, for example, at least one part, referred to as the "part of interest," considered relevant to a movement map. In some embodiments, the "part of interest" may be a one-dimensional geometric shape (such as a line or curve). It may be, for example, a boundary (called the "boundary of interest") that a moving object can cross or an area. In the example illustrated in figure 10 An encompassing geometric shape, considered equivalent to a vehicle, can be counted when it crosses a line of interest (428) on a movement chart (420). In some embodiments, the "area of interest" can be a two-dimensional geometric shape (such as a circle, triangle, square, rectangle, rhombus, etc.). It can, for example, be a surface or area (called an area of interest) into which a moving object can enter or exit. For instance, an encompassing geometric shape may only be counted when its centroid lies within an area of interest on a movement chart.
[0107] In some embodiments, several parts of interest (of the same or different dimensions) relating to the captured scene can be defined in the motion maps.
[0108] In some embodiments, the process can take into account (for 360 counting, for example) a direction of movement of the encompassing geometric shapes. In the example illustrated in figure 10 Thus, bounding geometric shapes treated as vehicles can only be counted if they cross the line of interest 428 from area 427 above line 428 of the movement map 420 to area 429 below line 428 of the movement map 420. The direction of movement of an object can, for example, be obtained (during the tracking of bounding geometric shapes, for instance) from the successive coordinates of the centroids of the bounding geometric shapes, to which the object's identifier is assigned in the movement maps. The centroid coordinates of a bounding geometric shape can, for example, be stored in the data structure representing the bounding geometric shape.
[0109] In some embodiments, the method may include a 370 rendering of at least one motion map on a user interface of said device 200. The rendering may further include, in some embodiments, at least one encompassing geometric shape of the motion map and optionally an object identifier to which it is associated. It may also optionally include at least one part of interest of the motion map.
[0110] At least some embodiments of the method of the present application may use only a small number N of images (for example 3) to construct a motion map and be implemented for applications where only a few objects are tracked in parallel, or target only a few types of objects, and may therefore be less memory-intensive than some prior art solutions.
[0111] In at least some embodiments, the resulting images may be low-resolution. Furthermore, the operations performed on these images may be simple algebraic operations available on all types of common processors and microcontrollers. Also, at least some embodiments of the method described in this application may be less demanding in terms of processing power and / or energy than some prior art solutions. For example, the implementation of at least some embodiments of the method described in this application does not require specific processors such as those specialized in graphics processing (e.g., GPUs) or in neural network inference (e.g., TPUs).
[0112] Consequently, certain embodiments of the process described in this application can be implemented on lightweight devices (such as Raspberry Pi 3), while still enabling the processing of 640 x 480 (or 480 x 640) pixel video at approximately 12 frames per second (whereas some prior art solutions, for example, can only process 416 x 416 video at 2 or 3 frames per second on this type of device). Certain embodiments of this application can thus provide a solution for identifying and counting objects with lower computational costs than some prior art solutions, and requiring little energy. Such embodiments can therefore help resolve the recurring problems of installing these counting systems (by avoiding costly construction work) and using them (in terms of energy cost and computing power).
[0113] Some examples of implementation of the method described in this application have been detailed above. However, the method can find numerous applications for identifying and counting moving objects in various fields, without requiring extensive infrastructure work or significant processing or communication capacities. In certain embodiments, the method described in this application can be adapted for the detection, tracking, identification, and / or counting of vehicles at very low cost, and can, for example, be implemented by a mobile device that is easily transportable as needed.
[0114] For example, the method can be applied, in at least some of its embodiments, in challenging environments (isolated or with limited infrastructure or computing resources) (as is sometimes the case in developing countries). The method can be applied in particular to roads, highways, the areas surrounding warehouses or industrial sites where it is necessary to count vehicles entering and leaving a zone. The method can also be applied, in certain embodiments, to the identification and counting of bicycles, animals, or pedestrians, as shown in the example below. figure 11(which illustrates pedestrian identification from images obtained from a top view of a scene) In some embodiments, the process can be implemented without using expensive artificial intelligence models (in terms of computing power, storage capacity and training time), and / or sophisticated remote systems that consume network resources.
Claims
1. Method for detecting motion, implemented in an electronic device, said method comprising - obtaining (310) at least two frames that are spaced apart temporally and that capture the same scene of a video stream; - obtaining (320) at least a first motion map divided into segments representative of differences between the values of at least two video channels for a region of at least two of said frames, said segments dividing said motion map into rows and columns along two different division axes, - obtaining a first histogram relating to said segments, the first histogram counting the number of moving segments per column in said first motion map; - obtaining a second histogram relating to said segments, said second histogram counting the number of moving segments per row in said first motion map; - obtaining at least one geometric shape bounding a set of moving segments of said first motion map and corresponding to an intersection in the motion map of columns comprising at least one non-zero element of the first histogram and of rows comprising at least one non-zero element of the second histogram.
2. Electronic device comprising at least one processor, said at least one processor being configured to: - obtain (310) at least two frames that are spaced apart temporally and that capture the same scene of a video stream; - obtain (320) at least a first motion map divided into segments representative of differences between the values of at least two video channels, for a region of at least two of said frames, said segments dividing said motion map into rows and columns along two different division axes, - obtain a first histogram relating to said segments, the first histogram counting the number of moving segments per column in said first motion map; - obtain a second histogram relating to said segments, said second histogram counting the number of moving segments per row in said first motion map; - obtain at least one geometric shape bounding a set of moving segments of said first motion map and corresponding to an intersection in the motion map of columns comprising at least one non-zero element of the first histogram and of rows comprising at least one non-zero element of the second histogram.
3. Method according to Claim 1 or electronic device according to Claim 2, wherein the value of a segment of said first motion map is a binary value reflecting differences between the values of each video channel for said region of said frames.
4. Method according to Claim 3, or electronic device according to Claim 3, wherein said value of a segment of said first motion map reflects the maximum of the differences of each video channel for said region of said frames.
5. Method according to Claim 1 or 3 or 4, wherein the method comprises obtaining one motion map for each video channel of said frames, or an electronic device according to any of Claims 2 to 4, wherein said at least one processor is configured to obtain a motion map for each video channel of said frames, and wherein - the motion map of a video channel of said frames is divided into segments representative of differences for a region of at least two of said frames, for the video channel associated with said motion map, - said first motion map is the motion map, among the motion maps associated with each of the video channels of said frames, that comprises the most segments of which the binary value is representative of a difference for a region of said frames, for the associated video channel.
6. Method according to any of Claims 1 or 3 to 5, or electronic device according to any of Claims 2 to 5, wherein said geometric shape is a rectangular shape.
7. Method according to any of Claims 1 or 3 to 6 wherein said method is implemented iteratively on successive frames of the video stream and comprises: obtaining a first motion map and a second motion map temporally preceding said first motion map, determining a conditional association between a first bounding geometric shape obtained for the first motion map and at least one second bounding geometric shape obtained for at least said second motion map, said association reflecting relative positions of said first and second bounding shapes in said first and second motion maps.
8. Method according to Claim 6 comprising assigning the same object identifier to said associated first and second geometric shapes.
9. Method according to any of Claims 1 or 3 to 8 comprising, or electronic device according to any of Claims 2 to 6 wherein said at least one processor is configured to assign an object type to at least one geometric shape, said assignment reflecting a size, a width and / or a length of said geometric shape.
10. Method according to any of Claims 1 or 3 to 9 comprising, or electronic device according to any of Claims 2 to 6 or 9, wherein said at least one processor is configured to - obtain a third histogram counting the number of moving segments per column in at least a third of said obtained bounding shapes; - obtain a fourth histogram counting the number of moving segments per row in said third obtained bounding shape; - obtain at least one fourth bounding geometric shape corresponding to the intersection in said first motion map of columns comprising at least one non-zero element of the third histogram and of rows comprising at least one non-zero element of the fourth histogram.
11. Method according to Claim 10, or electronic device according to Claim 10, wherein said third and fourth histograms are obtained successively using at least one bounding geometric shape obtained beforehand from said first motion map.
12. Method according to Claim 11, wherein the method comprises, or electronic device according to Claim 11 wherein said at least one processor is configured to achieve, prior to dividing said first motion map, a rotation of said first motion map.
13. Computer program comprising instructions for implementing, when the program is executed by a processor of an electronic device, the method for detecting motion according to Claim 1.
14. Recording medium readable by a processor and on which is recorded a computer program comprising instructions for implementing, when the program is executed by a processor of an electronic device, the method for detecting motion according to Claim 1.