PREVENTION OF SHIFT AND TOPOLOGY FOR A CAMERA NETWORK

DE602018087579T2Active Publication Date: 2025-12-03BULL SA
View PDF 0 Cites 0 Cited by

Patent Information

Application Number
DE602018087579
Authority / Receiving Office
DE · DE
Patent Type
Patents
Current Assignee / Owner
Priority Date
2017-12-29
Filing Date
2018-12-28
Publication Date
2025-12-03
Estimated Expiration
2038-12-28

AI Technical Summary

Technical Problem

Existing real-time and remote monitoring systems using video surveillance cameras face inefficiencies in target tracking due to large camera networks, camera malfunctions, and uncovered areas, leading to lengthy searches and uncertainty about target reappearances.

Method used

A real-time surveillance system utilizing an artificial predictive neural network that learns camera topology and predicts the next likely camera for target appearance by correlating target signatures across multiple cameras, enabling incremental learning and improved accuracy.

Benefits of technology

Enhances target tracking efficiency by predicting the most probable camera for target re-appearance, allowing quicker interception and providing valuable data for urban planning and traffic management.

✦ Generated by Eureka AI based on patent content.
Patent Text Reader
Need to check novelty before this filing date? Find Prior Art

Description

TECHNICAL FIELD

[0001] The present invention relates to the field of real-time site monitoring using video surveillance cameras. A site monitored by video surveillance cameras is understood to mean a site such as a city or neighborhood, or even a museum, stadium, or building, where a set of video surveillance cameras films areas of the site. This technology can be adapted to a surveillance system using multiple cameras with overlapping or non-overlapping fields of view, that is, including areas within the site outside the field of view of the video surveillance cameras.

[0002] The invention relates more particularly to a real-time site monitoring system, especially for cities or neighborhoods, accessible remotely via a communication network. It should be noted that such a communication network preferentially, but not exclusively, refers to an intranet or internet computer network. STATE OF THE ART

[0003] The state of the art already includes real-time and remote monitoring systems.

[0004] It is also known in particular for its CCTV camera maps of the site and detectors that can detect moving targets, as well as appearance pattern providers that can give a signature to a target per camera and record a list of target characteristics.

[0005] Specifically, the appearance model provider is known to perform target recognition using an artificial neural network that has learned to recognize a target and assign it a signature per camera, along with parameters also known as attributes. This target recognition allows the operator to initiate a search across all CCTV cameras to identify the target when it leaves a surveillance zone.

[0006] Such systems are disclosed for example in documents EP 2 911 388 A1 and WO 2008 / 100359 A1.

[0007] The drawback is that this search can be lengthy in the case of a large number of cameras, but also because the target must have already appeared on one of the CCTV cameras.

[0008] The drawback of the camera map also lies in the camera topology. Indeed, a camera may malfunction, or if an area of ​​the site is not covered by surveillance cameras, it can be difficult to know, if the target moves into this uncovered area, which camera the target might reappear on.

[0009] There is a need for operators to have a more efficient system to be able to track a target in order to be able to intercept the target, for example. SUMMARY OF THE INVENTION

[0010] The present invention aims to overcome the drawbacks of the prior art by proposing real-time monitoring of at least one target by means of an artificial predictive neural network learning camera topology and statistically the cameras that can probably identify the target when leaving an area filmed by a surveillance camera.

[0011] To achieve this, the invention relates to a real-time surveillance system comprising video surveillance cameras, the surveillance system being defined by claim 1.

[0012] Thus, the operator can determine the next likely camera to display. The predictive artificial neural network can learn the camera topology and suggest a camera or a list of cameras whose target is most likely to appear.

[0013] Unlike a support vector machine, better known by the acronym SVM for Support Vector Machine, the artificial predictive neural network allows for simplicity, and real-time learning allows for greater accuracy and speed, especially if there are many classes.

[0014] For example, the artificial predictive neural network can learn continuously or from time to time by receiving targets' positions and CCTV camera identifications filming those targets and then receiving the identification of the camera on which the target was detected, thus allowing the artificial predictive neural network incremental learning of the camera network topology.

[0015] This allows the artificial neural network for prediction to learn and thus be able to determine, based on a target's position, the next likely camera in which the target will be detected.

[0016] As an example, the camera management module receives a target data set from several tracking software programs and is able to correlate, via signatures, the movement of a target from one camera's field of view to another camera's field of view in order to transmit this data to the artificial predictive neural network for incremental learning.

[0017] The target can have a signature from each camera. The signature from one camera and the signature from another camera of the same individual will be close enough to each other that the camera management module can correlate the two cameras and thus recognize the target.

[0018] In other words, because of the signatures for each target in an image from the first camera, and then other signatures in another image from a second camera, and so on, the camera management module can correlate two similar signatures to identify the same target and thus the path taken by the target between the first and second cameras. Therefore, the module can use this data to perform incremental training on the artificial neural network by sending it at least: The identification of the first camera, the position at the exit state or the last position of the target in the first camera, the identification of the second camera.

[0019] To improve the probability, the module can also send: The class of the target and / or, the direction of the target and / or the speed of the target.

[0020] The surveillance camera management module also receives direction and sense data from the target and in that the target data for prediction sent to the artificial neural network for prediction also includes this direction data in the output state.

[0021] This improves the probability that the most likely camera identified by the artificial predictive neural network will identify the target within its field of view.

[0022] According to one embodiment, the surveillance camera management module further receives target speed data and in that the target data for prediction sent to the artificial neural network for prediction further include this direction data in the output state.

[0023] This improves the probability that the most likely camera identified by the artificial predictive neural network will identify the target within its field of view.

[0024] According to one embodiment, the surveillance camera management module further receives target class data and in that the target data for prediction sent to the artificial neural network for prediction further include this class data in the output state, the target class being able to be a pedestrian, a two-wheeler, a four-wheeler or other.

[0025] According to one example of this embodiment, the class of the four-wheeled vehicle is sub-classified as a car or a truck.

[0026] Thus, for a site such as a neighborhood or a road, teaching the artificial predictive neural network the probable paths per target effectively improves the probability that the most probable camera identified by the artificial predictive neural network identifies the target in its field of vision.

[0027] According to one embodiment, the state of the target in the area monitored by the camera may further be an active state in which the target is in an area of ​​the field of vision surrounded by the exit area.

[0028] In other words, the target, after being in the new state, goes into the active state as soon as the tracking software has received the target's signature from the appearance model provider, and as soon as the target enters an area of ​​the camera's field of view that may allow the target to be about to exit the camera's field of view, the target goes into the exit state.

[0029] Of course, tracking software can add exit zones in the middle of the camera's field of view when it detects a target's disappearance zone. This is particularly relevant, for example, when there is a tunnel entrance in the middle of a road.

[0030] This allows the operator to gather statistics on the target's likely route. This makes it possible to dispatch personnel, such as the police, to intercept the target more quickly and thus avoid a chase.

[0031] According to one embodiment, the surveillance camera management module is capable of receiving a request to predict a target, and in that the artificial neural network for prediction can send at its output a list of camera identifications by probability and in that the surveillance camera management module is further capable of sending to the interface machine the identification of the camera whose target is in the active state and a list ordered by probability of the possible probable camera identifications.

[0032] The prediction request can be made by an operator selecting a video from a camera displayed on a screen.

[0033] The request can also be made by sending an image of the target to the camera management module, which can query the appearance model provider directly to receive one or more probable signatures of the target filmed by one or more cameras.

[0034] In addition, the query can also be an attribute query, for example class = pedestrian, top clothing colour = red, hair colour etc... and thus select a number of probable targets filmed by the cameras.

[0035] According to one embodiment, the surveillance camera management module is capable of adding to an area of ​​the video whose target is identified in the active state or in the output state, a video filmed by a probable camera identified by the artificial predictive neural network.

[0036] According to one embodiment, the surveillance camera management module is capable of sending a list of surveillance camera nodes as well as the target path from the first camera that identified the target to the most probable camera of the target and for example another characteristic such as the class, the weight of the paths between the cameras.

[0037] For example, the weight can be calculated per class per camera.

[0038] This allows us to identify the most frequently used routes for a given class. This can help a city's urban planning department, for example, to determine which roads are the busiest and therefore most likely to require maintenance. Furthermore, it can also provide real-time traffic information.

[0039] In one embodiment, the management module and the predictive neural network are located on the same computer or separately connected via a public or private network, or on the same server or separate servers. The artificial neural network used is, for example, a multilayer perceptron for classification purposes, such as MLPCclassifier® from the Scikit-learn® library. The predictive neural network can be a deep learning network.

[0040] For example, the artificial neural network includes an input layer comprising neurons with an activation function of the rectified linear unit type also called "ReLU" (acronym for Rectified Linear Unit), including at least one position neuron and one camera identification neuron, a hidden layer comprising neurons with an activation function, for example, of the rectified linear unit type, and an output layer comprising at least one neuron with a loss layer activation function, also called Softmax, to predict the surveillance camera.

[0041] The input layer can include, for example, nine activation neurons. The input layer can include, for example, one neuron concerning the class of the target (pedestrian, car, truck, bus, animals, bicycle, etc.), four neurons concerning the position of the target (for example, two positions of the target in the image along two axes and two positions of the surveillance camera if it is mobile), two neurons concerning the direction of the target in the image along the two axes, one neuron for the speed of the target in the image and one neuron for the identification of the camera, i.e., nine variables.

[0042] The hidden layer may include, for example, one hundred neurons having the activation function of a rectified linear unit type.

[0043] The output layer can include one neuron per probability of the identified camera having the highest probability of the target appearing. For example, the output layer includes five output neurons for five probabilities of the five identified surveillance cameras having the highest probability of the target appearing.

[0044] In one embodiment, the system further comprises a set of target tracking software for surveillance cameras. The target tracking software may be installed on one or more computing devices, such as a computer.

[0045] The target tracking software is capable of following the target and deducing the target's direction and speed, identifying the target's state, and sending target data, including direction, speed, position, target state, and target signature, to the surveillance camera management module.

[0046] According to one example of this embodiment, the system further includes a detector that extracts a target image from the video, performs image recognition to assign the target class, and sends the extracted image and its class to the tracking software suite. The tracking software then sends the video stream to the detector.

[0047] In one example, the detector identifies the position of the target in the field of vision within the image.

[0048] In another example, it is the target tracking software that makes it possible to identify the position of the target in the field of vision.

[0049] According to one example of this embodiment, the system includes an appearance template provider that allows: to receive the image extracted by the detector along with its class, to assign a signature to the target for a camera, to identify a number of target characteristics such as color, etc., and to store the target characteristics and signature in memory. and finally to send to the tracking software suite the signature corresponding to the received image.

[0050] According to one example of this embodiment, the tracking software suite transmits the target's signature, speed, direction, class, and state to the surveillance camera management module.

[0051] According to an example of this embodiment, the appearance model provider uses a re-identification component using a RESNET 50 type neural network to perform signature extraction of detected targets and thus recognize a signature from the target's characteristics on an image and further identify and extract target characteristics also called attributes.

[0052] The signature can be represented as a vector of floating points. The appearance model provider can then search a database for a similarity between a target image from one camera and a signature from a previous image from another camera. The similarity between two signatures can be calculated by finding the minimum cosine distance between two signatures from two images taken from two different cameras.

[0053] The model provider can, for example, send the last calculated signature, for example, with a link to the first similar signature identified to the tracking software to inform it that these two signatures are the same target.

[0054] Thus, it allows the target to be re-identified when it is in a new state within a camera's field of vision.

[0055] Thus, the appearance model provider can record and recognize the target in the same camera but can also do so from one camera to another.

[0056] According to one embodiment, the system further includes detector software enabling the identification of moving targets on a video from a camera and in that the detector isolates the target, for example, by cutting out a rectangle in the video.

[0057] According to an example of this embodiment and the previous embodiment, the detector includes an artificial target recognition neural network capable of performing image recognition to identify the class of the target, the target being a pedestrian, a two-wheeler, a four-wheeled vehicle such as a truck or a car or others.

[0058] The two-wheeled target can include two sub-targets comprising a bicycle sub-target and a motorcycle sub-target.

[0059] In "other", the target could be, for example, a pedestrian using a means of transport such as rollerblades, a scooter, etc...

[0060] According to one example, the detector can include a feature extraction algorithm to abstract the image to send as input to the artificial neural network for target recognition.

[0061] According to another example, the detector includes a convolutional neural network of the type SSD, from the English "Single Shot multiBox Detector".

[0062] The invention also relates to a method of tracking a target using a surveillance system, the method being defined by claim 8.

[0063] According to one embodiment, the process may further include: A step of creating a list of probable camera identifications with their probability of success and adding to the camera a number of probable camera identifications with their probability.

[0064] According to one embodiment, the process comprises: A step of receiving a set of data from all target tracking software, A step of identifying a passage of the target from one field of view of one camera to another field of view of another camera by comparing the signatures, A step of incremental learning of the artificial predictive neural network from the last position on the field of view of the previous camera as well as the identification of the previous camera and the identification of the camera having an image representing the same target for example by receiving information that the signature on the target in the new state is close to a signature of the target in the exit state of the previous camera.

[0065] By proximity, the system can have a minimum distance between two signatures to accept them as similar, for example, for incremental learning. For instance, targets with rare (distinctive) attributes can be used for each class so that two signatures of the same class in two cameras are distinct from other signatures of other targets. For example, incremental learning for a pedestrian class can be performed with signatures, distinct from other individuals, representative, for example, of a pedestrian wearing a red top and bottom, easily recognizable compared to other pedestrians.

[0066] This allows the camera management module to perform incremental learning from targets with a high probability of being the same individual, i.e., the same car, truck, pedestrian, etc.

[0067] According to an example of this embodiment, the incremental learning step further includes as input information the last direction and last speed of the target received in the previous camera.

[0068] According to an example of this process, the process further includes: a step of detecting a target in its new state by the tracking software (target entering the camera's field of view), a step of sending an image of the video stream including the target in its new state to the detector, a step of detecting the class and position of the target in the surveillance camera's field of view, a step of sending the target's position in the received image, as well as the target's class, to the tracking software and an image of the target, a step of sending an image of the target to an appearance model provider, a step of signing the target, allowing the target to be identified, a step of sending the position, class, signature, and new state of the target to the camera management module, a step of detecting the target in its exit state, a step of sending the image of the target in its exit state to the detector.a step to determine the target's position within the surveillance camera's field of view, a step to send the target's position, class, signature, and output state to the camera management module.

[0069] In an example of the preceding embodiment, the process may further include a step of calculating the direction and velocity of the target in the output state as a function of the positions and date including the time of the target from the new state to the output state.

[0070] In another example, it is the camera management module that calculates the speed and direction and receives the date (including the time) for each state of the target and position.

[0071] According to one embodiment, the process includes a succession of sending images from the tracking software and a succession of receiving the position of the target from the new state to the exit state of the target. (That is to say, also to the new state).

[0072] This allows for greater precision in determining the target's speed and direction. Indeed, it enables the camera to account for changes in the target's direction, as well as any slowdowns, pauses, or accelerations as the target moves within its field of view.

[0073] According to one embodiment, the process includes a step of recording target characteristics determined by the appearance model provider and recording the signature and characteristics of the target in order to be able to re-identify the target with another image of another profile of the target in particular an image of the target from another camera having a different signature but close to the signature recorded previously.

[0074] The invention and its various applications will be better understood by reading the following description and examining the accompanying figures. BRIEF DESCRIPTION OF THE FIGURES

[0075] These are presented for illustrative purposes only and are in no way intended to limit the invention. The figures show: there figure 1 , represents an architecture according to an embodiment of the invention figure 2 represents an image flow diagram according to an embodiment of the invention; the figure 3 schematically represents a state cycle of a cycle within a camera's field of view; the figure 4 represents a schematic diagram of an example of an artificial predictive neural network. DESCRIPTION OF PREFERRED FORMS OF IMPLEMENTATION OF THE INVENTION

[0076] Unless otherwise specified, the various elements appearing on several figures will have retained the same reference.

[0077] There figure 1 represents an architecture of a monitoring system for an embodiment of the invention.

[0078] The surveillance system allows for real-time monitoring of a site comprising Xn CCTV cameras (n being the number of cameras X). In this instance, only two cameras, X1 and X2, are shown for simplicity, but the invention relates to a camera node comprising several dozen or even hundreds of CCTV cameras X.

[0079] The surveillance system also includes LT target tracking software, in this case LT tracking software by CCTV camera 1.

[0080] These LT target tracking software programs can be grouped in a machine such as a computer or distributed, for example in each of the cameras.

[0081] The surveillance system includes at least one PM management module for 1X surveillance cameras having at least one input to receive data from the X CCTV cameras from the LT target tracking software set.

[0082] The data sent to the PM management module includes at least one identification of one of the X cameras recording video over a monitored area of ​​the site, at least one signature corresponding to a moving target detected on the video as well as the position of the target in the video and at least one state of the target in the area monitored by the camera.

[0083] By signature, we mean a target identification that allows the target to be found in another camera. The designation of the signature and the identification of the same target in another camera are explained in more detail later.

[0084] There figure 3 represents the different states of the target.

[0085] The state of the target can be an input state E; this state is put on the target when the target has just entered the field of view of camera X1.

[0086] The target state can also be active L, which is the state in which the target is in the camera's field of view after the target in input state E has been sent to the camera management module.

[0087] The target's state can also be in an output state S, which is the state in which the target is at one end of the field of view. For example, within the 10% of the camera's field of view surrounding the 90% of the field of view in the active state A.

[0088] Finally, the target can be in a disappearance state O, when the target has moved out of the camera's field of vision (it no longer appears on the video).

[0089] The target's position can be identified, for example, by coordinates within the camera's field of view (based on an x-coordinate and a y-coordinate within the field of view). If the camera is mobile, the camera's angle can also be added to the coordinates to provide the target's position within the field of view.

[0090] It is evident that, in the case where the target is in the disappearance state O, either the LT tracking software does not send a position, or sends the position recorded in the exit state S or a position close to the latter if the target has moved between its position sent in the exit state and its last identified position.

[0091] In this example, the target position is defined by a detection component, also called detector D in the following.

[0092] The surveillance system also includes a CP (Predictive Targeting) neural network for target localization within an area monitored by camera X1. This CP neural network is connected to the PM (Program Management) module. The CP neural network includes a target information acquisition input. Specifically, the data sent to the CP neural network includes the identification data of the camera in which a target was detected, as well as the position data at the exit state, indicating the target's position at the exit state before it transitioned to the disappearance state.

[0093] The CP prediction neural network can be, for example, a multilayer perceptron-type neural network for classification purposes, in this case MLPCClassifier from the Scikirt Learn library.

[0094] The CP prediction neural network includes an output, containing at least one identification of a probable camera whose target will likely be identified, connected to the PM camera management module. The PM camera management module includes an output for transmitting the at least one probable camera identification to a human-machine interface device (HMI) including a display, such as a computer.

[0095] In addition, the camera management module can transmit video from the prediction camera.

[0096] In this particular embodiment, the CP prediction neural network receives additional data from the target. Specifically, it receives direction and / or sense data from the PM camera management module, as well as speed data.

[0097] The direction and speed data of the target are in this case calculated by the PM camera management module.

[0098] In addition, in this example, the predictive neural network receives target class data as input from the PM camera management module.

[0099] The class of a target is determined by detector D, which transmits the target class to the LT tracking software.

[0100] In other words, detector D provides the position and class of a target in an image. Specifically, detector D can provide the position and class for multiple targets in an image.

[0101] Detector D can be a single software for all LT tracking software or can be duplicated to work with each with a predetermined amount of tracking software.

[0102] Detector D comprises a neural network, specifically a convolutional SSD network. This neural network, which has undergone training, determines the class and position of each target in a received image. In this example, the class could be a pedestrian, a two-wheeler, a car, or a truck. Other classes include a bus, a taxi, or a tuk-tuk. Detector D can also segment the target within the image; in this case, it can segment multiple targets into rectangles. This is schematically represented in the diagram. figure 1 under the representation of a film I.

[0103] In what follows, an example of a target will be described, in this case having a class called "pedestrian".

[0104] The LT tracking software transfers the clipped images, and therefore in this case the clipped image of the "pedestrian" target, to an AMP appearance model provider.

[0105] It is this appearance model provider (AMP), referred to hereafter as the AMP provider, that transmits the target's signature to the LT tracking software. The AMP provider uses a Reid re-identification component, which in this case comprises a neural network, specifically a ResNet 50, but could also use a data acquisition algorithm. The AMP provider sends the target image to the Reid re-identification component, which calculates a target signature based on measurements taken from the target image. This will subsequently allow the target to be re-identified by correlating two signatures in another image received from a different camera.

[0106] In this example, the Reid re-identification component provided to the AMP provider also provides measured information on the target, also called "attributes", such as for a pedestrian the color of a top or bottom, the type of hair (brown or blond or red) and other characteristics... For a car, this could be the color of the car, the height etc... and if possible the reading of the car's license plate.

[0107] Thus, in the example of the figure 1From the image of the target, taken by camera X1 (a visible pedestrian camera), and sent to the AMP provider by the tracking software, the Reid re-identification component determines that the pedestrian is a tall, dark-haired woman wearing a red dress and carrying a beige handbag measuring between 30*40cm and 35*45cm. The Reid component calculates a signature of this target based on its measurements and further identifies its characteristics (attributes). The AMP provider then stores the target's characteristics (attributes) and signature in its database. The signature can be represented as a vector of floating-point numbers.

[0108] The AMP appearance model provider or the re-identification component can then search a database for a similarity to a signature from a previous image from another camera. The similarity between two signatures can be calculated, for example, by the minimum cosine distance between two signatures from two images from two different cameras. In this example of the pedestrian image from camera X1, the provider finds a similar target signature from camera X2, recorded in the database, for example, 3 minutes earlier. For instance, the signature differs because the identified size is medium. Knowing that a target can change from large to medium size and vice versa when it is close to the boundary separating the two sizes due to possible measurement discrepancies, the AMP provider deduces a similarity between these two signatures.In this example there are three size types: small, medium and large, but there could be more.

[0109] According to another embodiment, it is the management module that performs this search for similarity of the two signatures in the database.

[0110] The AMP provider in this example sends the last calculated signature, for example with a link to the similar signature identified, to the tracking software to inform it that these two signatures are potentially the same individual.

[0111] The LT tracking software then transmits the target's signature in its new state, the similar signature, to the PM camera management module.

[0112] The PM camera management module can thus search for the signature similar to the output state received from a tracking software of the X2 camera previously, make the possible correlation between these two signatures and send the probable neural network in incremental learning mode prediction data.

[0113] The learning prediction data can be: The class of the target, in this case pedestrian, The position of the target in the output state, in this case for example coordinates in the image, The speed of the target calculated in the output state, for example 5 km / h The direction of the target calculated in the output state, for example a function of a straight line The identification of the previous camera (i.e. that of the target in the output state), in this case X2, The identification of the new camera, in this case X2.

[0114] Thus the predictive neural network can learn that it is likely that a pedestrian leaving the field of vision of camera X2 at a certain position and speed and direction will then be visible on camera X1.

[0115] A single camera image stream is displayed on the figure 2 , according to an example of a target prediction method.

[0116] The process includes: a filming step 1 by cameras X, a detection step 2 by detector D of a target and identification of its class and position, a setting to the new state N by the tracking software LT (target entering the camera's field of view), a sending of the target image to the AMP provider, a targeting identification step 3 by the AMP provider by assigning it a signature and recording a list of its attributes (type, size, etc.), a change of state from new N to active A by the tracking software LT after receiving the signature from the AMP provider, a change of state from active A to the exit state S when the target's position obtained by the detector is in an exit zone, a sending to the camera management module PM of the target's exit state along with its class, signature, position, speed, and direction.a prediction step 4 by the neural network predicting the likely output camera of the target, a reception step by the PM camera management module of a data set from the entire LT tracking software and video processing 5 where the tracked target is requested by an operator by adding a list of likely cameras in order of probability and, for example, the video from the most likely camera, a display step by a human-machine interface device 6 of the live video including the identification of the likely camera, an event step 7 in which the requested target appears in a video from a camera identified in the new state, a prediction verification step 8 checking that the camera identification, in which the target reappears in the new state, is named in the camera prediction list, if the camera is not in the list,The PM state-change module performs a learning step 9 for the neural network by providing it with learning prediction data.

[0117] The artificial predictive neural network can thus learn the topology of CCTV camera X and suggest a camera or list of cameras whose target will likely appear when an operator sends a target tracking request.

[0118] For example, the CP prediction artificial neural network can learn continuously or intermittently. Thus, in the event of camera addition or degradation, the artificial neural network can suggest the likely camera(s) X, along with, for example, a percentage of subsequent cameras on the live video feed of the target in its output state. As another example, it can overlay smaller live video feeds onto the video showing the target in its output state until the target disappears.

[0119] There figure 3 represents a state cycle of a target within the field of view of a CCTV camera X.

[0120] A target identified by the LT tracking software, in this case entering a field of view of the X1 camera, changes to a new state and an image of the video is sent to the detector which cuts out an image of the target, identifies a class as well as the position of the target, then the cut-out image is sent to the AMP provider, which sends a calculated signature back to the LT tracking software.

[0121] The LT tracking software puts the target into active state A and sends video images to the detector, which returns the target's position until the target is in an exit zone of camera X1's field of view. At this point, the LT tracking software puts the target into state S. If the target returns to an active zone of the field of view—for example, if the target turns around and returns to the center of the camera's field of view—the tracking software puts the target back into active state A until it returns to an exit zone of camera X1's field of view. Finally, when the target leaves camera X1's field of view, the software puts the target into disappearing state D.

[0122] Because the LT tracking software knows the different positions of the target in the camera's field of view, it can calculate its speed and direction.

[0123] According to one embodiment, the LT tracking software sends target information data to the PM camera management module at each change of state, i.e. in this case the speed, position, direction, signature, class and state.

[0124] The Pm camera management module therefore receives this data from a set of LT tracking software. This is how the management module can identify the passage of a target, via signatures, from one camera's field of view to another camera's field of view, in order to transmit this information to the artificial neural network for prediction and incremental learning.

[0125] The PM surveillance camera management module can also receive, from a computer 6 used by an operator, a request to track a target, for example the operator clicks on a target in video X1.

[0126] Once the target is identified, it is tracked by the LT tracking software. When the target transitions to the exit state, the PM management module sends the following information to the CP prediction neural network: camera identification, target class, target position at the exit state, calculated target velocity at the exit state, and calculated target direction at the exit state. The neural network then provides a list of probable camera identifications and their corresponding probabilities.

[0127] The neural network therefore displays on the video the list of cameras with a high probability, for example the 5 most probable cameras.

[0128] The PM surveillance camera management module can also add to an area of ​​the video whose target is identified in the output state, for example top left, a video filmed by the most probable camera identified by the artificial neural prediction network.

[0129] According to one embodiment, the PM surveillance camera management module can send a list of surveillance camera nodes as well as the weights of the paths between the cameras calculated by class.

[0130] There figure 4 This diagram represents a schematic representation of an example of a CP (Predictive Control) artificial neural network. In this example, the CP artificial neural network comprises three layers. The CP artificial neural network includes an input layer (CE) containing neurons with an activation function, for example, of the rectified linear unit (ReLU) type. Specifically, the input layer (CE) comprises nine neurons: one for class, four for target position, one for velocity, and one for camera identification, for a total of nine variables.

[0131] The four target position neurons are, in this case, a target X position neuron along an X axis in the image, a target Y position neuron along a Y axis in the image, a camera X position neuron along a mobile axis for the camera filming the target, i.e. producing the image in which the target is detected and a camera Y position neuron along a mobile Y axis of the camera.

[0132] Of course, the CE input layer could contain more or fewer than nine neurons. For example, one more neuron could be a Z-position neuron for the camera moving along the camera's Z-axis. In another example, the number could be eight neurons, with each camera calculating the target's Z-position, and the two camera X and Y position input neurons would not exist.

[0133] The CP prediction artificial neural network also includes a hidden CC layer comprising neurons with an activation function, for example, of the rectified linear unit type. In this example, the CP prediction artificial neural network comprises one hundred neurons. The CP prediction neural network could have more than one hidden layer.

[0134] The CP prediction artificial neural network in this example includes an output layer CS comprising one neuron for each camera probability identified as having the highest probability of the target appearing. Specifically, the CS output layer comprises n neurons with a loss layer activation function, also known as Softmax, to predict the camera. For example, the CS output layer includes five output neurons for five probabilities of five surveillance cameras identified as having the highest probability of the target appearing. The neurons identified as having the highest probability of the target appearing form the list of camera identifications by probability.

[0135] In one embodiment, the CP prediction artificial neural network can provide a probable camera identification whose target will likely be identified based on a sequence of previously identified cameras in the input layer's identification neuron. For example, the list of camera identifications by probability is obtained based on this sequence.

[0136] Of course, the invention is not limited to the embodiments just described.

[0137] The present invention has been described and illustrated in the present detailed description and in the figures of the accompanying drawings, in possible embodiments. The present invention is not limited, however, to the embodiments shown. Other variations and embodiments can be deduced and implemented by a person skilled in the art upon reading the present description and the accompanying drawings, insofar as they fall within the scope of the attached claims.

[0138] In the claims, the terms "include" or "comprise" do not exclude other elements or steps. The various features presented and / or claimed may be advantageously combined. Their presence in the description or in different dependent claims does not preclude this possibility. Reference symbols shall not be construed as limiting the scope of the invention.

Claims

1. Surveillance system for at least one site comprising video surveillance cameras, the surveillance system comprising: • at least one surveillance camera management module having: - at least one input for receiving data: i of at least one identification of one of the cameras recording video in a monitored zone of the site, ii of at least one signature corresponding to a moving target detected on the video, iii of a state of the target in the zone monitored by the camera, being among: (a) an input state in which the target has entered the field of view of the camera, (b) an output state in which the target is at one end of the field of view, (c) a disappearance state in which the target has left the field of view, iv of positioning of the target in the monitored zone, v identification of a likely camera, - and an output for sending data: i of identification of the camera in which a target has been detected, ii of positioning of the target in the output state, iii direction data of the target, iv identification of a camera in which the target has reappeared in a zone monitored by the camera after the disappearance of the target v the at least one identification of the likely camera, • an artificial neural network for predicting the location of a target in a zone monitored by a camera, comprising: - a target information acquisition input connected to the surveillance camera management module, comprising data for prediction including data: i of identification of the camera in which a target has been detected, ii of positioning of the target in the output state, iii of direction in the output state received or calculated by the management module, - an output of at least one identification of a likely camera in which the target will probably be identified transmitted to the camera management module for transmission to a unit comprising a screen, - wherein the camera management module transmits to the prediction artificial neural network in automatic learning mode the data for predictions as well as the identification of a camera in which the target has reappeared in a zone monitored by the camera after the disappearance of the target.

2. Surveillance system according to the preceding claim, wherein the surveillance camera management module additionally receives speed data of the target and the target data for prediction sent to the prediction artificial neural network additionally comprise these direction data in the output state.

3. Surveillance system according to either of the preceding claims, wherein the surveillance camera management module additionally receives class data of the target and the target data for prediction sent to the prediction artificial neural network additionally comprise these class data in the output state, the target class can be a pedestrian, a two-wheeler, a four-wheeler or the like.

4. Surveillance system according to any of the preceding claims, further comprising: • a software package for target tracking by surveillance camera allowing: - the target to be followed and its direction and speed to be deduced, - the state of the target to be created • a detector allowing - a target image to be extracted from the video, - image recognition to be performed to assign the class of the target, - the extracted image and its class to be sent to the tracking software package, • an appearance model provider allowing: - the image extracted by the detector and its class to be received, - the target to be given a signature, - a number of characteristics of the target, such as color etc..., to be identified - the characteristics and signature of the target to be stored in a memory, - the signature corresponding to the image received to be sent to the tracking software package, and the tracking software package transmits the signature, speed, direction, class and state of the target to the surveillance camera management module.

5. Surveillance system according to any of the preceding claims, wherein the surveillance camera management module is able to receive a request for prediction of a target comprising a signature, and the prediction artificial neural network can send at its output a list of camera identifications by probability and the surveillance camera management module is further able to send to the interface machine the identification of the camera of which the target is in the active state and a probability-ordered list of possible likely camera identifications.

6. Surveillance system according to any of the preceding claims, wherein the surveillance camera management module is able to add, to a zone of the video of which the target is identified in the active state or in the output state, a video filmed by a likely camera identified by the prediction artificial neural network.

7. Surveillance system according to any of the preceding claims, wherein the surveillance camera management module is able to send a list of surveillance camera nodes as well as the path of the target from the first camera that identified the target to the most likely camera of the target and, for example, another characteristic such as class, path weights between cameras calculated by class.

8. Method of tracking a target using a surveillance system according to any of the preceding claims, wherein the tracking method comprises: • a step of requesting the tracking of a target on an image, • a step of recognizing a similarity of the target signature in a signature database, • a step of tracking the target on a video produced by a surveillance camera filming a zone of the site, • a prediction sending step in which the management module sends target data to the prediction artificial neural network, including at least the position of the target when said target is in a zone of the zone filmed by the surveillance camera corresponding to an output state, • a step of receiving from the prediction artificial neural network an identification of a likely camera in which the target may appear if it disappears from the field of view of the identified camera • a step of adding the identification of the likely camera to the video from the identified camera. • a step of sending the video and the identification of the likely camera to the unit comprising a screen.