Monitoring system and method for monitoring facilities, and computer program product

Through a multi-sensor system and machine learning algorithm, combined with facility models and logical relationships, the detection difficulties of the facility monitoring system under non-ideal lighting conditions are solved, and autonomous and reliable assessment of the status of facility components and anomaly detection are achieved.

CN114494768BActive Publication Date: 2025-09-26HEXAGON INNOVATION CENTER LTD
View PDF 15 Cites 0 Cited by

Patent Information

Application Number
CN202111627648.8
Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Filing Date
2018-10-29
Publication Date
2025-09-26
Estimated Expiration
2038-10-29

AI Technical Summary

Technical Problem

Existing image-based facility surveillance systems perform poorly under non-ideal lighting conditions, especially in darkness, where they struggle to effectively detect the presence, classification, and status of objects.

Method used

A multi-sensor system, including RGB cameras, depth cameras, infrared cameras, microphones, etc., combined with a central computing unit and a state inference device, combines and classifies state patterns through facility models and machine learning algorithms, and uses topological, logical and functional relationships to perform anomaly assessments.

Benefits of technology

It improves the accuracy and reliability of facility monitoring under non-ideal lighting conditions, can autonomously detect and evaluate status changes and anomalies of facility components, and provide timely safety assessments and notifications.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN114494768B_ABST
    Figure CN114494768B_ABST
Patent Text Reader

Abstract

The present invention relates to a monitoring system and method for monitoring a facility, and a computer program product. A monitoring system for monitoring a facility comprises: a monitoring robot having a main body, a drive system, and a motion controller for controlling the motion of the monitoring robot. The monitoring system further comprises: at least a first monitoring sensor designed to acquire first monitoring data of at least one object of the facility; a state detector configured to detect at least one state associated with the object based on the monitoring data, wherein the state detector is configured to: notice state ambiguity of the state, in particular by comparing it with a predetermined ambiguity threshold; upon noticing the state ambiguity, triggering, by the motion controller, an action of the monitoring robot, the triggered action being adapted to generate state verification information about the object, the state verification information being suitable for resolving the state ambiguity; and resolving the state ambiguity by taking into account the state verification information.
Need to check novelty before this filing date? Find Prior Art

Description

[0001] This application is a divisional application of the original invention patent application with application number 201880099180.X (International application number: PCT / EP2018 / 079602, application date: October 29, 2018, invention name: Facility monitoring system and method). Technical Field

[0002] The present invention relates to systems and methods for monitoring facilities such as buildings. Background Art

[0003] State-of-the-art image-based detectors rely on the rich color and texture information present in RGB images. However, their performance is known to degrade under non-ideal lighting conditions, particularly in darkness. In those non-ideal situations, supplementary information such as depth maps, point clouds (PCs), and / or infrared (IR) signals can be used to determine the presence, absence, classification, status, etc. of objects of interest in scenes where classic image detectors degrade or fail.

[0004] As shown in the above-mentioned prior art, the known systems therein focus on:

[0005] Apply detection in each of these modalities separately and independently, and then combine the results. So-called parallel approaches, for example by fusing the output of an RGB detector with the output of a depth detector, as in US 8,630,741, US 2014 / 320312, GB 2546486, [2] or others; or

[0006] Linking detectors for each modality. So-called serial or hierarchical approaches, for example, start by first clustering objects in the odometry data space and then use this clustering information to guide a second-stage visual detector to the found regions of interest, as described in KR 101125233, [1] or elsewhere.

[0007] Another approach, such as proposed in US 2009 / 027196, [4] or elsewhere, attempts to learn the best possible detector hierarchy from the data using a tree structure. Thus, a single general model is learned. Summary of the Invention

[0008] According to a first aspect, the present invention relates to a facility monitoring system for performing a safety-related assessment of anomalies or abnormal states of a facility, the facility comprising facility components, such as rooms, doors, windows, walls, floors, ceilings, electrical equipment, pipes, etc. "Facility" in a broad sense refers to buildings, such as houses, warehouses, industrial plants or other complex facilities, as well as entire properties or installations, such as (large) ships or aircraft, construction sites or another restricted or potentially dangerous area. The facility monitoring system comprises a central computing unit, which provides a model, preferably a building information model (BIM), of the facility. The facility model provides topological and / or logical and / or functional relationships of at least some of the facility components, such as an indication of the location of a specific window, the neighborhood of a specific room or a specific door allowing access to a specific room.

[0009] In addition, the system includes a plurality of monitoring sensors, which are particularly suitable for continuously monitoring at least a plurality of facility components and for generating monitoring data including monitoring information about the plurality of facility components. The plurality of monitoring sensors are configured to provide data about one or more facility components (either via corresponding sensors on the facility components themselves, or via two or more sensors that generate monitoring data in a working relationship), the plurality of monitoring sensors being suitable for monitoring or observing the status of the corresponding facility components and including, for example, one or more RGB cameras, depth cameras, infrared cameras, microphones, contact sensors, motion detectors, smoke detectors, key readers, ammeters, RIM cameras, laser scanners, thermometers. However, an input device suitable for allowing an operator to input monitoring data may also be included. Optionally, one or more monitoring sensors are mobile sensors, in particular as part of a mobile robotic platform.

[0010] The system also includes a communication device for transmitting data from the monitoring sensor to the central computing unit and a state derivation device for analyzing the monitoring data and deriving at least one state. Such a state is, for example, the position or orientation of a building component, its appearance, including an object associated with the building component (for example, the introduction of an object into a room or the removal of an object from a room) or any significant difference in the monitoring data of the facility component when comparing monitoring data generated at two different times or over a period of time. The state derivation device is preferably integrated into the monitoring sensor and / or the central computing unit. In other words, the state derivation device is a device capable of extracting the state of the facility component from the monitoring data related to the facility component. Such a state derivation device is, for example, a detector that detects a change from one state to another (for example, a sensor that detects that a window is open), or is implemented as software such as an image processing algorithm that is capable of determining the state of a device present in an image.

[0011] Furthermore, the system includes a recording device for recording state patterns by combining the states of one or more facility components based on topological or logical or functional relationships. In the case of a facility component, a priori information about this relationship is given. This is achieved by combining the states of components that are in a logical relationship (e.g., in the simplest case, combining the same monitoring features / data of the same monitoring sensor at subsequent points in time). In the case of two or more facility components, the relationship is, for example, that they are spatially adjacent or related to a common function. The recorder also assigns a timestamp to the state pattern, e.g., the time when one of the potential states was monitored.

[0012] Furthermore, the central computing unit is configured to provide a critical-noncritical classification model for classifying the recorded state patterns. The critical-noncritical classification model takes into account time stamps and topological and / or logical and / or functional relationships provided by the facility model and includes at least one "noncritical" state pattern of the "noncritical" category, and performs a criticality classification, wherein the recorded state pattern is classified as "critical" or "noncritical" based on at least one topological and / or logical and / or functional relationship and at least one time stamp (taking into account the time associated with the state pattern).

[0013] Optionally, the central computing unit is configured to provide a normal-abnormal classification model for the recorded state patterns, wherein at least one category is a normal category representing the classified state pattern as "normal", and the recorded state patterns are classified according to the classification model. Optionally, the classification model is part of the facility model. The central computing unit is configured to perform normality classification, wherein the recorded state patterns are classified as "normal" or "abnormal" according to the normal-abnormal classification model, and if the degree of deviation of the abnormal state pattern from the "normal" classification is higher than a certain threshold, the recorded state pattern that has been classified as "abnormal" is classified as "critical" according to the critical-non-critical classification model, thereby taking into account at least one timestamp.

[0014] In other words, the system is configured to classify detected facility status patterns as "normal" or "abnormal," and to examine, evaluate, or test abnormal status patterns with respect to facility safety / security, i.e., whether the status pattern is "critical" or "non-critical."

[0015] As another option, the central computing unit is also configured to establish a criticality classification model and (if present) optionally a normal-abnormal classification model, i.e., to start from a "blank" and to independently generate or create (and refine if necessary) the classification model by machine learning.

[0016] Optionally, the computer of the system is designed such that the criticality classification and optionally the normality classification take into account a schedule representing the times of human activities and / or automated activities associated with the facility, in particular wherein the schedule comprises planned working and / or operating times and / or comprises information about the types of planned activities and / or is embedded in the digital facility model.

[0017] As another option, the criticality classification and optionally the normality classification are implemented using at least one of the following: at least one of a rule-based system and a data-based system, the rule-based system is based on expert knowledge, specifically including decision trees, Bayesian decision networks, first-order logic, temporal logic and / or state models with hidden states, fuzzy logic systems, energy optimization-based systems, rough set-based systems, hierarchical decision systems, multi-agent systems, and the data-based system is based on previously generated states and / or state patterns, specifically including decision trees, decision tree ensembles, rule-based machine learning systems, energy optimization-based systems, probability generation or probability discrimination graphical models, probability generation or probability discrimination graphical models specifically including Bayesian networks, Markov random fields, conditional random fields, restricted Boltzmann machines; fuzzy logic systems, neural networks, specifically deep neural networks, especially recursive neural networks or generative adversary networks, case-based reasoning systems, especially instance-based learning systems using k-nearest neighbor methods, kernel methods, systems using supervised or unsupervised clustering, neuro-fuzzy systems, systems based on collective classification and / or collective regression.

[0018] If the recorded state pattern involves at least two installation components, at least one topological and / or logical and / or functional relationship of the two installation components may optionally be taken into account based on the installation model and / or the gravity classification provided by the installation model.

[0019] As another option, the central computing unit is configured to determine the probability of a false positive in the classification, in particular, whereby the central computing unit is further configured to trigger, if the probability is above a defined threshold, the acquisition of additional data using the communication means, in particular data from a database and / or acquired by at least one of the monitoring sensors, so that, taking into account the additional data, the subsequent classification results in a probability below the defined threshold. In other words, if the uncertainty or unreliability of the classification is too high, the system automatically retrieves additional data on the state pattern of one or more facility components, respectively, so that additional information describing the state (e.g. parameters or characteristics) is available, thereby allowing a more certain designation of "critical" or "non-critical" or possibly "normal" or "abnormal".

[0020] Optionally, the state is derived using at least one of detection and / or identification of a person, detection of door and / or window opening, detection of fire and / or smoke, detection of abandoned objects, or identification of activity.

[0021] Optionally, during the classification process described above, verification information for the assignment to a category is provided, preferably based on machine learning, in particular at least one of a confirmation of the assignment and / or a rejection of the assignment. Alternatively or additionally, the algorithm is designed to provide information on a change in the assignment of one of the multiple classes, in particular by removing the assignment and reassigning it to at least one of another class, or providing the definition of a new class, in particular at least one of modifying an existing class, dividing an existing class into at least two new classes, and merging multiple existing classes into a new class. As another option, the classification algorithm is designed to provide instructions based on machine learning for removing a class from the corresponding classification model, identifying a first selection of monitoring data to be used for classification, and / or a second selection of monitoring data to be ignored for further processing. As another option, a feedback loop mechanism is provided, in which a human operator can verify or discard notifications, and this information is fed back to the system once or periodically for relearning and improving the corresponding classification model.

[0022] Optionally, the system further comprises an update function for processing assignments and providing updated information of the classification model and / or the facility model.

[0023] Optionally, the central computing unit includes a state pattern recording device. As another option, the monitoring data and / or state patterns of one or more monitoring sensors at different locations and times are derived using at least one of person re-identification, person tracking, a hidden Markov model (HMM), a recurrent neural network (RNN), a conditional random field (CRF), a topological and / or logical and / or functional relationship of the measured facility components, whereby the relationship is based on and / or provided by the facility model. Taking into account the spatiotemporal relationship of the investigated facility components allows for enhanced state derivation associated with these facility components, for example, in order to better link the monitoring data of two adjacent facility components together to form a state pattern.

[0024] Monitoring based on state sequences / patterns has the advantage that more aspects are available (examined by the classification algorithm), leading to a more insightful assessment. In other words, instead of considering a single action or state in isolation, the context of an action or action string is considered, so that its classification allows for a more insightful or discriminating determination of (optional) anomalies and, primarily, criticality. For example, if changes in a facility component are investigated as part of a state and state sequence, it can be verified whether the subsequent state of the same or another facility component that was expected due to experience (machine learning) and / or due to logical, functional or topological relationships is indeed investigated. Using sequences and / or patterns instead of single states allows, for example, to consider the behavior of people at or in the facility represented by the state sequence and / or pattern. The present invention allows for an assessment of how unusual or unusual, or to what extent, an action or process affecting the state of a facility component is with respect to safety considerations.

[0025] Optionally, the facility model comprises sub-facility models representing at least a portion of the facility components, and the central computing unit comprises an assignment device configured to assign the monitoring data of each facility component to the corresponding sub-facility model, in particular wherein the monitoring data comprises location information, and the assignment of the survey data of each facility component to the corresponding sub-facility model is based on the location information, wherein the location information comprises information about the location of the facility component, an object or a person, or the location of the corresponding monitoring sensor. Preferably, the location information comprises coordinates in a global or local coordinate system, in particular in a coordinate system of the facility, and / or an identifier or code identifying a predetermined location or area in the facility, in particular a room.

[0026] As a further option, the system includes an output device for issuing a notification, in particular a warning and / or a command, if the status pattern indicates a "critical" overall state of the facility, in particular wherein the output is a graphical output of the facility model, in particular within a three-dimensional visualization, and / or wherein the notification includes a plurality of options for reacting to the detected state, wherein each option can be rated taking into account the reliability (probability) of the classification. The latter option means that not only are the reaction options presented by the system, but these options are also weighted taking into account the assessed reliability, so that, for example, a user of the system is assisted in better selecting from the possible reaction options.

[0027] In another optional embodiment of the present invention, the facility model includes topological facility data and the monitoring data includes time data. As another option, each state is represented by a combination of features that characterize the state at least with respect to time and location within the facility topology.

[0028] Optionally, the corresponding classification model comprises a database of stored state patterns, and the classification and evaluation of the detected state pattern is based on a comparison with at least one, in particular multiple, stored patterns and a determination of a level of deviation between the detected pattern and the stored patterns, in particular using a deviation threshold and / or statistical probability.

[0029] As another option, the criticality and, optionally, normality classification model comprises an n-dimensional feature space, wherein the state pattern is represented by an n-dimensional feature vector, and in particular, wherein the corresponding class is represented by a portion of the n-dimensional feature space. The state pattern is then located in the n-dimensional feature space, and its assignment to a class can be evaluated based on geometric principles or machine learning algorithms.

[0030] Optionally, the classification model comprises a (deep) neural network, in particular wherein the sequence and / or pattern of states are fed into different units of an input layer of the neural network.

[0031] In a further developed embodiment, person recognition is used to detect the state, and the classification and evaluation algorithm is configured to take into account the identity or type of the recognized person for the classification and evaluation.

[0032] Optionally, the classification is based on categories based on semantic properties, categories based on topological and / or geometric properties, linear classification, in particular based on Fisher linear discriminant, logistic regression, naive Bayes classifier or perceptron, support vector machine, in particular least squares support vector machine, quadratic classifier, kernel estimation, in particular k-nearest neighbor, boosting, decision trees, in particular random forests, sparse grids, deep learning, in particular based on neural networks, in particular convolutional neural networks and / or learning vector quantization.

[0033] Optionally, the central processing unit comprises at least one server computer, in particular a cluster of server computers operating as a cloud system, wherein the communication device is adapted to transmit data to the central processing unit via an Internet or intranet connection.

[0034] The present invention also relates to a facility safety monitoring method for performing safety-related assessments of abnormal states of a facility consisting of facility components, the method comprising the following steps: providing a model of the facility, in particular a dynamic model, which provides topological and / or logical and / or functional relationships of at least a part of the facility components, in particular a building information model (BIM), and / or wherein the building model includes a sub-building model representing at least a part of the building components; monitoring a plurality of facility components and continuously generating monitoring data relating to these facility components; analyzing the monitoring data and detecting at least one state thereof; recording a state pattern by grouping the states of one or more facility components such that the state pattern represents a sequence and / or pattern of states associated with one or more monitored facility components; classifying the at least one recorded state pattern based on a critical-non-critical classification model having at least one non-critical class for a "non-critical" state pattern, wherein the classification is based on the at least one topological and / or logical and / or functional relationship and at least one timestamp of the recorded state pattern.

[0035] Furthermore, the present invention relates to a computer program product comprising a program code which is stored on a machine-readable medium or implemented by electromagnetic waves comprising program code segments and which has computer-executable instructions for performing the steps of the method according to the invention, in particular when run on a computing device of a facility monitoring system according to the invention.

[0036] A second aspect of the present invention relates to a surveillance system for monitoring a facility, such as property inside or outside a building, an industrial or marine complex, a construction site, or another restricted or potentially hazardous area. The system includes a surveillance robot comprising a main body, a drive system or drive unit, and a motion controller for controlling the robot's movements. The robot can be ground-based or implemented as an aerial vehicle. The system also includes at least a first surveillance sensor, such as a part of the robot, configured to acquire first surveillance data of at least one object on the property, such as an object such as a room, a door, a window, a person, an electrical property facility, or a mobile property object such as a package or container.

[0037] The first monitoring sensor is preferably a sensor capable of measuring a larger area or a larger space. It is preferably implemented as a non-contact sensor and includes, for example, at least one or more cameras (photo and / or video, visible and / or other parts of the electromagnetic spectrum) and / or microphones and / or RIM cameras and / or laser scanners and / or LIDAR and / or RADAR and / or motion detectors and / or radiometers.

[0038] The detector further comprises a state detector, which is designed to evaluate the first monitoring data and detect the state of the object based thereon. The state of the object is, for example, its position, orientation, color, shape, properties (e.g., animate or inanimate), or operating state (e.g., open / closed, full / empty), or anomalies, i.e., deviations from a normal situation or state. This means that not every state can be recorded as a state, but only abnormal states, e.g., states that have not been previously monitored.

[0039] According to the present invention, the state detector is configured to be aware of state ambiguity in the detected state. Thus, the detection unit not only infers the state associated with the object, but also detects or identifies any ambiguity in the detected state. In other words, the ambiguity is determined to indicate, for example, the degree of uncertainty in the state.

[0040] The state detector or the underlying computational unit is further configured to trigger an action of the robot by the action controller if an ambiguity is noted. If not, the derived state is respectively estimated to be unambiguous, i.e., if the state detection leads to an unambiguous result, the derived state is considered to be certain without triggering the action. The triggering criterion is optionally an ambiguity indication, where the uncertainty of the event of the object is above a defined threshold; an uncertainty below this threshold is interpreted as an unambiguous detection.

[0041] The triggered action is suitable for generating state verification information about the object, and the state verification information is suitable for resolving event ambiguity.

[0042] The state detector is further configured to consider the state verification information to resolve the state ambiguity.

[0043] In other words, if the estimated state detection is ambiguous, an action of the robot is triggered, which enables the robot to verify (or, if applicable, falsify) the detected state and thereby resolve any ambiguity. The triggered action generates verification information about the object, causing the computing unit to take into account the state verification information to deduce an unambiguous object state.

[0044] As a preferred option, the state detector is further configured to plan an action to be triggered so that the action is optimized with respect to the generation of verification information. Based on the first monitoring data, the computing unit determines which of at least two possible triggerable actions is more effective or best at resolving the event ambiguity. The variables to be optimized are, for example, the ability of the action (and the verification information generated thereby) to repair, complete, supplement or improve the first survey data, a defined state-related or important feature of the measured object, the degree of force, the quality and / or quantity of the verification information, and / or the time to fully perform the action. Thus, this determination or selection is based on estimates of the values ​​of the variables and / or on values ​​stored in a database of the robot, for example values ​​provided in a machine learning phase, in which the robot learns the state ambiguity-related effects of different actions.

[0045] The optimization planning or determination depends, for example, on at least one of the object, the type or kind of the object, or the state of the object, state ambiguity, the (absolute) position and / or orientation of the object, the relative position and / or orientation of the robot / measurement sensor and the object, environmental conditions, the quality and / or quantity of the first monitoring data, and / or the time or date. Preferably, the planning of the action includes selecting at least one of the following: the collection of second survey data, interaction with the object, and / or the collection of external data.

[0046] Alternatively, the triggered action includes acquiring, by the robot, second survey data of the object. The second survey data is optionally acquired by a second survey sensor of the robot, preferably wherein the first survey sensor and the second survey sensor are different types of sensors, for example, the first sensor is passive and the second sensor is active, or the first sensor is an optical sensor and the second sensor is a tactile sensor. Alternatively, the first survey sensor is suitable for coarse (and rapid) overall monitoring, while the second survey sensor is suitable for fine, detailed monitoring.

[0047] As another option, the computing unit is configured to determine which of the first and at least one second sensor to utilize to acquire the second monitoring data such that the second monitoring data is optimized for generating event verification information. In other words, the state detector determines which of the at least two measuring sensors is the most efficient means for verifying the state of the object.

[0048] The calculation unit is optionally further configured to determine this optimization by also taking into account other parameters such as the available time for the second monitoring. If, for example, there is an emergency situation (e.g., as deduced by the state detector from the detected state, even if this is uncertain), leaving only a very limited time for verification and subsequent reaction, the calculation unit selects the monitoring sensor or measurement procedure, respectively, which is optimized with respect to acquisition time and perhaps only second-best with respect to measurement accuracy.

[0049] Alternatively, the triggered action includes changing the acquisition position and / or orientation, such that the second monitoring data is acquired using at least one acquisition position and / or orientation different from the position and / or orientation used to acquire the first survey data. In other words, the robot or its monitoring sensor has a first acquisition position and / or orientation when generating the first monitoring data. Prior to generating the second monitoring data, a change in the acquisition position and / or orientation of the robot (or at least the measurement sensor) is triggered, thereby establishing a second acquisition position and / or orientation different from the first acquisition position and / or orientation.

[0050] Preferably, the state detector assesses which additional data or information, respectively, are missing or which additional data or information are required to resolve ambiguities based on the first survey data, and the calculation unit selects the second acquisition position and / or direction such that the additionally required data (approximately) can be generated when measuring in or with the second acquisition position and / or direction. The second acquisition position and / or direction is thereby optimized with respect to the generation of verification information, since the second measurement reveals an optimal amount of second monitoring data about the state.

[0051] Optionally, the computing unit provides a database of state-related characteristics for a plurality of property objects, and the collection of the second survey data of the objects includes a specific measurement of at least one of the state-related characteristics thereof. For example, the second collection position and / or orientation is planned based on a Next-Best-View (NBV) algorithm so that one or more of the event-related characteristics of the objects can be measured by one of the monitoring sensors located in the second collection position and / or orientation.

[0052] Preferably, the system's computer provides a correlation map that correlates state ambiguity with acquisition positions and / or orientations. Thus, for example, the aforementioned changes in acquisition positions and / or orientations are based on a correlation map, such as a best view, that provides information about the correlation between the derived ambiguity of the corresponding object and the acquisition position and / or orientation, such as the best or optimal acquisition position and / or orientation for a plurality of derivable object states. Alternatively, the correlation map is based on defined criteria representing object states and / or is established through machine learning. In the latter case, the correlation map is established, for example, during a training phase of the robot, wherein a plurality of acquisition positions and / or orientations and a plurality of object states observed at the corresponding positions and / or orientations are recorded.

[0053] Alternatively or in addition to acquiring the second survey data as a triggered action, the triggered action for generating the state verification information includes interaction between the robot and the object, such as tactile contact, in particular to move the object and / or to acquire tactile sensor data of the object, and / or applying a material to the object, in particular applying a liquid and / or paint, and / or outputting an acoustic and / or optical signal directed toward the object, in particular when the object is a person. In other words, the robot interacts in such a way that it can collect information for verification that would not be available without the interaction.

[0054] Optionally, the computing unit of the system is configured to determine an interaction (e.g., by selecting it from a plurality of possible interactions) such that it can be optimized with respect to the generation of state verification information. In other words, the computing unit evaluates which interaction is (presumably) most effective for collecting information such that gaps in the robot's knowledge of the object's state can be closed.

[0055] Advantageously, there is a combination of triggered interaction and acquisition of secondary monitoring data. For example, the robot first interacts in that it moves the object, e.g., rotates it, and then acquires secondary survey data, enabling the measurement of event-related features of the object that were previously blocked, i.e., inaccessible during the first survey, without which it would not be possible to unambiguously deduce the event from the first monitoring data.

[0056] Optionally, the state detector provides a correlation map relating the state derivation fuzziness to the interaction position and / or direction, similar to the correlation map described previously. Alternatively, the correlation map is based on a defined criterion representing an event and / or is established through machine learning, and / or the correlation map includes optimal interaction positions and / or directions for a plurality of derivable object states.

[0057] As another option, the interaction is based on a Markov model of the object's state.A Markov model is a stochastic model of a randomly varying system, where the system is described by a state.

[0058] As a further alternative or in addition, the triggered action is to obtain data from a database of monitored objects, in particular data about the objects, and / or third monitoring data from a third monitoring sensor of the monitoring system, the third monitoring sensor being not part of the robot, in particular wherein the triggered action comprises triggering the acquisition of third monitoring data ("on-demand monitoring").

[0059] Preferably, the robot is implemented as an unmanned ground vehicle (UGV). Optionally, the UGV includes an unmanned aerial vehicle (UAV) as a subunit, wherein the UAV is detachable from the main body (so that it can fly freely over the property) and has surveillance sensors, such as the first and / or second surveillance sensors.

[0060] The present invention also relates to a monitoring method for a facility monitoring system comprising at least a first monitoring sensor and a mobile monitoring robot. The method comprises the following steps: acquiring first monitoring data with the monitoring sensor; evaluating the first monitoring data and detecting a state of the property therefrom; determining a state ambiguity of the detected state, in particular by comparing it with a predetermined ambiguity threshold; and triggering an action of the robot in the event of state ambiguity, wherein the action is adapted to generate state verification information about the object, the state verification information being adapted to resolve the state ambiguity, and resolving the ambiguity based on the state verification information.

[0061] Furthermore, the present invention relates to a computer program product comprising a program code which is stored on a machine-readable medium or implemented by electromagnetic waves comprising program code segments, and which has computer-executable instructions which, when run on a computing device of a central computing unit of a monitoring system comprising a mobile monitoring robot, perform the steps of the method according to the invention.

[0062] The present invention provides a monitoring system having a measuring robot for inferring an event of a monitored object, which advantageously is capable of determining (un)ambiguous indicators for the inferred or detected state, and further capable of taking action if ambiguous information about the state of the object is determined, so as to collect additional information that can be used to resolve the ambiguity. In other words, the present invention allows for an autonomous evaluation or assessment of the observed state of the measured property and the automatic remediation or resolution of any uncertainty about the observed object.

[0063] A further advantageous embodiment enables the robot to plan its actions so that the step of collecting additional information for verification is performed in an optimized manner. The action that is finally triggered is one of at least two possible options that is more efficient for the generation of such verification information, for example so that the additional information is best suited to repairing the first monitoring data from which the uncertain state was derived. For example, the collection of additional sensor data (secondary monitoring data and / or sensor data in the context of the robot's interaction with the object in question, such as data from the robot's arm rotation encoder) is selected or planned so that the effort to generate a specific amount of data is minimized and / or the quantity or quality of the data optimally complements the existing data on the state or optimally remedies or repairs deficiencies in the existing data. Thus, the planning takes into account, for example, the choice of sensor, the choice of position and / or orientation / direction for measurement, or the interaction or selection of other variables or parameters of the (monitoring) sensor or action controller, such as the measurement accuracy of the robot actuator, the measurement time or the applied power.

[0064] A third aspect relates to a patrol system suitable for patrolling areas within a facility. For example, it can perform inspection tasks within buildings, construction sites, or other confined or potentially dangerous areas, such as industrial plants, airports, or container ships. In particular, the system is suitable for autonomously monitoring areas and detecting and reporting incidents.

[0065] Autonomous inspection tasks performed solely by unmanned ground vehicles (UGVs) face several challenges, such as:

[0066] - Restrictions on UGV movement, such as stairs, roadblocks, etc.

[0067] -UGVs can only inspect objects near the ground, not in the air; and

[0068] - UGVs in some environments (eg pits or canyons) face limited availability of GNSS connectivity for determining their position and / or for receiving / sharing data with, for example, a command center.

[0069] Automated inspection tasks performed solely by unmanned aerial vehicles (UAVs) face several challenges, such as:

[0070] - UAVs are subject to weight restrictions and therefore suffer from the following:

[0071] - Limited operating time due to battery weight limitations;

[0072] -Limited data storage and processing capabilities due to hardware limitations;

[0073] - Limited ability to carry payloads (such as sensors or actuators);

[0074] - UAVs cannot operate in adverse weather conditions, such as heavy rain or strong winds; and

[0075] - UAVs may face connectivity issues when communicating with, for example, a remote command center.

[0076] It is therefore an object of this aspect of the invention to provide an improved mobile surveillance system.In particular, the invention provides an object that combines the advantages of a UGV and a UAV.

[0077] According to this aspect of the invention, the patrol mobile surveillance system is thus adapted as a ground-air sensor platform that combines at least one UGV and at least one UAV. The combination of the UGV and the UAV can at least overcome the above limitations with the following capabilities:

[0078] A UGV can serve as a launch and recharging platform for one or more UAVs. The UGV can be equipped with a large battery or multiple battery packs to recharge or replace the UAV's batteries. This allows for extended UAV operating times. In inclement weather, the UGV can provide shelter from rain, wind, and other factors. Potentially, the UGV can also perform some exploration missions during periods when the UAV is unable to operate.

[0079] The UGV can act as a data storage and processing platform. The UGV provides a means for storing data captured by the UAV and / or a means for processing said data. This allows for more data to be stored / processed “at the edge” than when using only the UAV.

[0080] UGVs and UAVs can serve as joint sensor and / or actuator platforms. Considered from a complementary perspective, UGVs and UAVs can explore and measure objects. This allows, for example, combining ground-based and aerial views of an object. UGVs and UAVs can inspect the same object with different sensors, allowing, for example, a UAV to perform a first, rough, exploratory inspection of many objects in a large area, but using only low-resolution and a limited number of sensors. The UGV can then take over the measurement of selected objects of particular interest by conducting its own inspection using high resolution and a large number of different sensors. This results in sensor fusion, resulting from data generated by different sensors, from different viewpoints, under different conditions, and at different points in time.

[0081] For example, a UAV could have only an onboard human detection camera for detecting people with high uncertainty, and a UGV equipped with additional sensors (such as an infrared camera) could then be deployed to get close enough to film the person. It's also possible, for example, for the UAV to perform an inspection task that results in an action being performed by the UGV, or vice versa. A specific example of this concept would be to task the UGV with approaching a person and requesting identification (e.g., "asking for" an ID card) if the UAV has already detected the person.

[0082] During operation, the UGV and the UAV can make their sensor information available to each other in real time (or near real time), thereby achieving or improving their positioning. In addition, the UGV can track the UAV based on the video stream captured by the camera, that is, determine the direction from the UGV to the UAV. In cases where the positioning accuracy of the UAV is poor (because there are weight restrictions and only a very limited number of sensors can be carried for positioning) and the positioning of the UGV is good, directional observation can improve the accuracy of the UAV's position. The tracking system can also be equipped with a rangefinder that measures the distance between the UGV and the UAV. In this case, the direction and distance of the UAV relative to the UGV can be determined, and the relative position (X, Y, Z) relative to the UGV coordinate system can be calculated. For example, the UAV can fly forward and provide sensor data to generate a map for path planning of the UGV.

[0083] On the other hand, for navigation purposes, the UGV can be equipped with more sensors, such as LIDAR (Light Detection and Ranging) and perform, for example, LIDAR SLAM (Simultaneous Localization and Mapping), thereby navigating very accurately. Sending this information to the UAV, for example as a sparse point cloud, enables visual SLAM of the UAV.

[0084] In addition, UGVs and UAVs can be used as a joint communication platform. UGVs and UAVs communicate during work, enabling the deployment of new workflows (e.g., modified mission objectives and / or paths) from the UGV to the UAV, and vice versa. The UGV (or UAV) can receive / send data to, for example, a command center (if the UAV (or UGV) is unable to do so). Specifically, in a scenario with multiple UGVs (and possibly multiple UAVs), the UAV can move from one UGV to another, for example, to recharge and / or send data to the nearest UGV or the UGV with the most stable data link.

[0085] The term "unmanned ground vehicle" (UGV) should be understood not to be limited to vehicles that are located on solid ground and move, but also includes unmanned marine vehicles, such as unmanned surface vehicles (e.g., ships), unmanned underwater vehicles, or unmanned amphibious vehicles (e.g., hovercraft). The UAV can be, for example, a multi-rotor helicopter (e.g., a quadcopter) or a lighter-than-air aircraft (e.g., an airship).

[0086] Therefore, the present invention relates to a mobile surveillance system (suitable for patrolling a surveillance area, in particular a surveillance area of ​​a building), the system comprising a plurality of sensors, for example at least two cameras, wherein the system comprises at least one unmanned ground vehicle, which is adapted to move autonomously on the ground of the surveillance area, the UGV comprising a housing, in which are enclosed: a first battery; a first sensor device, for example at least a first camera, the first sensor device being adapted to generate first sensor data; and a first computing unit, the first computing unit comprising a processor and a data memory, the first computing unit being adapted to receive and evaluate the first sensor data in real time, wherein the system comprises at least one unmanned aerial vehicle, the UGV and the UAV being adapted to collaboratively patrol the surveillance area, wherein the UAV comprises a second sensor device, for example at least a second camera, the second sensor device being adapted to generate second sensor data, the UGV comprising a first data exchange module and the UAV comprising a second data exchange module, the first data exchange module and the second data exchange module being adapted to exchange data; and the first computing unit being adapted to receive and evaluate the second sensor data in real time.

[0087] Optionally, the mobile surveillance system includes multiple UAVs for patrolling the surveillance area, the first data exchange module is adapted to exchange data (251, 252) with the second data exchange module of each of the multiple UAVs, and / or the first computing unit is adapted to receive second sensor data from each UAV and evaluate the second sensor data from the multiple UAVs in a combined holistic analysis method.

[0088] As another option, the computing unit is adapted to generate UGV control data for controlling functions of the UGV, in particular in real time and based on an evaluation of the first sensor data, and / or the computing unit is adapted to generate UAV control data for controlling at least one UAV, in particular in real time and based on an evaluation of the first sensor data and / or second sensor data, wherein the second sensor data can be transmitted from the UAV to the UGV via the first data exchange module and the second data exchange module, and the UAV control data can be transmitted from the UGV to the UAV via the first data exchange module and the second data exchange module.

[0089] Optionally, the first computing unit is adapted to generate mission data, the mission data including instructions for causing at least one UAV to perform a mission or workflow, wherein the mission data can be transmitted from the UGV to the UAV via the first data exchange module and the second data exchange module, in particular, wherein the mission data includes instructions for immediately executing the mission.

[0090] As a further option, the mission data comprises instructions to move at least one UAV to a defined position, in particular to a position inaccessible to the UGV, and / or the mission data is generated based on an evaluation of the first sensor data and / or the second sensor data.

[0091] Optionally, the system is adapted to perform a state detection function, in particular fully autonomously, comprising detecting a state in the monitored area based on an evaluation of the first sensor data and / or the second sensor data. States include, for example, detecting the presence of an intruder in the monitored area or, more generally, any object-related state, such as its position, orientation, color, shape, properties (e.g., animate or inanimate), or operational state (e.g., open / closed, full / empty), etc., or anomalies, i.e., deviations from a usual situation or state. This means that not all states can be detected as a state, but only abnormal states, such as states or changes that were not previously monitored.

[0092] Optionally, the first computing unit is adapted to perform a state detection function based on an evaluation of the first sensor data, wherein, if a state has been detected, the computing unit generates task data, the task data comprising instructions to move the UAV to a position of the detected state and / or to generate second sensor data related to the detected state, wherein the instruction data is transmittable from the UGV to the UAV via the first data exchange module and the second data exchange module, and the second sensor data related to the state is transmittable from the UAV to the UGV via the first data exchange module and the second data exchange module.

[0093] Optionally, the UAV comprises a second computing unit comprising a processor and a data memory, wherein the second computing unit is adapted to receive and evaluate second sensor data from a second sensor device and to perform a state detection function based on the evaluation of the second sensor data, wherein, if a state has been detected, state data and / or second sensor data related to the state are transmitted to the first computing unit.

[0094] As another option, the first sensor device includes sensors with better specifications relative to the corresponding sensors of the second sensor device, in particular allowing the generation of first sensor data with a higher resolution than the second sensor data, wherein, if a state or anomaly has been detected, mission data is generated, the mission data including instructions to move the UGV to a position with the detected state or anomaly and / or generate first sensor data related to the detection.

[0095] As another option, the state detection function includes using at least one machine learning algorithm to train the state detection model, and / or providing at least one machine learning algorithm in the first computing unit and providing the state detection model to at least the UAV, wherein the first computing unit is adapted to run the machine learning algorithm to update the state detection model, wherein the updated state detection model can be transmitted from the UGV to the UAV via the first data exchange module and the second data exchange module.

[0096] Optionally, the first computing unit comprises a graphics processing unit and is adapted to run a machine learning algorithm on the graphics processing unit, and / or the data storage comprises a database for storing data related to the detected state.

[0097] As another option, the UGV is adapted to send data related to the detected event to a remote command center, wherein the data storage is adapted to store the data related to the detected event when a data connection to the remote command center is unavailable, and to send the stored data related to the detected event to the remote command center when the data connection is available.

[0098] Optionally, the first computing unit is adapted to perform a workflow generation process for generating a workflow for performing a patrol task in a monitored area, during which the first computing unit is adapted to generate an optimized workflow for performing the patrol task, which workflow involves one or more of the UAVs, so as to generate workflow data for each of the UAVs involved, which workflow data allows the corresponding UAV to perform part of the patrol task, and provides the workflow data to the UAVs involved via the first data exchange module and the second data exchange module.

[0099] Optionally, the first computing unit is adapted to request and receive mission-specific data of at least one UAV via the first data exchange module and the second data exchange module, wherein the mission-specific data includes information about characteristics associated with the corresponding UAV (location and / or workload information), so as to evaluate mission-specific capabilities associated with each of the UAVs based on the mission-specific data, and generate an optimized workflow based on the patrol mission and the mission-specific capabilities.

[0100] As another option, the first computing unit is adapted to monitor the condition of the UAVs involved, including the ability of the UAVs to perform corresponding parts of the patrol mission, wherein, if the first computing unit determines that one of the UAVs involved has lost the ability to perform its part of the patrol mission, the first computing unit is adapted to generate an adapted workflow, comprising at least: reassigning the affected part of the patrol mission to one or more UAVs of the multiple UAVs so as to generate adapted workflow data for one or more UAVs of the multiple UAVs involved, and providing the adapted workflow data to one or more UAVs of the multiple UAVs via the first data exchange module and the second data exchange module.

[0101] Optionally, each of the at least one UAV is equipped with a software agent, wherein each software agent is installable on a computing unit of the UAV or on a communication module connected to the UAV and is adapted to exchange data with the UAV to which it is installed or connected, wherein, during the course of the workflow generation process, the first computing unit is adapted to request and receive task-specific data of the UAV from the software agent of the corresponding UAV and to provide workflow data to the software agent of the UAV concerned.

[0102] Optionally, the UGV and the at least one UAV are adapted to autonomously patrol the surveillance area, and / or the data storage is adapted to store a map of the surveillance area and the UGV is adapted to navigate through the surveillance area based on the map, and / or the first computing unit comprises a SLAM algorithm, which performs real-time localization and map construction based on the first sensor data and / or the second sensor data, in particular wherein the sensors comprise at least one LIDAR scanner, in particular wherein the SLAM algorithm is adapted to continuously update the map of the surveillance area, wherein the updated map is transmittable to the UAV via the first data exchange module and the second data exchange module.

[0103] Optionally, in order to collaboratively patrol the surveillance area, the UGV is adapted to move along a predetermined path, in particular along a sequence of defined waypoints; and the at least one UAV is adapted to explore the area around the UGV, moving within a maximum range around the UGV, in particular wherein the maximum distance is user-defined, depends on the maximum speed of the UAV and the patrol speed of the UGV when moving along the predetermined path, and / or is given by a requirement to exchange data via the first data exchange module and the second data exchange module.

[0104] Optionally, the first computing unit includes an algorithm for encrypting and decrypting data, and the first data exchange module and the second data exchange module are adapted to exchange encrypted data and / or the exchanged data include heartbeat messages, wherein the UGV and / or UAV is adapted to identify whether data exchange between the UGV and the UAV is available based on the heartbeat message, in particular, when it is identified that data exchange between the UGV and the UAV is not available, the UGV and / or UAV is adapted to change its position in order to reestablish a connection for data exchange, in particular to return to a previous position where data exchange was still available.

[0105] Optionally, the system comprises one or more radio communication modules adapted to establish a radio connection with a remote command center, in particular by means of a WiFi network or a mobile phone network, and adapted to send and receive data via the radio connection, and / or the system can operate from the remote command center, in particular in real time, by means of data sent to the radio communication modules of the system.

[0106] Optionally, the UAV includes a first radio communication module and is adapted to wirelessly transmit data received via the radio connection to the UGV, and / or the UGV includes a second radio communication module, and the UAV is adapted to wirelessly transmit data received via the radio connection to the UGV when the second radio communication module cannot establish a radio connection with the remote command center.

[0107] Optionally, upon identifying that the wireless connection is unavailable, the UGV and / or UAV is adapted to change its position in order to re-establish the wireless connection, in particular to return to a previous position where the radio connection was still available.

[0108] The present invention also relates to a mobile surveillance system adapted to patrol a surveillance area, in particular a building, the mobile surveillance system comprising: one or more radio communication modules adapted to establish a radio connection with a remote command center for sending and receiving data; and a plurality of sensors, for example comprising at least two cameras, wherein the system comprises at least one unmanned ground vehicle (UGV), at least one UGV being adapted to move autonomously on the ground of the surveillance area; the UGV comprising a housing enclosing: a first battery; and a first sensor device, for example comprising at least one first camera, the first sensor device being adapted to generate first sensor data, wherein the system comprises at least one unmanned aerial vehicle (UAV), the UGV and the UAV being adapted to collaboratively Patrolling a surveillance area, wherein the UAV comprises a second sensor device, such as at least one second camera, the second sensor device being adapted to generate second sensor data, the UGV comprises a first data exchange module, and the UAV comprises a second data exchange module, the first data exchange module and the second data exchange module being adapted to exchange data, the remote command center comprises a remote computing unit, the remote computing unit comprises a processor and a data storage, the remote computing unit being adapted to receive the first sensor data and the second sensor data via a radio connection, evaluate the first sensor data and the second sensor data in real time, and generate mission data, the mission data comprising instructions for causing at least one UGV and the UAV to perform a mission or workflow, wherein the mission data is sent via the radio connection.

[0109] Optionally, the UGV and the UAV are connected by a cable (with or without a plug), wherein the first battery is adapted to supply power to the UAV via the cable; and the first data exchange module and the second data exchange module are adapted to exchange data through the cable.

[0110] As another option, the UAV includes a second battery, wherein the first battery has a larger capacity than the second battery; the UAV is adapted to land on the UGV; and the UGV includes a charging station and / or a battery swap station, the charging station being adapted to charge the second battery when the UAV lands on the UGV, and the battery swap station being adapted to automatically replace the second battery when the UAV lands on the UGV.

[0111] Optionally, the UGV includes at least one robotic arm, which is capable of being controlled by a computing unit based on the evaluated first sensor data, wherein the computing unit is adapted to control the robotic arm to interact with multiple features of the environment, specifically including: opening and / or closing doors; operating switches; picking up objects; and / or positioning the UAV on a charging station.

[0112] Optionally, the UGV includes a plug connected to a power socket, wherein the computing unit is adapted to detect a power socket in the environment based on the first sensor data; and the UGV is adapted to insert the plug into the power socket, in particular by means of a robotic arm, and / or the UAV is a quadcopter or other multi-rotor helicopter, the helicopter comprising: a plurality of rotors, in particular at least four rotors; and a base comprising a plurality of skids or legs enabling the UAV to land and stand on the UGV.

[0113] Optionally, the UAV uses a lifting gas, particularly helium or hydrogen, to provide buoyancy, particularly wherein the UAV is an airship, and particularly wherein the UGV comprises a gas tank and a filling station adapted to refill the UAV with lifting gas.

[0114] Optionally, the charging station comprises at least one induction coil, wherein the charging station and the UAV are adapted to charge the second battery by induction.

[0115] Optionally, the mobile surveillance system includes at least a first UGV and a second UGV, wherein the UAV is adapted to land on the first UGV and the second UGV and to wirelessly exchange data with the first UGV and the second UGV.

[0116] As another option, the sensor includes at least one GNSS sensor for use with a global navigation satellite system, the at least one GNSS sensor being arranged in the UGV and / or the UAV, and / or the at least one GNSS sensor being arranged in the UAV, wherein the signal provided by the GNSS sensor can be wirelessly transmitted from the UAV (220) to the UGV, in particular as part of the second sensor data.

[0117] Optionally, the UGV includes: a space for accommodating one or more UAVs, the space providing protection from precipitation and / or wind, in particular wherein the UGV includes an extendable drawer on which the UAV can land, the drawer retractable into the housing of the UGV to accommodate the UAV within the housing; or a cover adapted to cover the UAV when the UAV is landed on the UGV.

[0118] Another fourth aspect of the present invention relates to a safety monitoring system for detecting the status of a facility.

[0119] In security monitoring systems, states are proposed by state detectors, i.e., using monitoring sensors such as person detectors, anomaly detectors, or special classifiers trained for security applications such as door or window opening detectors. These states can then be filtered by state filters, which determine whether the state is critical, non-critical, or previously unknown. States are, for example, the position, orientation, color, shape, properties (e.g., animate or inanimate), or operating state (e.g., open / closed, full / empty) of an object, or an anomaly, i.e., a deviation from the usual situation or state. This means that not every state can be recorded, but only abnormal states, e.g., states that have not been previously monitored.

[0120] Surveillance sites can have their own regularities or characteristic workflows and include various entities, such as human workers or machines. For example, workers may have specific access rights to different areas, buildings, or rooms of the surveillance site, while visitors may only have limited access rights. Furthermore, depending on the type of site, such as a construction site or a restricted military site, the surveillance site can include potentially dangerous areas, where people and / or machines could be injured or potentially violate laws, and therefore require warnings when entering these areas. Furthermore, open sites, in particular, are subject to diverse environmental influences, not to mention that the overall topology of different sites can vary significantly.

[0121] In other words: It is often difficult to provide a unified monitoring system for different facilities, because some statuses may only be relevant for a specific monitoring location, for example a monitoring location with a specific workflow or local characteristics, such as certain personnel who tend to start very early in the morning or work in the evening.

[0122] Therefore, on-site monitoring still requires human eyes and judgment. For example, algorithms typically alert operators to specific conditions, and operators decide whether the condition is truly critical. The method of choice is usually classification or anomaly detection, which alerts operators in the event of suspicious events. The operator then either ignores irrelevant alarms or takes action, such as calling colleagues, the police, or the fire department. Consequently, the same software often runs at different local locations, with operators adapting to the specific rules of each location.

[0123] There are many computational challenges faced when trying to automate this adaptation process through computer-implemented solutions.

[0124] It is therefore an object of the present invention to provide an improved security monitoring system that overcomes the above limitations.

[0125] A specific object is to provide a system and method which allows for a more versatile and robust monitoring and alarm system, wherein false alarms are reduced and only increasingly relevant alarms are brought to the attention of the operator.

[0126] These objects are achieved by realizing the features of independent claim 79. Features which further develop the invention in an alternative or advantageous manner are described in the patent claims dependent on claim 79.

[0127] The present invention relates to a safety monitoring system for detecting the status of a facility, such as an indoor and / or outdoor location, such as a building, a warehouse, an industrial complex, a construction site or another restricted or potentially dangerous area.

[0128] For example, such a security monitoring system may include at least one monitoring sensor configured to, in particular, continuously monitor a facility and generate monitoring data comprising information about the facility. In particular, such monitoring data may provide information about the status of a specific portion of the facility, for example, wherein the information may be provided substantially directly, i.e., based on a single sensor, in particular without the need for data processing, and / or two or more sensors may generate the monitoring data based on a working relationship.

[0129] For example, the monitoring sensors may include cameras, microphones, contact sensors, motion detectors, smoke detectors, key readers, galvanometers, RIM cameras, laser scanners, thermometers, but may also include input devices suitable for allowing an operator to input survey data. Optionally, one or more monitoring sensors are mobile sensors, particularly as part of a mobile robotic platform.

[0130] The system further includes a state detector configured to detect at least one state based on monitoring data of the facility, wherein the at least one state represents at least one state associated with the facility.

[0131] For example, according to one embodiment, the security monitoring system is configured to detect sequences and / or patterns of states associated with the facility, wherein typical states may be detection and / or identification of a person, detection of opening a door and / or window, detection of fire and / or smoke, detection of abandoned objects, identification of activities, and detection of anomalies. A state may represent a change in the position or orientation of a component of the facility, such as the opening of a door or window, a change in appearance (including a change or change in an object associated with the facility (e.g., an object is introduced into or taken out of a room, a person enters a room), or any significant difference in the monitoring data when comparing monitoring data generated at two different moments or time periods. For example, the state detection device may be integrated in the monitoring sensor and / or the central computing unit. In particular, the state detection may take into account the topological and / or logical and / or functional relationships of different entities associated with the facility, as well as a timetable representing the time of human and / or automatic activities associated with the facility, in particular wherein the timetable comprises planned work and / or working hours and / or comprises information about activity types and / or is embedded in a digital model of the facility.

[0132] In addition, the system includes a status filter, which is configured to perform automatic assignment, automatically assigning at least one status to a category of status, wherein the status filter is configured to assign at least one status to at least a first category indicating a status of "non-critical" and a second category indicating a status of "critical".

[0133] According to the present invention, the security monitoring system includes a feedback function configured to generate feedback information regarding the automatic assignment, the feedback information indicating confirmation or denial of at least one of the automatic assignment of at least one state to the first category or the second category and the manual assignment of at least one event to the first category or the second category. Furthermore, the security monitoring system includes a training function, particularly a training function based on a machine learning algorithm, configured to receive and process the feedback information and provide update information for the state filter, wherein the update information is configured to establish the first category and the second category.

[0134] Thus, the automatic assignment is dynamically improved and, for example, only requires a "coarse" initialization based on less prior information about the facility. In particular, the system dynamically adapts to the specific properties of the facility.

[0135] It may happen that a condition cannot be assigned to either the first or second category. In particular, it may be difficult for an operator to determine whether a condition is critical, for example if the condition has never occurred before or if the operator doubts the quality of the monitoring data. The operator can then manually request the retrieval of at least a portion of the monitoring data associated with the event, or request additional monitoring data using additional measurement sensors, such as by deploying a mobile sensor to the area.

[0136] Thus, as an example, other classes may exist, for example, wherein the state filter is configured to assign at least one state to a third class representing a state of "uncertain." In particular, the feedback function may be configured to trigger reacquisition of at least part of the survey data and / or acquisition of additional survey data if at least one state is assigned to the third class.

[0137] Different responses can be triggered based on different levels of assigned certainty. For example, if a state is critical or non-critical with high certainty, the system takes automatic action or ignores the state. However, if the assigned certainty level is low, and / or if the state filter cannot assign a previously unknown state, the operator must decide and take action, such as calling the police or fire department or ignoring the event. According to the present invention, the system dynamically learns new critical or non-critical events, thereby reducing false alarms and bringing only increasingly relevant alarms to the operator's attention.

[0138] Alternatively, all critical conditions (even with high certainty) may have to be acknowledged by the operator before an action is triggered.

[0139] In addition to a class representing "non-critical" events for each "critical" state, separate subclasses can be defined, for example, "fire", "intrusion", etc.

[0140] Taking into account the relationship of the measured components of a facility allows for enhanced detection of states associated with these components, for example in order to better link the measured states of two adjacent components together to form a state. Monitoring based on state sequences and / or patterns has the advantage that it has more aspects that can be examined by the state detection algorithm, resulting in a more in-depth assessment of the classification. In other words, instead of considering a single state in isolation, the context of the state or state string is considered, so that its classification as "critical" or "non-critical" can be better evaluated. If, for example, a single component is measured as part of a state sequence, it can be verified whether the subsequent state of the same or another component that was expected due to experience (machine learning) and / or due to logical, functional or topological relationships was indeed measured. The use of state sequences and / or patterns allows, for example, to consider the behavior of people towards the facility represented by the state sequences and / or patterns. Therefore, the present invention allows the system to learn how normal or to what extent the behavior or process of a facility is.

[0141] According to another embodiment, the automatic assignment is based on an n-dimensional feature space, wherein a state (e.g., "person detected" plus metadata such as time, location, etc.) is represented by an n-dimensional feature vector, and in particular, wherein the corresponding class is represented by a portion of the n-dimensional feature space or a neural network. In particular, the detected states representing sequences and / or patterns of changes associated with the facility are fed into different units of the input layer of the neural network.

[0142] In another embodiment, a security monitoring system includes a user interface configured to receive input from an operator of the security monitoring system and generate a feedback signal carrying feedback information based on the input. In particular, the feedback function is configured to interpret the absence of input received at the user interface within a defined time period as an automatically assigned confirmation or negation. The user interface is optionally provided by a mobile device (e.g., a smartphone application) that is part of the security monitoring system and processes the operator input and transmits the data between the smartphone and a computing unit of another system. This enables remote (feedback) control of the monitoring system.

[0143] By way of example, the user interface is configured such that the operator of the security monitoring system can manually cancel and / or confirm the automatic assignment, manually assign at least one state to a state category different from the category to which the at least one state was automatically assigned, generate a new state category, in particular by splitting an existing category into at least one of at least two new categories, for example, "critical" to "high criticality" and "low criticality", merge multiple existing classes into a new class, modify existing classes, and delete existing classes.

[0144] According to another embodiment, the security monitoring system is configured to provide a command output, in particular, to execute a warning or safety measure, when at least one state has been automatically assigned to at least one of the second and third categories. Similarly, such a warning can be sent to a mobile device such as a smartphone.

[0145] In another embodiment, the security monitoring system is configured to have a release function, the release function being configured to prevent the command output during at least a defined waiting period, wherein the release function is configured to at least one of: release the command output based on a release signal from a user interface; release the command output at the end of the defined waiting period; and delete the command output based on a stop signal from the user interface. For example, the security monitoring system can be configured to interpret the absence of a release signal during the waiting period as a confirmation of the automatic assignment and / or to interpret the deletion of the command output as a negation of the automatic assignment through the training function.

[0146] Different facilities have specific patterns, and a critical condition at one location may not be worth mentioning at another. Therefore, monitoring locations and assigning conditions to different categories and subcategories still requires a time-consuming process, human eyes, and human judgment.

[0147] The application of machine learning algorithms allows for the automation of various processes involved in classifying measurement data. Compared to rule-based programming, this classification framework, a subcategory of general machine learning (ML), provides a highly effective "learning approach" for pattern recognition. Machine learning algorithms can handle highly complex tasks, utilize implicit or explicit user feedback, thus becoming adaptive, and providing "per-point" probabilities for classification. This saves time, reduces processing costs, and reduces manual effort.

[0148] In so-called “supervised ML,” the algorithm implicitly learns which representational attributes (i.e., combinations of features) define target properties of an event (such as class membership, affiliation to a subclass, etc.) based on definitions made by the user when labeling the training data.

[0149] On the other hand, in so-called "unsupervised ML," algorithms find hidden structure in unlabeled data, for example by finding groups of data samples that share similar properties in feature space. This is called "clustering" or "segmentation."

[0150] Probabilistic classification algorithms also use statistical reasoning to find the best class for a given instance. Instead of simply determining the "best" class for each instance, probabilistic algorithms provide the probability of the instance being a member of each possible class, typically selecting the class with the highest probability. This has several advantages over non-probabilistic algorithms, namely, associating a confidence value to weight its selection and thus providing the option to discard a selection if the confidence value is too low.

[0151] The use of machine learning algorithms requires large amounts of training data. In the case of supervised machine learning, labeling information is also required (i.e., assigning object classes to the data). Data acquisition, preparation, and labeling require significant effort and time. However, by storing and feeding back operator information and decisions (especially those of many operators across many different facilities) to the underlying algorithms in an active learning setting, the algorithmic models are improved and gradually adapt over time.

[0152] Therefore, according to another embodiment of the present invention, the training function is based on at least one of the following: linear classification, in particular based on Fisher linear discriminant, logistic regression, naive Bayes classifier or perceptron; support vector machine, in particular least squares support vector machine; quadratic classifier; kernel estimation, in particular k-nearest neighbor; boosting; decision tree, in particular decision tree based on random forest; hidden Markov model; deep learning, in particular based on neural network, in particular convolutional neural network (CNN) or recurrent neural network (RNN); and learning vector quantization.

[0153] Furthermore, a monitoring system installed locally at a particular facility may be part of an extended network of many local monitoring systems operating at multiple different monitoring locations, with each local security monitoring system providing updated information to the global event detection algorithm and / or global event classification model.

[0154] On the one hand, such a global algorithm and / or model (e.g., a global algorithm and / or model stored on a central server) can then be used to initialize the local monitoring system installed at a new monitoring site. On the other hand, the global algorithm and / or model can be updated by operator feedback similar to the local model.

[0155] Thus, during the working hours of the local system, in the event that an event is unknown to the local model, the global state detection and / or filtering model can be queried by the local security monitoring model before prompting the operator to make a decision. Thus, the local model can utilize information collected at other facilities.

[0156] Therefore, in another embodiment, the security monitoring system includes at least one further state filter, specifically foreseen for use on a different facility, and configured to detect at least one state associated with the facility and automatically assign the state to at least a first category representing a "non-critical" state and a second category representing a "critical" state. Furthermore, the security monitoring system includes a common classification model comprising a first common class representing a "non-critical" state and a second common class representing a "critical" state, wherein the security monitoring system is configured to initially establish, for each state filter, a respective first category and a respective second category based on the first common class and the second common class. Furthermore, the security monitoring system is configured to generate, for each state filter, respective feedback information regarding the respective automatic assignments, each respective feedback information indicating at least one of: confirmation or denial of the respective automatic assignment of the respective at least one state to the respective first category or the respective second category; and manual assignment of the respective at least one state to the respective first category or the respective second category.

[0157] Therefore, according to this embodiment, the security monitoring system is configured to provide corresponding update information for each (local) state filter, wherein the corresponding update information is configured to (locally) establish a corresponding first category and a corresponding second category, and the security monitoring system is configured to: if relevance to the common classification model is identified, provide at least a portion of each of the (local) corresponding feedback information and / or at least a portion of each of the (local) corresponding update information as common update information to the common classification model, so that the first common category and the second common category are established based on the common update information. In other words, the system evaluates whether the additional feedback or update information is meaningful for the common model, for example based on known common characteristics of different facilities, and if so, provides it as common update information.

[0158] It may be beneficial to update the public classification model based on operator filtering, ie where the operator decides to take over the updating / feedback to the public classification model. Alternatively, the updating may be performed automatically, eg based on rules, in particular based on consistency checks.

[0159] Furthermore, local state filters can be dynamically updated based on updated public classification models, i.e., representing global state filters. For example, a local operator familiar with local prerequisites can decide which differences between the global event filter and the local state filter should be incorporated into the local event filter. Alternatively, this update can be performed automatically, for example based on rules, particularly consistency checks.

[0160] Therefore, in another embodiment, the security monitoring system includes a first filter function configured to feedback information and / or a selection from the update information of each state filter and provide the selection as common update information to the common classification model.

[0161] In another embodiment, the security monitoring system is configured to update each state filter based on at least a portion of the public update information, i.e., for each event filter, a corresponding first category and a corresponding second category are established based on the public update information. In particular, the security monitoring system includes a second filter function configured to select from the public update information and provide the selection to at least one of the state filters.

[0162] In another embodiment, local and / or global state detection algorithms can also benefit from operator feedback, and thus, similar to improvements to event filtering, state detection can also become more sophisticated over time. Thus, the feedback loop can work in a dual manner, namely, providing feedback to the state filter to answer questions such as "is the state critical?" and "is action required?", and providing feedback to the state filter to answer questions such as "is the event triggered correct?" and "is human detection working correctly?".

[0163] Therefore, according to this embodiment, the security monitoring system is configured to provide feedback information corresponding to one of the state filters, update information corresponding to one of the state filters, and at least a portion of one of the public update information to the event detector as detector upgrade information, that is, the (local) state detector detects future events based on the detector upgrade information. For example, the local event detector can be automatically upgraded based on the detector upgrade information, for example, once new detector upgrade information is available or based on a defined upgrade interval.

[0164] The present invention also relates to a computer program product comprising a program code stored on a machine-readable medium or embodied by an electromagnetic wave comprising a program code segment, and having computer-executable instructions for performing the following steps, in particular when run on a computing unit comprising a state detector and / or a local state filter of a security monitoring system according to the present invention: detecting at least one event based on survey data of a facility, wherein the at least one state represents at least one change associated with the facility; automatically assigning the at least one state to a class of states, wherein at least two classes of states are defined, namely, a first class representing a state of "non-critical" and a second class representing a state of "critical"; processing feedback information about the automatic assignment, the feedback information indicating confirmation or negation of the automatic assignment of the at least one state to the first class or the second class and manual assignment of the at least one state to at least one of the first class or the second class; and providing update information for performing the automatic assignment, wherein the update information is configured to establish the first class and the second class.

[0165] In an automatic surveillance system as described in this paper, one problem is that automatic detection of objects or people relying solely on RGB or IR image information from a camera is not robust enough, especially under challenging conditions such as at night or in unfamiliar lighting situations, in unusual postures, in non-stimulating situations, etc.

[0166] Therefore, some previous works have proposed using multiple signal sources to solve the object detection problem. For example, regarding the task of person detection, the paper [1]: Spinello and Siegwart's "Human Detection Using Multimodal and Multidimensional Features", ICRA 2008 proposed the fusion of laser ranging and camera data. Based on clustering based on the classification of geometric features in the range data space, image detection (HOG features + SVM) is performed on the restricted image area defined by the clusters obtained from the ranging data. This assumes that the calibration parameters of the sensor are available. However, since the ranging data is used to define the search space of the image-based classifier, objects outside the range cannot be simply detected.

[0167] Alternatively, the paper [2]: “Human Detection in RGB-D Data” by Spinello and Arras, IROS 2011 proposes a fusion of RGB data and depth data obtained from a Kinect-device. The resulting depth data is robust to illumination variations but is sensitive to low signal strength returns and suffers from limited depth resolution. The resulting image data is rich in color and texture and has high angular resolution, but it quickly decomposes under non-ideal lighting conditions. This method exploits both modalities thanks to a principled fusion method called Combo-HOD that relies neither on background learning nor on a ground plane assumption. The proposed person detector for dense depth data relies on a novel HOD (Histogram of Oriented Depth) feature. Combo-HOD is trained separately by training a HOG detector with image data and a HOD detector with depth data. The HOD descriptor is computed in the depth image, while the HOG descriptor is computed in the color image using the same window that requires calibration. When no depth data is available, the detector gracefully degrades to a conventional HOG detector. When the HOG and HOD descriptors are classified, they are ready to be fused using informative filters.

[0168] Paper [3]: “Multi-modal Person Localization and Emergency Detection Using the Kinect” (IJARAI 2013) by Galatas et al. proposes a combination of depth sensors and microphones to locate people in the context of an assisted intelligent environment. Two Kinect devices are used in the proposed system. The depth sensor of the primary Kinect device is used as the main source of positioning information due to skeletal tracking (using the MS Kinect SDK). People are tracked while standing, walking, or even sitting. Sound source localization and beamforming are applied to the audio signal in order to determine the angle of the device relative to the sound source and acquire the audio signal from that specific direction. However, one Kinect can only provide the angle of the sound source but not its distance, which hinders the localization accuracy. Therefore, a second Kinect is introduced and used only for sound source localization. The final estimated position of the person is the result of combining the information from both modules. In addition to the position information for emergency detection, automatic speech recognition (ASR) is also used as a natural means of interaction.

[0169] The paper [4]: ​​“On the Use of a Low-Cost Thermal Sensor to Improve Kinect People Detection in a Mobile Robot” by Susperregi et al., Sensors 2013 describes the fusion of visual cues, depth cues and thermal cues. The Kinect sensor and the thermopile array sensor are mounted on a mobile platform. The achieved false alarm rate is significantly reduced compared to using any single cue. The advantage of using thermal morphology is that there are no large differences in appearance between different people in thermal images. In addition, infrared (IR) sensor data does not depend on lighting conditions and people can also be detected in completely dark conditions. As a disadvantage, some phantom detections may occur near heat sources such as industrial machines or radiators. After building a single classifier for each cue, these classifiers are assembled hierarchically so that their outputs are combined in a tree pattern to obtain a more robust final person detector.

[0170] GB 2546486 shows a multimodal building-specific abnormal event detection and alarm system with a plurality of different sensors.

[0171] KR 101125233 and CN 105516653 mention safety methods based on the fusion technology of visual signals and radar signals. CN 105979203 shows multi-camera fusion.

[0172] Therefore, it is an object of a fifth aspect of the present invention to provide an improved automatic monitoring system. In particular, it is an object to provide such a system that automatically detects critical events, in particular with a low number of false positives and missed negatives due to erroneous detection under non-optimal sensing conditions. It is also an object to provide a system that can substantially automatically adapt to a specific environment, preferably with a lower on-site training workload than the prior art. It is also an aspect that the system can be machine-learned, with substantially lower human interaction than existing methods, in particular while still being specifically trained for a specific building or any other monitoring location.

[0173] It is also an object to provide a corresponding monitoring system that is able to handle complex multi-sensor and multi-modality monitoring applications, in particular wherein the contextual correlation of the modalities can be derived automatically with substantially low manual programming effort.

[0174] These objects are achieved by realizing the features of independent claims 93, 104 and 107. Features which further develop the invention in an alternative or advantageous manner are described in the patent claims dependent on these independent claims.

[0175] This aspect of the invention (which itself may also be construed as a specific invention) relates to the idea of ​​combining multiple sensors and / or modalities (e.g., like color, depth, thermal imaging, point clouds, etc.) in an intelligent manner for object detection in surveillance systems. Considered independently of each other, each modality has its own strengths. For example, prior art image-based detectors rely on the rich color and texture information present in RGB images. However, their performance is known to degrade under non-ideal lighting conditions, especially in darkness. In those non-ideal situations, supplementary information (such as depth maps, point clouds (PCs), and / or infrared (IR) signals) can be used to determine the presence, absence, classification, status, etc. of an object of interest in a scene where classic image detectors degrade or fail.

[0176] As shown in the above-mentioned prior art, the known systems therein focus on:

[0177] Apply detection in each of these modalities separately and independently, and then combine the results. So-called parallel approaches, for example by fusing the output of an RGB detector with the output of a depth detector, as in US 8,630,741, US 2014 / 320312, GB 2546486, [2] or others; or

[0178] Linking detectors for each modality. So-called serial or hierarchical approaches, for example, start by first clustering objects in the odometry data space and then use this clustering information to guide a second-stage visual detector to the found regions of interest, as described in KR 101125233, [1] or elsewhere.

[0179] Another approach, such as proposed in US 2009 / 027196, [4] or elsewhere, attempts to learn the best possible detector hierarchy from the data using a tree structure. Thus, a single general model is learned.

[0180] In contrast, this aspect of the invention proposes learning context-aware detectors that can follow a parallel or serial architecture, particularly a hybrid of selective learning of subsets of those learned on real-world and / or synthetically generated training data. A particular aspect can be seen in the introduction of contextual learning in detection to leverage the strengths of each modality, which is done based on the context provided and used in training and learning. The context, which can specifically include environmental information, can include, for example, location, image region, time of day, weather information or forecast, etc.

[0181] Therefore, the present invention is not just a simple independent side-by-side use or layered monitoring sensor method of the prior art, which can overcome many shortcomings of the prior art and can improve the overall detection rate and reduce the false alarm rate, especially considering at least partially contradictory monitoring sensor results.

[0182] This aspect of the invention relates to an automated monitoring system for automatically detecting conditions, particularly critical conditions, abnormalities, or anomalies, at a facility. "Facility" is broadly defined as not only a building, for example, but also an entire property or installation, such as an industrial plant or other complex facility, such as a (large) ship or aircraft. In particular, in mobile and / or fixed automated monitoring systems as described herein, for example in combination with one or more aspects of other embodiments. This state detection can be implemented, for example, based on information from a mobile patrol system adapted to patrol an area, such as a building to be inspected, but can also be implemented based on information from fixed monitoring equipment, such as surveillance cameras, various intrusion sensors (e.g., light, noise, shattering, vibration, thermal imaging, motion, etc.), authentication terminals, personal tracking devices, RFID trackers, distance sensors, light barriers, etc. In particular, a combination of fixed and autonomous monitoring of an area for automatically detecting and reporting abnormal and / or critical conditions can be implemented, for example, adapted to a specific environment.

[0183] The system includes a plurality of monitoring sensors, at least two of which operate in different modes. Thus, there is at least one first monitoring sensor for at least one first modality and at least one second monitoring sensor for at least one second modality, thereby sensing at least two modalities. The term "modality" herein may specifically refer to the sensing of different physical properties sensed by the sensor, such as sensing visible light radiation, invisible light radiation, audible sound waves, inaudible sound waves, distance, range, temperature, geometry, radiation, vibration, shock, earthquake, etc. The monitoring sensor may be adapted to monitor at least a portion of a property or building and to generate real-world monitoring data in these multiple modalities.

[0184] In an automated monitoring system, anomaly detection is derived from a combination of one or more sensings from those multiple monitoring sensors. This can be implemented, for example, by a computing unit designed for evaluation of monitoring data, preferably comprising a machine learning evaluation system that includes automatic detection of anomalies and / or classification of those anomalies (e.g., as potentially critical or non-critical) or derivation of their status patterns, such as a machine learning classifier unit and / or a detector unit.

[0185] According to this aspect of the invention, a combination of multiple monitoring sensors is provided by a machine learning system (e.g., implemented in a computing system, such as in an artificial intelligence unit), for example, by supervised training of the artificial intelligence system, which is trained at least in part based on training data including contextual information of the property or building.

[0186] In anomaly detection during normal use of the monitoring system, the combination is selected from at least a subset of monitoring sensors based on a machine-learned dependency on contextual information. The contextual information may include real-world sensor measurements from a dedicated context sensor, real-world sensor measurements derived from an evaluation of details from one or more monitoring sensors of the same or another modality, or real-world sensor measurements derived from an abstracted virtual state model of a building or property in the automated monitoring system, including, for example, power status, lighting status, heating status, occupancy status, etc.

[0187] In other words, the present invention relates to an embodiment of a machine learning combination of multiple sensors of different modalities in an automatic monitoring system. Learning can be based on training data including contextual information, which contextual information can, for example, lead to a classifier that classifies the monitoring sensor modalities into different context categories. During monitoring, the actual field environment is taken into account and the sensed modalities are then selected for evaluation in a learning manner. The present invention thus learns contexts in which different types of detectors with different modalities are most suitable. For example, with respect to contexts such as spatiotemporal segments or clusters, lighting, temperature, etc., the best detector and / or classifier combination is learned from the training data for these contexts. Based on this learned context-specific best detector and / or classifier combination, the security system can obtain automatic detection of states or state patterns and / or classification of critical states based on real-world survey data from multiple measurement sensors with understanding of the contextual conditions. There can be at least one dedicated context sensor for obtaining information about the physical contextual conditions of a building or property or the internal state of the monitoring system.

[0188] In addition to the binary selection of one or more measurement sensors to be evaluated, the results from multiple measurement sensors can also be weighted, resulting in a combination that serves as the basis for determining the security state. Thus, the training results in the classification of measurement sensor availability can depend on contextual conditions, such as environmental conditions, for example, by machine learning classification of strategies for evaluating one or more measurement sensors based on content conditions.

[0189] According to the present invention, it is possible to learn how and when to interpret surveillance sensors in a specific context, such as when certain environmental conditions apply, or the spatial context in which a particular classifier must be applied (which may not be applicable in another context). For example, people should only be detected on floors, while open windows or doors are not detected at the floor level, but are generally substantially vertical. Or, for example, it can be learned that visual images are primarily evaluated under certain lighting conditions, such as in the context of daytime or in the environmental context of lights on. Another example is that IR images are primarily used when the lights are off, while depth images can be used regardless of the lighting context. Another example can be the context of the occupancy state of the monitored room, such as the context of the occupancy state known to the surveillance system by tracking people entering and exiting or by mapping their trajectory. By providing this occupancy context information in the training data, the system can, for example, learn that presence detectors or motion sensors can provide useful security information to issue a security alarm if the room is in an "empty" occupancy state, but their information is not very useful if the room is occupied by one or more people. According to the present invention, the monitoring system can learn such sensor evaluation strategies from training data and does not require manual programming and modification of complex correlations for each case and site to reflect complex, large, and site-specific conditions.

[0190] The combination can be selected from at least a subset of monitoring sensors using a modality-specific machine-learned weighting function, particularly where the weighting function can include modality-specific weighting factors that depend on contextual information. An information filter can then derive one or more safety states based on the contextually weighted results from each of the plurality of monitoring sensors.

[0191] The weighting factors may be clustered, for example, by a classifier that is machine-learned using training data including contextual information.

[0192] The classifier or detector can also be trained at least in part with training data containing contextual information, the training data being at least in part synthetically generated and derived from the virtual model, in particular as described in detail elsewhere in this document. The modalities can, for example, include at least a visible image modality, an infrared modality, and a depth modality. In one embodiment, multiple modalities from different sensors can be combined into a single dataset, with which the machine learning system is trained separately and then evaluated at runtime by the monitoring system. These modalities are particularly derived from different sensors, which can, for example, be arranged in different locations, have different viewpoints, different fields of view, different resolutions, etc. Thus, such combination can particularly include (preferably digitally applied) compensation or correction for differences in viewpoint, field of view, resolution, etc., for example to achieve a substantially pixel-for-pixel accurate match of the different modalities in the combined dataset. In other words, combining the different modalities into a single dataset can include a geometric transformation of the data of at least one modality, performed in such a way that the combined modality is referenced to a single common coordinate system of the single dataset. In particular, modalities can be combined with pixel-to-pixel correspondences of modalities established by numerical transformation of sensed modality data (e.g., multi-channel images). For example, a visual RGB image modality and a depth modality captured by nearby but separate sensors can be combined into a red-green-blue-range-image (RGBD) or a hue-value-range-image (HVD) as a single dataset.

[0193] In exemplary embodiments, the position and / or orientation offsets between sensors can be exploited to capture modalities from different sensors, and / or the sensors can have different fields of view, such as in the case of a visual image sensor mounted next to a depth sensor. This can result in different offsets, focal lengths, and distortions in the data captured for the same scene. When combining these modalities into a single dataset, a pixel-to-pixel correspondence can be established by applying transformations or image processing to one or more of the modalities. For example, a four-channel RGBD image can be generated by combining pixel-to-pixel corresponding RGB data from a visual RGB sensor and depth data D from a depth sensor into a single image with four channels. In another example, a color conversion of the visual image from RGB to HSV (hue, saturation, value) can be established, which can then be combined with the depth data D (by omitting the saturation channel information and replacing its depth with the depth channel of another modality from another sensor) to form a three-channel "image," such as an HVD (hue, value, depth) "image," where a corresponding transformation of the depth data must be established to fit the view of the RBG data. To be presented to humans, such HVD "images" obviously must be processed into understandable representations, but artificial intelligence systems or machine learning detectors and / or classifiers can act directly and efficiently on such HVD data. The combined RGBD or HVD images can be fed directly into event detection / classification algorithms, particularly into convolutional neural networks, for example, to detect and / or identify people, to detect open doors and / or windows, to detect fire and / or smoke, to detect abandoned objects, to identify activities, to detect anomalies such as people running, sneezing, crawling, walking on tiptoe, etc.

[0194] The context information may be obtained at least in part from a first monitoring sensor of a plurality of monitoring sensors, in particular wherein the context information included in the weighting of the first modality is obtained at least in part from another second modality different from the first modality, the first monitoring sensor acting on the other second modality.

[0195] The context information may also be derived at least in part from specific environmental sensors configured to sense the context information.The context information may also be derived at least in part from an internal virtual state space state of the monitoring system.

[0196] The context information may, for example, include one or more of the following:

[0197] o lighting context (e.g. day, night, artificial and natural light, lighting sensors, brightness and / or contrast of camera images, etc.),

[0198] o Temperature context (e.g. thermal camera image, temperature sensor, internal and / or external temperature, heating system status, etc.),

[0199] o Time context (e.g. timer, weekday and hour calendar, monthly and / or calendar calendar, etc.) and / or

[0200] oVibration or earthquake context, audio context, etc.

[0201] The contextual information can be divided and / or segmented into classes by training a machine learning classifier configured to filter monitoring sensor evaluations based on the contextual information. Evaluations from different monitoring sensors can be filtered, or the actual evaluations of monitoring sensors can be disabled or specifically reconfigured based on the contextual information. For example, by disabling sensor evaluations in specific environments, computational time and energy can be saved.

[0202] In embodiments of the present invention, spatial priors and temporal filtering can also be applied to data sequences. For example, in video data, spatiotemporal consistency can be exploited. For example, in detection, the system can restrict the search space for objects based on their previously known positions and motion models. This restriction can be used not only to reduce the search space, but also to temporarily fill gaps in the detection results due to false detections, obstacles, insufficient coverage, etc. In embodiments with a mobile data acquisition platform such as a robot or drone, the movement of the platform can also be compensated, for example by using ego-motion estimation techniques.

[0203] In another embodiment, the present invention can be implemented using the output of a detector and / or classifier from a first monitoring sensor of a first modality as input to an annotator for a second monitoring sensor of the same or another modality. This can be implemented in fully automated annotation and / or can also include annotations fed back by an operator. This principle can also be implemented in the augmentation of a training database using information from another detector and / or classifier or from an operator, where the accuracy of the detector and / or classifier can be improved over time through supplementary training based on this information, which can then be processed in an at least partially automated process.

[0204] For example, in order to avoid lengthy and tedious annotations (based on which detectors and / or classifiers for all modalities are learned), especially when those annotations are not automatic and require manual user interaction, the training process can be started by applying a first detector, preferably a first detector for which a pre-trained model is available, such as an open source RGB detector from the prior art or another pre-trained detection and / or classification system. The output of the first detector can then be used to select an area previously detected by the first detector for a second detector to extract features from other modalities, such as from a depth sensor, an IR sensor, a point cloud sensor, etc. In addition to the area, classification information from the first detector can also be provided to the second detector, which classification information can be used to train and / or test the second detector. The process can involve at least partially manual feedback to filter out possible false positives, but can also be partially or fully automated by a computing system, especially detection with a sufficiently high probability of detection. In addition, the principles of generative adversary networks (GANs) can be applied in the context of the present invention, for example implemented as a deep neural network architecture comprising two networks pitting relative to each other, the pitting also being called an adversary.

[0205] Such a learning system according to aspects of the present invention can be implemented, for example, by using a first annotation output from a first detector and / or classifier as a training input for a second detector and / or classifier of the same or different modality. This can provide information about the first detection and / or classification of a new instance of a newly detected object to a human administrator for verification, while automatically processing all subsequent detections of the same instance, particularly taking into account the motion model of the object as described above. In addition to new objects, objects that are suddenly lost or lost during detection or tracking can also be provided to a human operator for confirmation or for supervising the training routine. In training, another possibility in this regard is to deliberately insert false negatives (false detections) in the training data, which can be done automatically and / or manually, for example to improve robustness and / or fault tolerance and to avoid overtraining. Once additional detectors for other modalities have been learned based on the results from the first detector, these detectors can in turn be used as annotators to improve the pre-trained model used in the first detector, for example in one or more iterative cycles, to improve the system preferably in an automatic manner.

[0206] The present invention also relates to an automated monitoring method for machine learning detection of abnormal conditions in a building or property. In one embodiment, the method may include at least the following steps:

[0207] obtaining a plurality of monitoring data from at least two monitoring sensors for the building or property, the monitoring sensors operating in at least two modalities;

[0208] Obtaining at least one contextual information, in particular environmental information, at the building or property by means of a dedicated context sensing device or from one of the monitoring sensors or from an internal state of the monitoring system;

[0209] • Combining data from one or more of the monitoring sensors for detecting the critical condition with a machine-learned automated information filter built through training with training data that includes the contextual information.

[0210] In other words, monitoring sensor data of different modalities can be classified based on contextual information in machine learning automatic information filters to become critical or not critical and / or an assessment of the severity of the status mode.

[0211] In one embodiment, this may include (at least in part) training a detector and / or classifier using training data that includes contextual information (e.g., to enable detection of one or more abnormal conditions and / or classification of conditions or anomalies as being abnormal and / or potentially critical), which is included in at least a portion of the training data. In particular, this may be accomplished via machine learning of an artificial intelligence unit.

[0212] The method may further include deploying such a detector and / or classifier to a computing unit for analyzing real-world surveillance data from at least two sensors at a building or property. Thus, a method for detecting the possible presence of abnormal state instances within the real-world surveillance data and / or classifying such states as critical or non-critical may be established, in particular based on a critical or non-critical classification model, wherein different modalities are weighted according to real-world contextual information obtained at the building or property during monitoring. As described above, the weightings may be learned, in particular by supervised machine learning.

[0213] In one embodiment, this can include a parallel and / or hierarchical evaluation structure for data derived from surveillance sensors based on machine learning in multiple contexts. An embodiment of the hierarchical evaluation can be, for example, a first hierarchical stage in which an image from a visual sensor is evaluated and (particularly only) if the image is dark (e.g., at night), a next hierarchical stage in which an image is taken with an IR camera is evaluated instead or in addition. The weighting or parameterization of the parallel and / or hierarchical evaluation structure can be established, in particular, by machine learning of context-dependent weighting factors.

[0214] Thus, a machine learning automatic information filter can be configured by machine learning for a combination of at least two modalities that rely on contextual information to obtain detection of abnormal or critical conditions.

[0215] This aspect of the invention may also include a method for obtaining a machine-learned automatic monitoring system, which machine-learned automatic monitoring system may be implemented, for example, in a computing unit of a monitoring system. The obtaining may include an automatic classifier and / or detector for security issues based on at least two monitoring sensors operating in at least two different sensing modalities. Therein, the training of the automatic classifier and / or detector includes providing training data that at least partially includes contextual environment information corresponding to the training data. The training data may in particular be or include real-world training data from the monitoring sensors. Based on the training data, a combination of at least two different sensing modalities is machine-learned, which combination can at least partially segment the different modalities based on the context information.

[0216] The combination can be established, for example, by a hierarchical and / or parallel structure for merging detections of at least two monitoring sensors, wherein the structure and / or weights of the different sensing modalities can be obtained by machine learning and depend on contextual information.

[0217] The present invention or at least one inventive part thereof may be embodied as a computer program product comprising a program code stored on a machine-readable medium or implemented by electromagnetic waves, the program code comprising program code segments and having computer-executable instructions for performing the steps of the method described herein, in particular those method steps comprising data calculation, value determination, scripting, respectively operating an artificial intelligence unit or the artificial intelligence unit itself, configuration of the artificial intelligence unit, etc.

[0218] A sixth aspect of the invention is the automatic detection of anomalies or states, which can also be combined with an assessment and classification of whether the state is critical. Such anomalies can be achieved based on information from a mobile patrol system adapted to patrol an area such as a building to be inspected, but can also be achieved based on information from fixed surveillance equipment such as surveillance cameras, various intrusion sensors (e.g., light, noise, debris, vibration, thermal imaging, motion, etc.) authentication terminals, personal tracking devices, RFID trackers, distance sensors, light barriers, etc. In particular, a combination of fixed and autonomous surveillance of an area for automatic detection and reporting of events can be achieved, which can then be classified to obtain a critical or non-critical surveillance state. In the case of a critical state, appropriate actions can be automatically triggered or suggested to human security personnel. All of this can, for example, be adapted to the specific environment under surveillance.

[0219] To date, human operators have played a crucial role in security applications, and such automated inspection tasks have been quite rare. Automated devices are often error-prone, so while they can at best assist human operators, they struggle to act autonomously. Whether it's a guard patrolling an area or an operator monitoring a security camera, human intelligence is essential in situations where high security risk exists. Automated mobile security systems attempt to support human operators with modules specifically trained to detect people, open doors, misplaced packages, and so on. These modules must be carefully prepared if such automated systems are to be fully implemented. First, considerable effort is invested in data acquisition. The datasets used to train the automated decision makers to distinguish between correct and incorrect actions must encompass a vast variety of event, appearance, and recording locations, encompassing hundreds or thousands of data samples, due, for example, to the countless possible surface characteristics, lighting conditions, and local obstructions. For security applications, this requires extensive field data recording and interaction with critical elements, such as doors, which must be captured in various stages of opening. Even with this effort, achieving reasonable generalization of the automated decision makers beyond those used for training on the test field remains a significant challenge for such autonomous systems. Typically, for automatic detection and / or classification, the deployed model is only able to accurately detect what it has seen during training, but covering all possible aspects during training is very challenging or practically impossible. For example, in a warehouse scenario, there are too many variations in possible combinations of people, open doors, or misplaced packages, but this can be crucial when detecting misplaced packages in public areas that could be used in a bomb attack, for example.

[0220] Therefore, an object of this aspect of the present invention is to provide an improved automatic monitoring system. In particular, the object is to provide such a system that automatically detects conditions and / or automatically classifies conditions as critical or non-critical, particularly with a low number of false positives and false negatives, and that can also handle rare or single conditions.

[0221] According to this aspect of the present invention, an automated system for detecting security conditions is provided. The system may at least partially include artificial intelligence or machine learning methods for detecting and / or classifying security conditions. A key aspect of the present invention is supporting the training process of a decision-making system for security applications using synthetically created examples of rare but critical conditions, particularly for high-risk security applications. Conditions may include one or more monitoring data sets, optionally combined with additional information such as time and location. To accurately detect such critical conditions, a decision maker or classifier must be adequately trained to detect them. Because collecting sufficient data for these rare conditions is not only time-consuming but sometimes impossible, the present invention includes a method for augmenting the training data set with artificially generated samples or positive samples altered to reflect malicious behavior. Positive samples can be samples of non-critical conditions that are then virtually transformed to artificially generate critical conditions. Artificially generated critical conditions, such as those involving door openings or unauthorized building access, can be used to train a decision maker, such as one implemented by an artificial intelligence system, which may include a classifier such as a neural network, support vector machine, or other model. The method according to the present invention can also be used to pre-train classifiers and decision layers for a multimodal classifier, which can then be deployed to a field system.

[0222] For example, the present invention may be directed to embodiments of a monitoring system for automatically detecting anomalies at a facility, such as a property, building, power plant, or industrial plant. The system includes at least one monitoring sensor adapted to measure at least one or more components of the facility and to generate real-world monitoring data comprising information about the facility. The system also includes a computing unit configured to evaluate the real-world monitoring data and to automatically detect conditions, including a machine-learned detector component. The detector component is at least partially trained using training data, the training data being at least partially synthetically generated and derived from a virtual model of the facility and / or at least one of its components.

[0223] For example, the present invention may involve embodiments of a monitoring system for automatically detecting anomalies at a facility. The system includes at least one monitoring sensor adapted to monitor at least one or more components of the property or building and to generate real-world monitoring data comprising information about the property or building. The system also includes a computing unit configured to evaluate the real-world monitoring data and automatically classify conditions, including a machine-learned classifier. The classifier is trained at least in part using training data that is synthetically generated and derived at least in part from a virtual model.

[0224] In another embodiment, the two aforementioned embodiments of the present invention can be combined. For example, a sensor senses real-world data, and an artificial intelligence computing system is configured to detect specific conditions (e.g., an open / closed door, the presence / absence of a person, etc.), for example, by applying person detection, object detection, trajectory detection, or another event detection algorithm, such as applied to image data from a camera. According to this aspect of the present invention, the detection is trained at least in part using syntactically generated events derived at least in part from a virtual model. There is also an artificial intelligence computing system configured to classify states based on one or more states, which can specifically be classified as critical or non-critical. States can also include data from multiple locations or a history of states, as well as contextual data such as lighting, temperature, time, etc. According to this aspect of the present invention, the artificial intelligence computing system is trained at least in part using training data synthetically generated based on a virtual model, which generally automatically simulates a large number of events and states, preferably under a wide variety of options, environmental conditions, etc. In the event of a critical state, an alarm or automated response can be triggered, for example, through a rule-based or expert system.

[0225] Anomalies to be detected as events in the monitoring system may be simulated in a virtual model, wherein a plurality of instances of training data are generated using variations of at least one environmental parameter in the virtual model.

[0226] The monitoring system may include an automatic classification unit comprising a machine-learned classifier configured to classify at least one of the events into different states, in particular, to classify at least critical or non-critical state information of a property or building. The machine-learned classifier is at least partially trained using training data, which is at least partially synthetically generated and obtained from a virtual model, through which multiple instances of states are simulated. In particular, the classifier may be trained using simulations of multiple virtual scenarios of critical and / or non-critical states provided in the virtual model. The training data for the classifier may, for example, include state information of a building, property, or building component, including at least one synthetic state and supplementary contextual information including at least time information, location information, and the like. In one embodiment, the training data for the classifier may include state sequences or state patterns, which are virtually synthesized by the virtual model and used to train the classifier, in particular, wherein the state sequences or state patterns are defined as critical or non-critical.

[0227] In other words, aspects of the present invention relate to a monitoring system for automatically detecting anomalies at a property or building, which anomalies may, for example, be included in an object state of an object such as the property or building. The monitoring system comprises at least one measuring sensor adapted to monitor at least one or more parts or building components of the property or building and adapted to generate real-world survey data comprising information about the property or building, preferably but not necessarily in a continuous or quasi-continuous manner. The system also comprises a computing unit designed to evaluate the survey data and automatically detect and / or classify anomalies, including a machine-learned classifier and / or detector, which may be implemented as a dedicated classifier unit and / or detector unit. Such a computing unit may be implemented as a central computing unit or a distributed system of multiple units. According to this aspect of the invention, the classifier or detector is at least partially trained with data, which is at least partially synthetically generated and obtained by a virtual model, preferably processed substantially automatically by the computing unit.

[0228] For example, this may include a digital representation of an event or anomaly or a digital representation of a model. Specifically, the computing unit may provide a classification and evaluation algorithm configured to classify and evaluate the detected state, wherein the normal-abnormal classification model includes at least two classes of states, at least one class being a normal class representing a classified state as "normal" and at least another class being an abnormal class representing a classified state as "abnormal," as described in detail elsewhere. Preferably, when obtained from the virtual model, meta-information about the state categories within the training data is also included in the partially synthetically generated training data.

[0229] Therein, the synthetically generated training data may be obtained, for example, by digitally reproducing a virtual image from a virtual model.

[0230] The synthetically generated training data may include multiple instances of training data obtained by varying parameters of the virtual model and / or varying the parameters generated. In particular, these variations may reflect environmental parameters (such as varying lighting) and presentation parameters (such as varying field of view, etc.). The virtual model may include a digital 3D model of at least a portion of a property or building or object and / or a physical model of the property or building or object or a portion thereof. For example, such a portion may include an instance of a specific monitored object, or may be embodied as such an object, such as a door, window, corridor, staircase, safe, cash register, keyboard, ventilation shaft, drain, manhole, gate, etc.

[0231] The surveillance sensor may, for example, include at least one camera providing digital images as real-world survey data, which camera may, for example, be implemented as a visual 2D image camera, a range image camera, an infrared camera, or some combination thereof.

[0232] As mentioned, the surveillance sensors may be embodied as mobile surveillance robots used to survey a property or building, or as fixed surveillance devices located at the property or building, or as a combination of both.

[0233] The synthetic training data may also partially include some real-world images. The synthetic training data may also include modified real-world images augmented with synthetically generated digitally rendered virtual security alarms derived from the virtual model. For example, a real-world image may be augmented to virtually depict security-related anomalies of an object, such as certain types of damage.

[0234] The training data may in particular at least partially comprise real-world monitoring data from sensors, which may be at least partially augmented by computer-synthesized virtual data.

[0235] In one embodiment, the classifier may also be implemented as a pre-trained general classifier that has been trained at least in part based on synthetic training data and deployed to the automated security system, which is then post-trained by real-world survey data of the property or building, for example during a commissioning phase and / or use of the automated security system.

[0236] The present invention also includes a monitoring method for detecting abnormal events in an automated safety system for a facility comprised of facility components. One step involves providing a virtual model of at least one facility component. Based on this, a computing unit, in accordance with software development, obtains at least one synthetically generated training data from the virtual model. This result is used to train a detector and / or classifier based at least in part on the synthetically generated training data. The detector and / or classifier can then be used to detect and / or classify abnormal conditions included in at least a portion of the synthetically generated training data. This can also be described as machine learning by an artificial intelligence unit based at least in part on the virtualized training data.

[0237] By deploying detectors and / or classifiers to a computing unit used to analyze real-world surveillance data from at least one sensor at a building or property, detection of the potential presence of instances of abnormal conditions within the real-world survey data can be established, particularly in combination with applying a normal / abnormal classification model as described elsewhere.

[0238] The acquisition of the at least one synthetically generated training data may comprise a plurality of variations of the virtual model, in particular variations of parameters of its building components and / or representations of its building components.

[0239] In particular, the present invention relates to a method for training a computing unit of a surveillance system according to the present invention, comprising an automatic classifier and / or detector for security issues based on at least one measurement sensor. The training comprises synthetically generating virtual training data by virtually obtaining training data from a virtual model, and training the classifier and / or detector based at least in part on the at least partially synthetically generated training data. The results of the training (e.g., the trained classifier and / or detector or their configuration data) can then be used on real-world data from the measurement sensor to detect security alerts.

[0240] In an example of one possible implementation with a surveillance camera, the virtual model may be a 3D model of an object that is the target of surveillance, the synthetically generated training data may include digitally rendered images obtained from the 3D model, and the real-world data from the measurement sensors may include real-world pictures of the location from the surveillance camera.

[0241] The present invention can be embodied as a computer program product comprising a program code stored on a machine-readable medium, or as an electromagnetic wave comprising a program code segment and having computer-executable instructions for performing the steps of the above-described method, in particular those steps comprising data calculation, value determination and rendering, scripting, training of a machine learning system or an artificial intelligence unit, running of a machine learning system or an artificial intelligence unit, etc. BRIEF DESCRIPTION OF THE DRAWINGS

[0242] The present invention will now be described in detail by way of exemplary embodiments with reference to the accompanying drawings, in which:

[0243] Figure 1 A first exemplary embodiment of a building monitoring system 1 according to the invention is shown;

[0244] Figures 2a to 2d Examples illustrating classification and evaluation according to the present invention are illustrated;

[0245] Figure 3 A topological or logical representation of a building model is shown;

[0246] Figure 4 Another example of classification and evaluation according to the present invention is illustrated;

[0247] Figure 5 Another example of classification and evaluation according to the present invention is illustrated;

[0248] Figure 6 An example of state detection is shown;

[0249] Figure 7illustrates another example of classification and evaluation according to the present invention; and

[0250] Figure 8 Another example of classification and evaluation according to the present invention is illustrated.

[0251] Figure 9 A first exemplary embodiment of a monitoring system according to the present invention is shown;

[0252] Figure 10 The measurement process according to the present invention is illustrated;

[0253] Figure 11 A first example of an action triggered to verify the state of an object is illustrated;

[0254] Figure 12 Another example of an action triggered to resolve state ambiguity is shown;

[0255] Figure 13 Another example of an action triggered to resolve state ambiguity is shown;

[0256] Figure 14 Another example of actions for resolving state ambiguity is shown;

[0257] Figure 15a and Figure 15b Another example of actions for resolving state ambiguity is shown; and

[0258] Figures 16a to 16c Another example of actions triggered to resolve state ambiguity is illustrated.

[0259] Figures 17a to 17b A first exemplary embodiment of a combined ground-air sensor platform system is shown;

[0260] Figures 18a to 18b Shown Figures 17a to 17b The system patrols the warehouse;

[0261] Figures 19a to 19c A second exemplary embodiment of a combined ground-air sensor platform system is shown;

[0262] Figure 20 A patrol of a combined ground-air sensor platform system with two ground vehicles is illustrated;

[0263] Figure 21 A fourth exemplary embodiment of a combined ground-air sensor platform system is shown;

[0264] Figures 22a to 22b A fifth exemplary embodiment of a combined ground-air sensor platform system is shown;

[0265] Figures 23a to 23b A sixth exemplary embodiment of a combined ground-air sensor platform system is shown;

[0266] Figure 24 It is the implementation method of UGV and multiple UAVs;

[0267] Figure 25 This is an example of workflow generation for UGV;

[0268] Figure 26 A general security monitoring system is schematically depicted;

[0269] Figure 27 A first embodiment of a safety monitoring system with feedback and training functions according to the present invention is schematically depicted;

[0270] Figure 28 A second embodiment of the security monitoring system of the present invention is schematically depicted, wherein the feedback loop comprises a common classification model;

[0271] Figure 29 A third embodiment of the security monitoring system of the present invention is schematically depicted, wherein the event detector is updated in the same feedback loop as that used to train the event filter;

[0272] Figure 30 shows exemplary implementations of warehouses to be measured in different contexts;

[0273] Figures 31a to 31c Example from Figure 1 Examples of different modalities of real-world surveillance data for warehouses;

[0274] Figure 32 A first example of a context-based machine learning result or structure for a monitoring system is shown;

[0275] Figure 33 A second example of context-based machine learning results or structures for a monitoring system is shown;

[0276] Figures 34a to 34c An example of a contextual learning monitoring system with multiple modalities according to this aspect of the invention is shown;

[0277] Figure 35a and Figure 35b A first example of real-world surveillance data in different modalities is illustrated;

[0278] Figure 36a and Figure 36b A second example of real-world surveillance data in a different modality is illustrated.

[0279] Figures 37a to 37b An exemplary embodiment of a floor plan of a site or building to be measured is shown;

[0280] Figures 38a to 38c Examples of real-world images from a monitored building or property are illustrated;

[0281] Figures 39a to 39b An example of a 3D model and virtual rendering is shown, taking a door as an example;

[0282] Figures 40a to 40b A virtual model of an autonomous surveillance vehicle's patrol and a synthetically generated IR view of an intruder are illustrated;

[0283] Figure 41 An example of a floor map with monitoring of people's trajectories is shown;

[0284] Figure 42 shows an exemplary embodiment of a virtual model and rendering of a security event, which may also be augmented to a real-world picture; and

[0285] Figures 43a to 43b An exemplary embodiment of synthetic training data generation is shown. DETAILED DESCRIPTION

[0286] The illustrations in the accompanying drawings should not be regarded as drawn to scale. Where appropriate, the same reference numerals are used for the same features or for features with similar functions. Different reference numerals in the accompanying drawings are used to distinguish different embodiments of the features shown by way of example. The term "substantially" is used to indicate that a feature can, but is not generally required to be implemented exactly to 100%, but only in a way that can produce similar or equivalent technical effects. In particular, minor deviations may occur due to technical, manufacturing, structural considerations, etc., but still fall within the scope of the present invention. The term "at least partially" includes such embodiments in which the following features hereby declared are used only for their purpose in the sense of the present application (in the sense of "completely"), as well as such embodiments in which the following features hereby declared are included but can be implemented in combination with other options to obtain similar or equivalent purposes in the sense of the present application (in the sense of including or partially).

[0287] Figures 1 to 8 A facility monitoring system with multiple measuring sensors and a "critical-non-critical" classification for recording status patterns is provided.

[0288] exist Figure 1 In FIG. 1 , a first exemplary embodiment of a facility monitoring system 1 according to the invention is depicted.

[0289] The system 1 includes a central computing unit 2. As shown, the central computing unit 2 is a single server computer 2, or, for example, a server cluster, a cloud, or the like. The central computing unit 2 provides a model 3 of a building to be measured (e.g., an office building or a warehouse); in this example, the building includes one floor with three rooms 50a-50c, and the building model 3 is illustrated as a topological 2D representation. The facility model 3 provides relationships of facility components, such as a topological relationship of room 50a to room 50b, a logical relationship of room 50c connecting door 51c to room 50b, or a functional relationship of door 51a giving room 50a (50a-c, 51a-d, 52a, 52b) an entrance. Preferably, the building model 3 is implemented as a building information model (BIM).

[0290] The system 1 also includes a plurality or a number of monitoring sensors 4, which are adapted to monitor a plurality of building components 5. In this example, for simplicity, the building components 5 and the monitoring sensors 4 are shown in a 2D representation of the building model 3. Such building components 5 are, for example, rooms 50a-50c, doors 51a-51d, windows 52a and 52b, but may also include, for example, building equipment such as electrical (optical) equipment or workstations or computer networks. Monitoring sensors 4 are distributed throughout or within a building and include, for example, cameras (such as cameras 40a, 40b, particularly those adapted for personal identification or labeling); range imaging cameras; laser scanners; motion; light; infrared; heat or smoke detectors (such as detector 41); ammeters or voltmeters (e.g., for detecting whether a light source or another electrical device is on or off or for monitoring the building's total power consumption); contact sensors, such as beam break detectors or magnetic contact sensors (e.g., sensors 42a, 42b adapted to detect whether windows 52a, 52b are open or closed); thermometers; hygrometers; key readers; and input detectors that allow operator input. Measurement sensors 4 can be installed in a building to operate autonomously or by an operator. Building components 5 can be measured by a single monitoring sensor 4 or by more than one monitoring sensor 4, for example, regarding different characteristics of building components 5 (e.g., by two cameras with different viewing angles or by a camera and an IR sensor). The monitoring data of the monitoring sensor 4 includes, for example, color images, depth images, thermal images, videos, sounds, point clouds, signals from door sensors, elevators, or key readers, and is transmitted to the central computing unit 2 by, for example, wired and / or wireless and / or Internet communication means (indicated by arrow 70). The monitoring data is assigned to (sub)objects of the building model corresponding to the respective building components 5 and stored in the database of the central computing unit 2.

[0291] The monitoring sensors 4 are mounted at fixed locations and / or are mobile, such as a monitoring drone or monitoring robot 43 shown in the figure, which serves as a mobile station for the several types of monitoring sensors described above (e.g., cameras, microphones, IR sensors, etc.). The central computing unit 2 knows the location or position of each monitoring sensor 4 or the position of the building component 5 measured by the sensor 4, so in the case of mobile monitoring sensors 43, the monitoring sensors 4 send their position data to the central computing unit 2 via the communication device 5. In this example, the positions of the monitoring sensors 4 are integrated or incorporated into the building model 3.

[0292] System 1 also includes a state derivation device 6 for deducing a state at or within a building, e.g., whether a person is present and / or the light in room 50a is turned on or off at time T1, door 51c is opened at time T2 and closed at time T3, etc. State derivation device 6 can be integrated into monitoring sensor 4, for example, implemented as contact sensors 42a, 42b or a monitoring sensor of monitoring robot 43, which detects, for example, motion as an event. In this example, state verifier 6 is integrated into central computing unit 2, analyzing a monitoring data stream, e.g., from camera 40, and detecting changes in its (significant) characteristics or video or image stream as states. Specifically, it detects states such as a person in a building component (e.g., one of rooms 50a-c) or an object, such as a suitcase or package, introduced into room 50c, and thus measures room 50c in the image or video of camera 40a. As another example, monitoring sensor 4 is located in or integrated into a light switch, and state verifier 6 detects whether the light is on or off, controlling the light (or current) of the building component "room" or "light switch." As shown in the example, a state represents, for example, at least one physical property of a building component 5, such as the temperature of rooms 50a-50c or the orientation / position of windows 52a, 52b or doors 51a-51e, as well as at least one property associated with the building component 5, such as a change in its environment or contents, such as a person leaving or entering the building component "room 50a." Instead of handling specific states, the system can also manage non-specific states that simply indicate whether the building component 5 is in a normal or abnormal state. A state can be "semantic," such as a person detected, a door opened, etc., or it can refer to at least one simple change in the data, such as a significant difference between a point cloud or image of a sub-object acquired a few hours ago and a point cloud or image acquired now. These states can be hierarchically structured; for example, a person detected in a hallway can be standing still, walking, running, or moving slowly, etc., instead describing the events "person detected," "running person detected," "walking person detected," etc.

[0293] States are combined or grouped together, taking into account the topological and / or logical and / or functional relationships or links between building components. Such combined states form a state pattern. A state pattern can be formed by the states of the same building component at different times, such as the states of door 51c at times T1, T2, and T3. Alternatively or additionally, a state pattern can be formed by grouping or recording the states of at least two building components, where the states are derived from monitoring data generated simultaneously or at different times. For example, a state pattern is a combination of the state of window 52a and the state of a light in room 50a at time T1, where window 52a and room 50a are topologically related. As another example of a state pattern, the state "open" of door 51c is linked to the state of the light in room 50c ("on" or "off") and the state of one or both doors 51a, 51b ("open" or "closed"). When a person entering door 51c would (normally) turn on the light in room 50c and then also open one of doors 51a, 51b, these three building components 51c, 50c and 51a, 51b are in a topological relationship as well as a functional / logical relationship. A state pattern thus has at least one timestamp, such as the time of one of its underlying states or the start and / or end of a state sequence.

[0294] The central processing unit 2 further provides a critical-non-critical classification model and a classification algorithm for classifying the detected state patterns or state sequences with respect to their criticality, in order to answer the question of whether the recorded state patterns are to be considered critical or non-critical with respect to safety considerations. The criticality classification is thus based on at least one topological and / or logical and / or functional relationship and at least one timestamp of the corresponding state patterns.

[0295] To give a very simple example, if door 51c is open and no one is subsequently (or previously) seen in room 50c by camera 40a, then this pattern of the state "door 51c is open" and the state "no one in room 50c" (or: "no one in room 50c near door 51c") is classified as critical based on the logical relationship between room 50c and door 51c, as a person should enter (or leave) room 50c when door 51c is opened. A slightly more complex example is that if either of the two cameras 40a and 40b does not detect a person before or after detecting the opening of door 51c, only the state pattern including the state "door 51c" is designated as "critical." In other words, since the state "no person detected / no person present" of only one camera 40a (40b) is still classified as non-critical, and only the "verification" of "no one" by the second camera 40b (40a) results in the classification as "critical," there is a higher "obstacle" than in the previous example. The first example will result in more assignments to "critical" at the expense of more risk of false or unnecessary classifications and corresponding alerts, the second example, with a finer underlying state pattern, is more complex but more robust.

[0296] The criticality classification of the usage status pattern is further illustrated with reference to the following figures.

[0297] Figures 2a to 2d An example of classification according to the present invention is shown. All three figures show a classification model 7 implemented as an n-dimensional feature space. For simplicity, the feature space is a two-dimensional state space, however, a higher-dimensional state space is preferred. The state pattern is represented by an n-dimensional vector in the feature space.

[0298] Figure 2a A normal-abnormal classification model 7 with two status pattern classes 8b and 8a provided by the central computing unit is shown. The first class 8b is a normal class of status patterns classified as "normal", while the second class 8a is events classified as "abnormal" or "not normal". From the outset, two classes are predefined based on a database with stored or predetermined states or states entered by a manual operator or according to a planned schedule (with reference to planned activities in the measured building) (for example, as shown in the figure, based on the status pattern 9a shown as a small line segment (for readability reasons)). Alternatively, the central computing unit does not provide a predefined classification model, but is designed to establish a classification model based on detected events without using predefined events. In this example, the status pattern 9a is classified at least mainly with respect to date and time. That is, the status pattern is recorded with the date and time of occurrence. In addition, topological information can be included.

[0299] exist Figure 2b In , additional state patterns 9b are recorded with respect to date and time and inserted into the 2D feature space. Figure 2b In the example, since the computing unit is in the learning phase, the state pattern 9b has not yet been (finally) classified as normal. The classification model 7 or the underlying classification algorithm is preferably trained and updated by a machine learning algorithm, such as a support vector machine, random forest, topic model or (deep) neural network.

[0300] exist Figure 2c , illustrates the final normal-abnormal classification model at the end of the training phase. Compared to the initial normal-abnormal classification model, the "normal" and "abnormal" classes have been refined into classes 8a' and 8b'. The computing unit has learned, for example, that the occurrence of state pattern 9a / 9b at 6:00 is "normal" (except on Saturdays and Sundays) and that it is "normal" at certain times on Saturdays. On the other hand, from Monday to Thursday, records of the corresponding state pattern at certain times in the afternoon that were previously classified as "normal" in the initial classification model will now be classified as "abnormal."

[0301] Figure 2d The final step in the classification of the criticality of the status patterns is shown. The status patterns assigned to the "normal" class in the previous normal-abnormal classification are classified as "non-critical". The recorded abnormal status patterns 9c, 9c', 9c" are classified as "abnormal" states based on their properties that lie outside the "normal" class. The central computing unit now evaluates how serious the occurrence of these status patterns 9c, 9c', 9c" is and therefore whether they are "critical" with regard to the safety of the measured building. For the assessment of the criticality, the degree D, D', D" of deviation of the individual status patterns 9c, 9c', 9c" from the "normal" class 8b' is determined as the basis for the assessment.

[0302] In this example, taking into account the features provided by the facility model and the resulting topological and / or logical and / or functional relationships, the degrees of deviation D' and D" of the state patterns 9c' and 9c" are low, as indicated in this simplified example by the short distances / arrows to (the boundaries of) the "normal" class 8b'. Consequently, these two events 9c, 9c" are considered "non-critical" by the calculation unit according to the critical-non-critical classification model that depends on the degree of deviation.

[0303] On the other hand, the degree of deviation D from the normal class 8b' of the state pattern 9c is determined to be high or significant. Since the degree of deviation D is considered relevant, i.e., falls into the "critical" class of the critical classification model, the state pattern 9c is classified as "critical".

[0304] Optionally, the system then reacts to the detected "critical" event, for example by issuing a user alert or notification, preferably with instructions or options for reaction and / or embedded in the 3D visualization of the building model.

[0305] Thus, the monitoring system first determines whether the monitored state pattern, which is a potentially safety-relevant state, is "abnormal" (not a normal state pattern), and then determines whether this "abnormality" must be considered "critical." This process of having a criticality classification including a normality classification provides an enhanced safety architecture, enabling a more reliable and robust assessment of the building state, e.g., reducing false alarms.

[0306] As mentioned above, the classification is not only based on a single state / change or modification of a building element 5, but also on a sequence and / or pattern of states of one or more building elements 5 (see Figure 1 In other words, the system 1 is adapted to interpret at least two survey data from one measurement sensor 4 at two different times, or at least two survey data from two measurement sensors 4 at the same or two subsequent times. As a simple example, using the survey data of the camera 40a and the contact sensors 42a, 42b (see Figure 1 ), the state pattern represents the sequence "door 51a is open" and "window 52a is closed". In order to detect such sequences and / or state patterns, algorithms such as hidden Markov models (HMM) or recurrent neural networks (RNN) or conditional random fields (CRF) can be applied, which will be further summarized below.

[0307] This complex approach using state patterns representing relevant sequences and / or patterns of conditions of building components has the following advantages: the greater the complexity or number of parameters / characteristics / dimensions of the state pattern, the more robust the assessment of the criticality. For example, a large number of parameters possessed by the state pattern is advantageous for determining the degrees of deviation D, D', D" from the "normality" class as a criterion for criticality.

[0308] In the preferred case of a status pattern involving two or more building components, the severity assessment takes into account the topology and / or functional and / or logical relationships of these building components, so that even more test criteria for the "criticality" or non-criticality of the status pattern are available, resulting in an even broader assessment basis. Figure 3 This is further outlined.

[0309] Figure 3A topological or logical representation 12 of a building, as part of a building model, is shown in diagram form. Dashed lines 11 represent proximity relationships; i.e., the dashed line between door 51a and door 51b indicates that door 51a is immediately adjacent to door 51b. Solid lines 10 represent connections; i.e., a person can follow these edges from one room 50a to another 50c. Based on these logical or topological relationships, as an example procedure, if a person wants to enter room 50a at a certain time of night and open window 52a from outside the building, they would enter the building through entry door 51c, activate the light in room 50c, walk along room (corridor) 50c, pass through door 51a, activate the light in room 50a, and open window 52a. This connected flow of actions is then represented as a state pattern, detected using surveillance data from cameras 40a, 40b, light detector 41, and contact sensor 42a at window 52a. It can be seen that not all actions or building components need to be recognized / checked by the surveillance system to form a state sequence that represents and indicates the basic process described above.

[0310] Compared to a "single" state, any criticality of a state pattern and / or sequence can be detected with greater robustness, thereby detecting the critical overall state of a building. For example, assume that the above sequence is used as the basis for a criticality class. Now, if only the state "window 52a is open" is considered at night, then according to a simple monitoring system / model, this would result in a classification as "critical", because opening a window at night usually indicates an intruder such as a thief. Instead, the state "window 52a is open" is loaded into the investigation event chain and grouped together as a state pattern to be classified and evaluated. If a state pattern similar to the above action flow is detected, the recorded state pattern including the state "window 52a is open" is classified as "non-critical", even though it might be classified as "abnormal" in an optional normality classification model.

[0311] If any significant deviation is assessed, such as the window 52a and the door 51a being open, but the light detector 41 detecting no light, there is a high deviation from "normal" and the detected pattern is classified not only as "abnormal", but also as "critical" (as can be seen from this example, the status pattern can also include zero components, such as "no light detected" or no change, where the change would normally be detected separately or would normally be expected given a known sequence of events and / or pattern). This means that not every "abnormal" status pattern will be classified as "critical". In other words, every status pattern classified as "normal" may be classified as "non-critical", because certain status patterns may be detected regularly, yet still pose a risk to the building. According to the present invention, criticality can be tested without considering normality-abnormality or considering normality-abnormality, and further a critical-non-critical distinction can be made between these categories.

[0312] The classification is achieved by taking into account the topology or logical links provided by the model 12 as additional parameters or features. Any state related to at least two building components 50a-52b is evaluated, as the relationship is checked to see if it complies with the known relationship between these building components 50a-52b. As a result, the classification becomes more reliable.

[0313] In more complex methods, not only the timestamp of the state pattern or the timestamp associated with one of the underlying states is considered, but also the time points of the individual occurrences that constitute the state pattern. For example, the time interval between "door 51c is open" and "door 51a is open" is determined, and the degree to which this time interval of the detected pattern deviates from the previously measured time interval or does not deviate, and / or the degree of criticality itself can be evaluated, such as when the time interval exceeds a certain time limit. This consideration of the time interval increases the probability of correct classification as "critical."

[0314] As another example, a typical pattern may be generated by a cleaning crew as they clean one room after another. It can be expected that, in the case where rooms 51a, 51b have nearly identical dimensions, the cleaning of rooms 51a, 51b will take more or less the same amount of time, which is called CT. Therefore, the sequence acquired with the door sensor can only be seen as follows: 51a (door open) - [CT later] 51a (door closed), 51d (door open) - 51d (door closed) - [CT later] 51b (door open), etc. Instead of a door sensor, the mobile security robot 43 (see Figure 1 ) can patrol along corridor 50c, detecting people or objects or open / closed doors. The patrol of robot 43 can result in the following state pattern, for example, when patrolling from the entrance to the end of corridor 50c while a cleaning person is in room 50a: 51c (door closed) - 51a (door open) - [moving into room 50a] 50a (person detected) [continue patrolling, moving out of room 50a] - 51b (door closed). The idea is to combine all the detections from different sensor systems to produce this state sequence, for example, combining the input from the door sensor with the input from the mobile monitoring robot 43.

[0315] In addition to human detection, other useful information can be obtained from computer vision and machine learning algorithms, such as human identification, tracking, and re-identification, which can be utilized to, for example, extract trajectories by tracking people (e.g., in corridor 50c monitored by surveillance camera 40a), and then classify and evaluate these trajectories as normal or abnormal. Alternatively, identification of a person or a class of people can be achieved, such as using surveillance camera 40b to identify cleaning staff who have access to the building at night and can be identified from their colored uniforms, or using a high-resolution camera to perform biometric person identification based on facial and / or iris features.

[0316] Another option is to re-identify a person who appears on the network of surveillance cameras 40a, 40b. For example, surveillance camera 40b detects a person entering. Some time later, camera 40a detects the person leaving room 50a. Re-identification can help determine whether it is the same person, and therefore if the room is still occupied, or if a person is detected and tracked in hallway 50c, a model of the person's appearance is learned from this data. If a person is detected and tracked at building exterior door 51c, re-identification attempts to match appearance descriptors and estimate whether it is the same person previously seen in hallway 50c.

[0317] Furthermore, state patterns can be detected, including by monitoring power consumption as previously described, as turning on lights or computer workstations in rooms 50a-50c leaves a fingerprint in the building's overall power consumption. While this survey data cannot be directly assigned to a specific room 50a-50c, it can be combined with other survey data to make an assignment. For example, if a person is detected entering room 50b and a small increase in power consumption is detected, this indicates that the person in room 50b has turned on a light or workstation. A person entering rooms 50a-50c at night without turning on the lights could indicate an intruder, and therefore this state has a high probability of being classified as "abnormal" or "critical."

[0318] Figure 4 An example of a classification of a sequence and / or pattern of states of one or more building components is shown. Figure 1 ) image, a first state 27a, such as "person detected", is determined. Based on the monitoring data 26b and 26c of two different sensors, such as a color image and a depth image taken from the mobile security robot 43 and / or from one sensor at different times (see also below for Figure 6 ), a second change or state 27b of the building component is derived, such as "door 51a is open." Both states 27a and 27b are fed into the classification model 25. The classification model 25 outputs a classification result 28 indicating the overall state of the monitored object, for example, with a probability of 80% "not critical" to 20% "critical."

[0319] exist Figure 5, a more complex example of classification using a classification model 30 of sequences and / or patterns of states of one or more building components is shown. A state 27a is derived from monitoring data 26a of a first building component (e.g., a door, utility, corridor, etc.). This state 27a, along with data 26b input to a classification sub-model 25a, delivers the state of the first building component as output 28a of the sub-model 25a. Data 26b and data 26c related to a second building component (e.g., survey data and / or data from a database) are fed into the sub-model 25b to determine the state 25c of the second building component 2. The states 28a and 28b of the two building components, along with additional monitoring or database data 26d, are input to the classification sub-model 25c, which classifies and evaluates the inputs 28a and 28b and ultimately outputs a classified state pattern 31 of the monitored object—"critical" or "non-critical."

[0320] exist Figure 6 , a portion of a building model 3, including partial rooms 50a-50c and a door 51b, is shown together with a mobile measurement robot 43. The model also shows three different positions P1-P3 of the measurement robot 43, representing its position within the building (corridor 50c) at three different times T1, T2, and T3, respectively. Also shown is a monitoring camera 40c, one of the monitoring sensors of the robot 43. At each time T1-T3, the camera 40 is aimed at the door 51b and captures a first image 29a at time T1, a second image 29b at time T2, and a third image 29c at time T3.

[0321] Image processing of images 29a and 29b by the state deduction means (of the robot 43 and / or the central computing unit) reveals that at time T1, door 51b is closed, whereas at time T2, door 51b is open, so that the state (mode) "door 51b open" is detected and can be classified and evaluated, fed into the classification model, as described above. Figure 4 As described.

[0322] The state detection with respect to images 29a and 29c shows the state "a package was placed close to door 51b (between times T1 and T3)". In a simple monitoring system where only this single state is classified, the detection of package 53 would most likely result in a classification as "critical" and the sounding of an alarm. However, package 53 may not pose any risk to the building. In a more sophisticated approach, where the detection is part of a detection pattern or state pattern, the classification can be assessed with a higher level of confidence, for example because a large amount of measurement / image data of door 51b is collected over time (e.g. throughout a day, a week or a month), enabling the classification algorithm to learn (e.g. by a machine learning algorithm) the normal and therefore non-critical state patterns associated with objects 51b that are similar to doors 51b.

[0323] If, for example, package 53 or a specific package (identified by sufficiently detailed image processing) appears repeatedly, its detection can be classified as "normal." Alternatively, if, as part of the corresponding state pattern, it is discovered that package 53 was not placed at a normal time, but was placed by someone believed to be a building employee, then classification as "abnormal" but "non-critical" will occur.

[0324] However, if the package 53 is dropped repeatedly, but at times very different from the "normal" detection times, or if the package 53 is dropped by an unidentified person, the corresponding event is classified not only as "abnormal", but also as high severity.

[0325] The normal state can be inferred from detected motion, such as door 51b being closed at time T3, or from a state determined directly from survey data, such as in image 29c, door 51b appears to be in the same "closed" state as in image 29a. This state similarity can be determined, for example, based on depth and / or point cloud data analysis.

[0326] Figure 7 An example of using a hidden Markov model (HMM) 16 to classify state patterns is shown. In the case where a person enters a building through an entrance door 51c (start S), the hidden Markov model 16 is applied to determine whether the person is in a corridor 50c or in a room 50a or room 50b (see Figure 1 ), which uses the starting probabilities 13 for the following states: Ra: “Person in room 50a”, Rc: “Person in corridor c”, and Rb: “Person in room 50b”: interconnection probabilities 14 and output probabilities 15: Fa: “Door 51a is open”, Fc: “Person detected in corridor 50c”, and Fb: “Door 51b is open”. In addition, the HMM 16 can evaluate the extent to which the person's behavior corresponds to a typical pattern or an atypical pattern, thereby determining the probability of assigning the corresponding pattern (representative event) to the “critical” or “non-critical” category.

[0327] Figure 8 An example of classifying state patterns using a neural network 20 is shown. The neural network 20 includes an input layer 17, a hidden layer 18, and an output layer 19. Each layer 17-19 has a set of units 22, 23, 24. Arrows 21a, 21b represent weights between units 22-24.

[0328] The monitoring data underlying the detected sequence or pattern (e.g., detection of a person, door opening, etc.) is fed to the input layer 17, for example, together with some other data (e.g., day of the week, location information, etc.). Figure 3The topological relationship, functional relationship or logical relationship of the building model 3 shown is input into the neural network 20, for example, door 51d connects room 50a with room 50b, door 51a is adjacent to door 51b, etc.

[0329] The cell 24 in the output layer indicates whether the object is in a non-critical state (cell 24b) or a critical state (cell 24a), for example, 80% probability of non-critical and 20% probability of critical. In more complex models, the output layer 19 may contain cells 24 for individual building components, indicating whether the particular building component is in a normal state or an abnormal state.

[0330] The classification can also be a combination of several concepts, such as HMM, neural network, etc. For example, a sub-method can indicate whether the detected person is an employee, a visitor, or an intruder, such as a person detected in corridor 50c at 8:00 (high output probability of an employee), a person detected in corridor 50c at 10:00 (high output probability of an employee or visitor), or a person detected at 2:00 (high output probability of an intruder). The output of the algorithm (i.e., the person detected at 8:00 is 80% an employee, 17% a visitor, or 3% an intruder) can be input to another step, where this information is combined with other state / state patterns or sensor data.

[0331] Figures 9 to 16c A monitoring system capable of monitoring a robot and resolving state ambiguity is provided.

[0332] Figure 9A surveillance system 103 is shown having a mobile surveillance robot 100 for monitoring a facility 102. Facility 102 (e.g., an office building or warehouse) includes three rooms 150, 151, and 152, an entrance door 106, two room doors 107, 108, and other facility objects, such as a box or container 105 within room 151. Robot 100 is implemented as an unmanned ground vehicle (UGV) in this example and includes a main body 100b, a drive unit 100a, a computing unit 100c, and, in this example, at least a first surveillance sensor 110 and a second surveillance sensor 111 for collecting survey data. At least first surveillance sensor 110 is implemented as a non-contact sensor suitable for large-area or large-scale measurements. Surveillance sensors 110, 111 are implemented, for example, as one or more cameras, IR cameras, laser scanners, or motion detectors, but may also include measurement sensors such as smoke detectors or sensors for personal identification such as ID card readers or key readers. Advantageously, the two monitoring sensors 110, 111 are of different types, such as two cameras covering different parts of the electromagnetic spectrum, or a sensor pair comprising a passive sensor such as a thermal or IR sensor and an active sensor (radar or lidar) that emits a measurement signal. The first and second sensors 110, 111 may also differ in their measurement accuracy, such as the first sensor 110 being designed for coarse but rapid monitoring, while the second sensor 111 is designed for fine and detailed monitoring.

[0333] The robot 100 also includes a motion controller (not shown) for controlling any of the robot's 100 actions, such as movement or acquisition of survey data. The robot 100 is configured to autonomously patrol rooms 150-153, continuously changing its position and orientation, thereby also allowing it to survey the exterior of the facility 102. Its mission is to observe the status of property 102 / objects 150-152, 105, 106-108 on property 102 (e.g., the status of doors 106-108 or container 105) and detect security-related events associated with property 102 / objects 150-152, 105, 106-108. For example, the robot 100 may be configured to check whether doors 106-108 are open or closed, or to determine the location and / or appearance of container 105. As another example, the robot 100 may need to check whether a window is open or closed, whether the heating is on or off, whether a faucet is leaking, or any other security or safety-related event on property 102.

[0334] The computing unit 100c has a state ambiguity recognition and remediation function to note or recognize ambiguity in a detected state and optionally take action to generate state verification data suitable for resolving the ambiguity and inferring the state of a particular object.

[0335] For example, the state detector obtains some sensor input (e.g., an image) from surveillance sensors 110, 111, and / or 104 and, based on a person detection algorithm applied to the image, determines whether room 150 is occupied, unoccupied, or unsure. In practice, the person detector may know that there is a 51% probability that a person is present and a 49% probability that a person is not present. Because this difference is very small, i.e., below a predefined threshold, it is interpreted as unsure and an ambiguity exists that needs to be resolved by performing an action to collect additional information about room 150 and the person.

[0336] Robot 100 can be fully autonomous and constitute a (complete) surveillance system in its own right, or, as shown in the example, part of a (master) surveillance system or network 103, including a central computer 101 and, as shown in the example, at least a third surveillance sensor 104. Third surveillance sensor 104 can be part of another surveillance robot or, as shown, a fixed sensor, such as a camera 104 installed in a corner of room 150. System 103 includes communication means for surveying devices 104 and 100 to communicate with each other and / or with central computer 101 (indicated by arrow 109). As shown, central computer 101 is a single server computer 2, or, for example, a server cluster, cloud, or similar device. The means for detecting state ambiguity and / or notifying state ambiguity (computing unit 100c) can be located at robot 100, central computer 101, or another unit of system 103, or can be distributed across two or more system components.

[0337] exist Figure 10 , illustrates a monitoring process performed by a robot 100 according to the present invention. A portion of a room 150 of a building (including a door 108) and a flowchart are shown. At step 112, the monitoring robot 100 uses a first sensor 110 at an acquisition position P10 to acquire first survey data 108a of the door 108. In this example, the first sensor 110 is implemented as a camera, and the survey data 108a is an image of a portion of the door 108 and a portion of an adjacent wall.

[0338] At step 113, the robot's computing unit evaluates the survey data 108a and deduces a state 114 of door 108; in this example, deduction 113 results in the following state: "Door 108 is closed." Furthermore, the computing unit determines an uncertainty 115 in deduction 113 of door state 114, which serves as an indicator of the ambiguity of deduced state 114. In other words, it estimates the extent to which deduced state 114 actually corresponds to the true state of door 108.

[0339] In this example, door 108 is actually close but not completely closed; there is a small opening or gap 160. From the perspective of robot 100 (i.e., viewpoint P10 of its camera 110), opening 160 is barely visible, and therefore, image processing of image 108a (image 108a acquired at acquisition position P10) by the computing unit allows only a vague inference about the state of the door, with detected state 114 having a high uncertainty, represented in this example by an uncertainty value or (probabilistic) ambiguity of 70%.

[0340] In step 116, a check is performed to determine whether the determined uncertainty 115 is above a defined threshold. If the result is "no," i.e., there is a low uncertainty and the detected state 114 can be considered correct or unambiguous, the robot 100 continues its patrol, moving on to the next object to be measured (step 118). However, if the ambiguity indicator 115 is above the threshold ("yes"), the computing unit triggers an action 117, which is suitable for verifying the state of the door 108. The triggered action 117 may, for example, be the acquisition of another image of the door 108 at a second imaging position different from the first position P10, the interaction of the robot 100 with the door 108 (e.g., attempting to push the door 108 to test whether it can be opened without using its joystick), or the acquisition of additional data about the door 108 by another device, such as an image of the door 108 taken by the measurement camera 104. All of this will be explained in detail below.

[0341] In other words, the detected state 114 can be considered as a preliminary derivation state. If the detection at step 113 is evaluated as reliable (low uncertainty), the state 114 detected at step 113 is considered to be the final result. If the derivation at step 113 is evaluated as unclear or ambiguous (high uncertainty), the derivation of the preliminary state is verified by a verification action 116 of the robot 100 to resolve the ambiguity determined at step 113.

[0342] Of course, the state of door 108 does not yet need to be determined in step 113 in such a way that the output "door 108 closed" is explicitly generated by the calculation unit, but only to the extent that an evaluation or rating of detection 114 is possible. In other words, a simple decision about the true state of door 108 is not required in step 113, but rather the evaluation of the survey data allows the (uncertainty) of the evaluation to be determined with respect to the deduction 113 of the state 114 of the door.

[0343] In this example, a two-step process 117p, 117e is used to resolve the determined ambiguity based on the action 117. Before triggering the action 117, the computing unit plans the action to be triggered in step 117p and implements the action in step 117e to resolve the event ambiguity 115. Considering the object 108, the detected event 114, the ambiguity indicator 115, the first survey data 108a, or other conditions such as the equipment or environmental conditions of the robot 100, there are usually more than one possible action 117 for generating the verification information. In the planning step 117p, the generation of the verification information is optimized by selecting the most efficient or best action 117 from the various possible actions 117.

[0344] As described above, the decision process, for example, considers the detected state 114 as the specific state in question for a specific object or type of object 108 and assigns a specific optimal action, for example, predefined in an event-action assignment database or calculated by a computing unit based on the first monitoring data 108a. Alternatively, the first monitoring data can be used as a criterion for selecting the action to be triggered, such as selecting the first or second measurement sensor for generating the second survey data, or selecting a measurement resolution based on the size of missing data points in the first survey data. Preferably, planning step 117p is based on more than one criterion, for example, based on multiple criteria, such as the type of object 108, ambiguity, and the number of verification data points to be acquired. Further examples of optimization of verification information generation are provided in the following figures.

[0345] An advantage of the present invention is that the mobile surveillance robot 100 can handle survey situations where the object state of the property being surveyed is unclear or ambiguous, which is common in reality. Ambiguous detection or inference of event 114 can, for example, be the result of an uncertain or ambiguous state of the object (such as a door between two or more possible states, or an object state / condition / event previously unknown to the robot 100, such as door 108). Ambiguous inferences can also be the result of adverse conditions in survey data collection, such as low light conditions when generating surveillance data, unrealistic robot position P10, or interfering environmental influences. According to the present invention, the robot 100 can take actions to compensate for these shortcomings and provide authenticity regarding object events.

[0346] Compared to a time-consuming solution where the robot 100 first conducts a detailed survey of the status of each object 108, the present invention advantageously enables a rapid survey of all locations on the property, wherein only those objects 108 are surveyed in detail and their status is assessed as requiring a review based on the results of the previous rapid survey, which is more time-consuming. In other words, the robot 100 sweeps over the objects under surveillance and, based on the assessment results of these sweeps, selects those objects that require a more in-depth survey (rather than conducting a detailed survey of all objects or failing to conduct a detailed survey at all).

[0347] Figure 11 The first example of the action triggered in order to generate the verification information of the object is shown. The triggered action 120 is to change the survey position of the robot 100 from the initial position P10 (see Figures 2a to 2d ) is changed to a new survey position P20. In addition, the orientation of the first monitoring sensor 110 is changed so that a new field of view 119 is generated. In summary, the acquisition position and direction P20 are changed so that the state of the door 108 can be unambiguously deduced, thereby verifying the state under the surface. In the position and orientation P20, since the viewing direction points to the edge between the door 108 and the wall / door frame 121 and the position close to the gap 160, the opening 160 of the door 108 can be better perceived compared to the original position P10. In other words, the computing unit or state detector of the robot or system triggers the robot to enhance or optimize the survey position and / or alignment (of the entire robot or at least of the measurement sensor in question) for surveying the door 108. A second image as second survey data is taken at the position P20 as verification information. Taking into account this verification information, the computing unit deduces that the real state of the door 108 is "open".

[0348] In other words, if the robot 100 is uncertain or unsure about the door state detected based on the first monitoring data, the robot 100 plans a different position P20 from which the door state can be better observed (e.g., more efficiently or optimally because the position allows for better or optimal generation of survey data of the door 108 with higher resolution, higher sensor signal / less noise or with less interfering / obstructing environmental influences), or whereby additional or other components and / or features of the door 108 are surveyable that are suitable for state verification. In particular, if the door plane 123 or the door frame 121 obstructs the door gap 160, the robot 100 moves backward and approaches the door 108 from an optimized (observation) angle to get closer to the door gap 160. Advantageously, the robot 100 repeats this action, establishing a third survey position and / or direction, etc., until the door state is identified with high certainty. As an option, it then sends a notification to the central computing unit (see Figure 1 ).

[0349] The acquisition position and / or orientation used to generate this state verification information is optionally determined similarly to the Next Best View (NBV) method known in the art. The purpose of the Next Best View planning according to the present invention is to identify state-relevant features with high certainty (therefore, it is not necessary to identify or survey the door 108 in its entirety). In this example, the relevant features may be the lock of the door 108 and / or the edge of the door panel 123 or the adjacent wall 121, or the presence or form of a shadow or light cone. Preferably, the computing unit of the robot 100 is provided with a database of event-relevant features of a plurality of objects of the property to be monitored.

[0350] For example, the robot 100 learns the optimal or best positioning and / or orientation by associating robot positions P10 or P20 with object states. One or more criteria are defined for the object state, which are then optimized with respect to robot positions P10, P20 or the robot's waypoints (e.g., arranged for a reconnaissance patrol). For example, the state of a door is defined by the size of the door gap 160. During the training phase, many robot positions or waypoints P10, P20 are sampled around the door 108, where the door 108 has many different opening angles, i.e., many different sizes of door gaps 160. Based on all robot positions for the defined door opening angles, the optimal robot positions P10, P20 are defined. The optimal view is learned from this data. This can be done for a variety of objects and various object states. For example, if the robot is performing a task at waypoint No. 2 on a patrol and some ambiguity is detected, there may be several defined alternative waypoints in the planning, such as No. 2a, No. 2b, etc.

[0351] More generally, a mapping must be learned that relates the survey sensory input to the state. Assuming a quasi-static environment, changes in the sensory data are due to the robot 100's own motion. Therefore, the mapping to be learned is between changes in the survey sensory data and object state ambiguity. Accounting for this difference in survey sensory input allows for a more object-independent mapping, meaning the robot 100 can be empowered to survey not only the door gap 160 but also gaps in general.

[0352] Figure 12 Another example of an action triggered to resolve state ambiguity is shown. In this example, the robot 100's second survey sensor 111 is used to trigger the acquisition of second surveillance data of the door 108. Compared to the first surveillance sensor, such as a camera with a relatively limited resolution, the second sensor 111 is used for finer surveillance. Triggering the second survey with the second surveillance sensor 111 rather than the first surveillance sensor 110 can be a result of selecting the more efficient of the two surveillance sensors 110, 111 to resolve the ambiguity. For example, if evaluation of the first surveillance data reveals that the measurement resolution is too low to deduce a clear state of the door 108, the second surveillance sensor 111 is selected because it is capable of achieving a higher measurement resolution than the first surveillance sensor 110.

[0353] In this example, the second sensor 111 is an active monitoring sensor that transmits a measurement signal 122, specifically aimed at a boundary door / wall 121 (i.e., a door panel 123 and a portion of the door frame 121), and receives the signal 122 reflected from the boundary door / wall 121. Such a second sensor 111 is, for example, a laser scanner or a laser-based rangefinder.

[0354] Verification information available from the fine measurements of the second survey sensor 111 is, for example, a high-resolution 3D point cloud representing the door-wall edge. Such a 3D point cloud provides detailed information about the state of the door 108. As an alternative, two precise distance values ​​for the distance from the robot 100 to the door panel 108 and the distance from the robot 100 to the wall (i.e. the door frame 121) are provided by measurements made by the second measurement sensor 111 which is implemented as a laser rangefinder. Comparing these two values ​​gives a high-certainty indication whether the door 108 is still closed or has been opened. Thus, the 3D point cloud or the two precise distance values ​​allow the state "door open" to be detected with very low uncertainty or without ambiguity. Similar to the Figure 3 The process described optionally creates (ie, stores in a database in the robot computing unit) a correlation map of the second monitoring data acquisition position and / or orientation and the object state.

[0355] As an alternative example of status verification by acquiring the second monitoring data, the object being viewed is a person whose identity is to be checked in the building being viewed. Figures 2a to 2d ) to clearly determine the identity of a person, then verify his or her identity using the second surveillance sensor 111, which is a high-resolution camera with enhanced human recognition capabilities. As another option, the second surveillance sensor 111 is used as a key reader, and the actions triggered thereby include requesting the "uncertain" person to insert or display an identification key and reading the key. This surveillance robot 100 (i.e., surveillance system) can provide the following advantages: for example, staff in an office building do not need to identify themselves when entering the building (for example, by generating an ID card at a key reader fixed adjacent to the door), but can enter without such a cumbersome procedure because the robot 100 simultaneously surveys all people, for example, by acquiring the first surveillance data from each camera. Inside the building, when patrolling the building, and only when the identity cannot be clearly determined from the first surveillance data, is it necessary to generate an ID card for verification information.

[0356] Figure 13 Another example of triggering an action in the case of an uncertain state of an object is shown. In this example, the triggered action is to obtain external additional information about the object. In this example, the robot 100 is part of a monitoring system that includes a third monitoring sensor (a measurement camera 104 installed in a corner of a room 150 (see also Figure 1Due to its position, the camera 104 has a better field of view 124 of the door 108 at its current position P10 than the robot, and therefore the image captured by the camera 104 as the third survey data allows for better detection of the opening 160 of the door 108. According to the communication device 109, the robot 100 either loads a survey image of the door 108 (an image captured automatically by the camera 104 in the form of an image stream captured at regular intervals), or the robot 100 instructs or triggers the camera 104 to capture an image at the current moment (for example, when the third monitoring sensor 104 has a long acquisition interval or is in sleep mode).

[0357] The third monitoring data (the image of the camera 104) provides verification information with which to unambiguously deduce the state of the door 108. The door gap 160 is identified with low uncertainty, so the survey robot 100 has verified that the door 108 is open.

[0358] Other forms of external verification data that are actively obtained are, for example, data from a database stored in a central computing unit of a monitoring system including the measuring robot 100. For example, if it is detected that the box 105 is not present in the room 151 during the previous patrol of the robot 100 (see Figure 1 ), the robot can "query" the central computing unit whether there is available data regarding the delivery of such a box. If the robot obtains confirmation data that box 105 has been recorded or officially delivered, the event "box present in room 151" can be confirmed as "OK." If no data is available in the database regarding box 105, the object's status is deduced to be "risky," and, for example, an alarm is output. Alternatively or additionally, in the event of an ambiguous or "risky" status of box 105, further actions are triggered to further verify this status, such as performing a detailed or meticulous survey of the object using the robot's 100 monitoring sensors, similar to the additional survey procedure described above, and / or triggering interaction between the robot 100 and the object, similar to the interaction procedure described below.

[0359] Other data is data provided by the operator of the monitoring system. For example, the robot 100 reports to the human operator that the detected state has a high degree of uncertainty, and the human operator provides the robot 100 with additional data about the event or related objects, or sends data with instructions for further actions of the robot 100, until the human operator temporarily controls the robot 100 or some parts of the robot 100 (for example, one of the monitoring sensors 110, 111) by sending a control command.

[0360] Figure 14Another example of an action for resolving state ambiguity is shown. In this example, monitoring robot 100 interacts with an object (door 108) to verify whether door 108 is open or closed. In the simple method shown, the robot moves toward door 108 (indicated by arrow 125) and pushes door leaf 123 (without pushing the door handle). If door leaf 123 can be pushed open and door 108 is pushed open, it is verified that door 108 is not completely closed. If door leaf 123 blocks when placed in the door lock, this serves as verification information for the "door closed" state.

[0361] This interactive behavior is learned, for example, by associating robot actions with object states. Similar to the triggered exploration actions described above, criteria are defined for the object states. For example, the state of the door 108 is defined by being able to push the door 108 a certain amount. During the training phase, many action points and action directions are sampled around the door 108, where the door 108 has many different opening angles. For each door opening angle, the optimal action point / direction is defined for the best (lowest) action effort, i.e., where to push the door 108 and in which direction is the best / most indisputable verification information. The optimal action map is learned from these data.

[0362] In the simplest case, the action graph is "move forward." If the robot 100 hits the door 108 and no resistance is detected, the door 108 is expected to open. Possible resistance is detected, for example, by sensors of the robot's drive unit and / or by touch or tactile sensors at the front of the robot or at the robot's 100 arms.

[0363] Action maps can be learned for various objects and various object states. In fact, one can again assume a quasi-static environment and then learn the changes in sensory input and the interactive actions to be performed to infer a certain object state.

[0364] A suitable representation for controlling the robot 100 may be a Markov decision process. A Markov model is a stochastic model of a randomly changing system, where the system is described by states. The states are measured by input data from the robot's sensors, and changes in state depend on the robot's movements. If the logical relationships between states are known, the robot 100 can predict the outcome of any movement and, therefore, the object's state.

[0365] In the case of a monitoring robot 100 equipped with appropriate manipulation tools, interaction with the object whose state is to be deduced may also include pressing or turning a button, knob, or handle of the object, such as the handle of the door 108. In some cases, such more detailed or random actions may be more suitable for verifying events associated with the object via the simpler methods described above.

[0366] Another example of an interaction for generating state verification information ( Figure 6 (not shown) is the output of an acoustic and / or optical signal directed by the robot 100 toward an object. This signal is (potentially) suitable for eliciting a reaction from the object. For example, the object is (meaning the robot 100 suspects it is) a human. The acoustic or optical signal, for example, a message (text) displayed on the robot 100 screen or spoken through the robot's speaker, can then prompt the person to react. Another example of a triggering acoustic signal is the sounding of the robot's horn or siren, which can cause some movement by the person, which can be observed or detected by the robot 100 and used as verification information for verifying the state. This is useful, for example, if the uncertain state is uncertainty about the nature of the object, i.e., if the object is a human (or animal) or an object without the ability to move (such as a human-like statue). Alternatively, the monitoring robot 100 can verify whether the detected state is (not displaying a reaction) or not "the person lying unconscious on the floor" (e.g., reacting by turning their head or opening their eyes). Furthermore, the type of reaction detected by the robot 100 can be used to verify the state. If there is ambiguity in the sense that the status of a person is uncertain as an "intruder" (or "staff / visitor"), the robot 100 issues an alarm signal. If the robot 100 observes that the person reacts by running away, this can be used as verification information for the event being an "intruder" (preferably, not the only verification information, as law enforcement staff may also run away out of fear).

[0367] As already indicated in the above examples, the generation of verification information optionally includes a combination of the robot 100 interaction with the object and subsequent acquisition of second monitoring data (eg, surveying a human's reaction to the robot-human interaction).

[0368] Figure 15a 、 Figure 15b Another example of such an action sequence of interactions and measurements for state verification is given. In this example, a first survey of the door 108 with the first monitoring sensor 110 is accompanied by detection uncertainty if the door 108 is "closed" or "open" because the first survey data (e.g., a picture taken with a camera serving as the monitoring sensor 110) is quite noisy or low-contrast.

[0369] To generate the verification information, the robot 100 first interacts with the door 108 as it applies a (volatile) paint 128 to a portion of the door 108 via a nozzle 127, respectively. Figure 15a As shown. The coating 128 is selected to enhance contrast. Alternatively, the coating 128 is, for example, ultraviolet light or UV-ink (fluorescent coating).

[0370] Then, if Figure 15bAs shown, the robot 100 uses the monitoring sensor 110 to obtain second monitoring data, such as a second image of the door 108, which now has low noise, allowing the state of the door to be detected with no (or almost no) uncertainty (or ambiguity). For example, if the robot 100 has applied ultraviolet ink, a photo is taken while the door 108 is illuminated by UV light 129 from the UV light source of the robot 100 as the second survey data.

[0371] Other examples of applications for such materials are paints or liquids that can highlight the surface of an object, allowing for clear investigation of surface structure or highlighting the outline or shape of an object.

[0372] Figures 16a to 16c Another example of an action triggered for resolving state ambiguity is illustrated. In this example, the robot 100 obtains first monitoring data of the door 108 using the sensor 110. Figure 16a As shown, the first survey data is distorted due to the presence of another object 131 in the room 150 and between the robot 100 and the door 108, and the determination of the door's state uncertainty reveals that the event detection is highly uncertain. For example, if the first survey data is an image of the door 108, the disturbance volume 131 is also imaged, covering the features necessary to determine the door 108 to unambiguously deduce the door state.

[0373] Figure 16b It is shown that an interaction with an object (door 108) is first triggered, wherein a perturbation object 131 is pushed aside by the robot 100 (indicated by arrow 130) in order to improve or free up the robot 100's field of view towards the door 108. It can be seen that the triggered interaction may also include an interaction with another object (perturbation body 131) that is dependent on the object to be measured (door 108).

[0374] Figure 16c A further triggered monitoring action is shown. As object 129 moves away from door 108 due to the triggered interaction, robot 100 uses survey sensor 110 to obtain second survey data, such as a second image, of door 108. Since disturbance body 131 no longer obstructs deduction of the door's state, the state (closed) can be deduced with low or no uncertainty based on the second survey data (state verification data).

[0375] Figures 17a to 25 The invention relates to a patrol system adapted to patrol an area, for example a building. In particular, the system is suitable for autonomous monitoring of the area and detection and reporting of anomalies.

[0376] exist Figure 17a and Figure 17b, a first exemplary embodiment of a combined ground-air sensor platform system 200 is shown. System 200 includes an unmanned ground vehicle (UGV) 210 and an unmanned aerial vehicle (UAV) 220.

[0377] Figure 17a The drone 220 is shown when the UAV is positioned on top of the car robot 210 as a UGV. Figure 17b Drone 220 is shown flying away from robot 210 .

[0378] The robot 210 is equipped with wheels 212 or other means, such as tracks or legs, that allow the robot 210 to move on the ground. The robot 210 includes a housing 211 that houses the internal components and is configured to allow a drone 220 to land on the robot 210. The robot 210 is preferably adapted to serve as a launch and recharging platform for the drone 220. In this embodiment, the housing 211 includes a protrusion 214 on the top of the robot 210 that includes a charging station 216 for the drone 220. The charging station 216 allows the drone 220 to be charged, for example, by means of an induction coil, when the drone 220 lands on the robot 210.

[0379] Alternatively, the UGV 210 may include a battery swap station where multiple battery packs for the UAV 220 are provided and the batteries of the UAV 220 may be automatically replaced with new batteries.

[0380] Drone 220 is a quadcopter that includes rotors 222 that allow the drone to fly, legs 224 that allow the drone to stand on the ground or on robot 210, and batteries 226 for powering the rotors and surveillance equipment of drone 220. The relatively small batteries 226 can be charged while the drone is standing on robot 210. Legs 222 are adapted to properly stand on top of robot 210 to allow batteries 226 to be charged via the robot's charging station 216.

[0381] Figure 18a and Figure 18b 2 shows a combined patrol of two parts of the system 200. A warehouse with a plurality of racks 280 is shown in a top view. Figure 18a In FIG, system 200 autonomously patrols through a warehouse, with the system's camera having a field of view 260 that is blocked by shelf 280 , so an intruder 270 is able to hide behind shelf 280 .

[0382] exist Figure 18b In FIG, a robot 210 and a drone 220 patrol a warehouse, continuously exchanging data 251, 252. The robot 210 and the drone 220 together have an extended field of view 265, allowing detection of an intruder 270.

[0383] The advantage of the drone 220 is that it can fly at different altitudes and thus pass obstacles that block the path of the ground robot 210. Thus, the drone 220 can also fly above the shelves, for example to reduce the time of crossing the warehouse diagonally.

[0384] After detecting a (suspected) intruder 270, the robot 210 can autonomously receive a task to facilitate identification of the person and the request, such as through voice output. The robot can include an ID card reader and / or facial recognition software for identifying the person. The system 200 can also include a user interface, such as for entering a PIN code for identification.

[0385] Instead of a quadcopter or other multi-rotor helicopter, the UAV 220 can also be adapted as an airship that uses a lift gas (e.g., helium or hydrogen) for buoyancy. This is particularly useful for outdoor use. Furthermore, a UAV 220 that uses a combination of rotors and a lift gas for buoyancy is possible. If the UAV 220 uses a lift gas, the UGV 210 can optionally be equipped with a gas tank containing compressed filler gas and a refill station suitable for refilling the UAV 220 with the lift gas when the UAV 220 lands on the UGV 210.

[0386] exist Figure 19a and Figure 19b , a second exemplary embodiment of a combined ground-air sensor platform system 200 is shown.

[0387] The system 200 includes an unmanned ground vehicle (UGV) 210 and an unmanned aerial vehicle (UAV) 220. The UAV 220 can land on top of the UGV's housing 211, in which a charging station 216 is provided that is adapted to charge a battery 226 of the UAV 220 while the UAV 220 is landed.

[0388] In one embodiment, the UGV and the UAV may include a landing system including a light emitter and a camera, wherein the landing system is adapted to guide the UAV 220 to a designated landing pad on the UGV 210. A universal landing system for landing a UAV on a fixed landing pad is disclosed in US 2016 / 0259333 A1.

[0389] The UGV 210 includes a battery 217 inside its housing 211, which provides energy to a charging station 216 for charging the UAV's battery 226 and other electrical components of the UGV 210. These electrical components include a computing unit 218 with a processor and data storage, electric motors for driving the UGV's wheels 212, sensors such as a camera 213, and a communication unit 215 for wirelessly exchanging data with a corresponding communication unit 225 of the UAV 220. Additional sensors may include, for example, a LIDAR scanner, an infrared camera, a microphone, or a motion detector.

[0390] The computing unit 218 is adapted to receive and evaluate sensor data from the UGV's sensors, in particular from the camera 213, and to control functions of the UGV 210, in particular based on the evaluation of the sensor data. Controlling the UGV includes controlling the motors of the wheels 212 to move the UGV through the environment.

[0391] In particular, the computing unit 218 may be adapted to perform simultaneous localization and mapping (SLAM) functionality based on sensor data while moving through an environment.

[0392] The UAV 220 includes a relatively small battery 226 that can be charged while the UAV 220 is standing on the UGV 210. The battery 226 provides power to the components of the UAV 220. These components include the electric motors that drive the rotors 222, sensors such as the camera 223, and the UAV's communication unit 225.

[0393] The camera 223 and other sensors of the UAV 220 generate data 252, which is provided wirelessly to the computing unit 218 of the UGV 210 via the communication units 215, 225 of the UAV 220 and UGV 210. The sensor data from the UAV's sensors is stored and evaluated by the computing unit 228 and can be used to control the UAV 220 in real time by generating control data 251 that is sent to the UAV's communication unit 225.

[0394] Alternatively, the communication units 215, 225 of the UAV 220 and the UGV 210 can be connected by a cable (not shown here). A solution for connecting a tethered UAV to a ground station is disclosed in US2016 / 0185464A1. With this connection, which can include one or more plugs, the UAV can be powered from the battery of the UGV and data 251, 252 can be exchanged between the UGV and the UAV. It is also possible to connect more than one UAV to one UGV. In this case, it is preferred to control the UAV taking into account the position of more than one cable in order to prevent these cables from eventually becoming tangled. A solution for determining the position of the cables of a single UAV is disclosed in US2017 / 0147007A1.

[0395] exist Figure 19c In the embodiment, UGV 210 and UAV 220 include radio communication modules 219 and 229. These modules are adapted to establish a data link to a remote command center to transmit sensor data and receive commands from a central computer or human operator at the remote command center. The data link can be wireless, i.e., a radio connection that can be established, for example, via a WiFi network or a mobile phone network.

[0396] Alternatively or additionally, a communication module can be provided that allows tethered communication with a remote command center. A socket needs to be provided in the environment to which the communication module can establish a connection. The location of the socket can be stored in a data memory of the computing unit. Furthermore, the computing unit 218 can be adapted to detect the socket based on sensor data, in particular in images captured by the system's cameras 213, 223. The UGV 210 and / or UAV 220 can then be positioned relative to the socket so that the plug to which the communication unit is connected can be autonomously inserted into the socket to exchange data with the remote command center.

[0397] Similarly, system 200 can be adapted to connect to a power outlet to recharge the battery 217 of UGV 210. The location of conventional power outlets can be stored in the computing unit's data memory. Furthermore, computing unit 218 can be adapted to detect the power outlet based on sensor data, particularly in images captured by the system's cameras 213, 223. UGV 210 can then be positioned relative to the outlet so that the plug connected to the UGV's battery 217 can be autonomously inserted into the power outlet to charge the battery 217.

[0398] To connect to a data or power outlet, the UGV 210 may include a robotic arm (not shown here). The robotic arm may be operated by the computing unit 217 based on evaluated sensor data, such as images captured by the cameras 213, 223. The arm may include a plug or be capable of guiding the plug of the UGV 210 to the outlet or socket.

[0399] The arm can also be used for other purposes, such as manipulating or picking up objects in the environment. This can include moving obstacles that are blocking the path or opening and closing doors or windows to allow the UGV 210 and / or UAV 220 to continue moving. Switches can also be used, such as to turn lights on and off in a room or to operate an automatic door.

[0400] The arm can also be used to rescue a disabled UAV 220 (e.g., one that has crashed or is out of power) and position it on the charging station 216 of the UGV 210 or transport it to a service station for repair.

[0401] Figure 20 A third exemplary embodiment of a combined ground-air sensor platform system for patrolling a two-story building is shown. Because UGV robots cannot navigate the stairs between the first and second floors, the system includes two UGVs, with a first UGV 210a patrolling the first floor and a second UGV 210b patrolling the second floor. Because the unmanned aerial vehicle (UAV) 220 is capable of navigating stairs, only a single UAV is required. The UAV 220 can patrol the stairs, the first floor, and the second floor, and can be guided by and land on both UGVs 210a, 210b to recharge its batteries. Furthermore, the memory unit of the UAV 220 can be used to exchange sensor data and / or command data between the two UGVs 210a, 210b in different parts of the building.

[0402] Figure 21 A fourth exemplary embodiment of a combined ground-air sensor platform system 200 is shown. UAV 220 is capable of overcoming obstacles (e.g., walls 282) and capturing images of locations 283 that are inaccessible to UGV 210. Sensor data 252, including image data of hidden locations 283, can be sent to UGV 210 for evaluation.

[0403] In the illustrated embodiment, the UGV 210 includes a laser tracker 290 adapted to emit a laser beam 292 onto a retroreflector 291 of the UAV 220 in order to determine the relative position of the UAV. Alternatively or optionally, a camera image-based tracking function may also be provided.

[0404] If the environment is unknown, the UAV 220 may fly ahead of the UGV 210 and provide sensor data 252 to generate a map for path planning of the UGV 210 .

[0405] Both the UGV 210 and the UAV 220 are equipped with GNSS sensors to determine position using a global navigation satellite system (GNSS) 295 (e.g., GPS). To conserve power in the UAV 210 (which has only a small battery), the UAV's GNSS sensor can be selectively activated only when the UGV's GNSS sensor has no GNSS signal.

[0406] Similarly, the UGV 210 and UAV 220 may be equipped with a radio connection to a command center to report unusual or significant events or to receive updated instructions via a wireless data link (see Figure 19b ). To save power in the UAV 210, the radio connection can optionally be activated only when the UGV's radio connection has no signal. For example, the radio connection can be established via WiFi or a cellular / mobile phone network.

[0407] exist Figure 22a 、 Figure 22b and Figure 23a 、 Figure 23b Two further exemplary embodiments of a system according to this aspect of the invention are shown in , wherein a UGV 210 is adapted to provide shelter for a UAV 220 .

[0408] Figure 22a and Figure 22b The UGV 210 includes an extendable drawer 230, which, when unfolded ( Figure 22b ), the UAV 220 can land on the drawer and take off from the drawer. When the drawer 230 is retracted into the housing 211 of the UGV 210, the UAV 220 on the drawer 230 is contained in the space inside the housing 211, thereby providing protection from adverse weather conditions such as precipitation or wind. A charging station or battery exchange station can be provided to allow the UAV 220's battery to be charged or replaced while it is positioned in this space.

[0409] Figure 23a and Figure 23b The UGV 210 includes a cover 235 that is suitable for shielding the UAV 220 when the UAV 220 lands on top of the UGV 210.

[0410] exist Figure 24In FIG, the system includes a UGV 210 and a plurality of UAVs 220, 220a, 220b, 220c. The UAVs are all connected to the UGV 210, and the UGV is adapted to preferably receive and evaluate sensor data from all UAVs (in a combined holistic analysis approach). The UAVs may be of the same type or different in size and kind, as well as in terms of the sensor systems installed. As an example in Figure 24 The four UAVs depicted in FIG2 include two small quadcopters 220 and 220c, one larger quadcopter 220a, and an airship 220b. The larger quadcopter 220a can carry more or heavier sensors than the smaller quadcopters 220 and 220c. For example, the small quadcopters 220 and 220c can be equipped with a simple camera setup, while the larger quadcopter 220a can have a camera with superior optics and resolution, as well as an additional IR camera or laser scanner. The airship 220b requires less power than a quadcopter to maintain position and can be used outdoors, for example to provide an overview camera image of a monitored area from a high vantage point.

[0411] The UGV 210 may control the UAVs 220, 220a, 220b, 220c based on the received sensor data, either directly or by sending commands for specific actions, such as moving to a certain location and taking images or other sensor data of a certain object.

[0412] The UGV 210 may also be adapted to generate a workflow including itself and one or more of the UAVs 220, 220a, 220b, 220c to jointly perform patrol tasks in the surveillance area. Figure 25 The flowchart in FIG. 4 shows an example of such a workflow generation.

[0413] In a first step, the UGV receives a mission to patrol a surveillance area, for example from a remote control center or directly through user input from the UGV. This mission can include further details about the patrol mission, such as which specific sensors should be used or what actions should be performed, for example if a specific defined event such as an anomaly is detected.

[0414] The UAVs of this system are self-describing. The UGV requests mission-specific data from multiple UAVs (and optionally other UGVs) that are available (i.e., within communication range). This mission-specific data includes information about the characteristics of the UAV that may be relevant to the mission. For example, the mission-specific data may include information about the propulsion type, installed sensor components, overall dimensions, battery status, etc.

[0415] In this example, there are three available UAV1, UAV2 and UAV3, so that the method includes three steps that can be performed essentially simultaneously: requesting mission-specific data for the first UAV, the second UAV and the third UAV. Subsequently, the requested mission-specific data of UAV1, UAV2 and UAV3 are received by the UGV, and the mission-specific capabilities of the three UAVs are evaluated by the device. After the capabilities of the three UAVs have been evaluated, a workflow can be generated by the computing unit of the UGV. Workflow data for each UAV involved in the workflow is generated and then sent to the UAV involved. In the example shown, as a result of the capability evaluation, the generated workflow only involves the UGV and two of the three UAVs, so only these two UAVs need to receive the corresponding workflow data to perform their part of the mission.

[0416] In an alternative embodiment, the workflow is generated on an external device and the UGV is treated like the three UAVs of the depicted embodiment. Specifically, the external device is located at or connected to a command center to which the UGV and the UAV's systems are connected via a wireless data link.

[0417] In particular, if the UGVs and UAVs are not all from the same manufacturer, the devices may have incompatible software standards that often prevent them from working together. To address this issue, a software agent can be provided at a single device that converts the data transmitted between the UGV and the UAV into the corresponding machine language. The software agent can be installed directly on the computing devices of the UGV and the UAV, or on a module that can be connected to the UGV or UAV (especially if direct installation is not possible). Suitable software solutions are known in the art and are disclosed, for example, in EP 3156898A1.

[0418] A workflow generation method based on this is disclosed in EP 18155182.1. The compilation of the workflow can be performed on one of the devices, advantageously the device with the most powerful computing unit. Typically, the UGV has the most powerful computing unit and will therefore be used as the commander. In this case, in order to generate the workflow, task-specific data for each UAV is requested from a software agent installed on or connected to the UAV, and the workflow data is provided to the UAV via the software agent operating as an interpreter. Alternatively, the compilation can be performed on an external computing device (for example, a command center located on or connected to the UGV and UAV system to which it is connected via a wireless data link). In this case, in order to generate the workflow, task-specific data for each UGV and UAV is requested from a software agent installed on or connected to the UGV or UAV, and the workflow data is provided to the UGV and UAV via the software agent operating as an interpreter.

[0419] Figures 26 to 29 The invention relates to a safety monitoring system including a state detector, a state filter and a feedback function.

[0420] Figure 26 A general safety monitoring system without feedback functionality is schematically shown.

[0421] State 301 is detected by a state detector associated with one or more survey sensors (not shown). As an example, multiple monitoring sensors (e.g., a person detector and an anomaly detector) can be arranged and linked together into a survey group, such that the group is configured to monitor a specific area of ​​a facility and detect events associated with that specific area. Alternatively or additionally, a state detection algorithm (i.e., a common event detector associated with all monitoring sensors of a monitored site) can be stored on a local computing unit 302 and configured to process survey data associated with the monitored site.

[0422] The incoming state 301 is then classified by a local state filter 303, for example, where the state filter 303 provides an initial assignment 304 of the state 301 into three categories: a "critical state" 305, for example, automatically triggering an alarm 306; a "non-critical state" 307, for example, not triggering an automatic action 308; and an "uncertain state" 309, where an operator needs to be consulted 310 to classify the state 311.

[0423] Alternatively, the "critical status" is also forwarded to the operator for confirmation before sounding the alarm.

[0424] As an example, the state filter 303 may be based on a normal-abnormal classification model in an n-dimensional state space, wherein a state is represented by an n-dimensional state vector, and in particular wherein a corresponding class is represented by a portion of the n-dimensional state space.

[0425] In case of an uncertain state, the operator 310 classifies 311 the state as a critical state 305 and, for example, issues an alarm 306 such as calling the police or fire department, or the operator 310 classifies the state as a non-critical state 307, wherein no action 308 is performed.

[0426] Alternatively or additionally (not shown), the operator may also reconsider the initial assignment 304 of the categories "critical state" 305 and "non-critical state" 307 and reassign the initially classified events to different categories based on the operator's experience and / or certain new rule characteristics for the monitored location.

[0427] Figure 27 The safety monitoring system according to the present invention is schematically shown, i.e., has feedback and training functions.

[0428] According to the present invention, a feedback loop 312 is introduced, wherein the marking 311 of the uncertain state 309 and / or the reassignment of the initially assigned state by the operator 310 is fed back to the state filter 303, as in a typical active learning environment.

[0429] In particular, marking 311 and reassignment of events can occur explicitly, such as based on manual operator input specifically directed to marking a state or changing an assignment, or implicitly, such as where an operator directly issues an alarm or performs a specific action, i.e., without explicitly addressing the state assignment to a specific class.

[0430] Feedback information, such as operator flags and reassignment status, is processed by a training function stored, for example, on the local computing unit 302 or a dedicated computer or server (not shown).

[0431] For example, the state filter 303 can be trained by a machine learning algorithm. Compared to rule-based programming, machine learning provides a very efficient "learning method" for pattern recognition and can handle high-complexity tasks, utilize implicit or explicit user feedback, and is therefore highly adaptive.

[0432] Furthermore, the described monitoring system installed locally at a particular monitoring site may be part of an extended network of many such local monitoring systems operating at multiple different monitoring sites, each local security monitoring system bidirectionally sharing 313 its update information with a global model 314 (e.g., a global state detection algorithm / global state classification model), which may be stored on a central server unit or one of the local computing units.

[0433] Thus, the initial local detection and classification model may have been derived from the global model 314, which contains knowledge about critical states and is a paradigm for all local models. Furthermore, during the operating hours of the locally installed system, the global model 314, including, for example, the global state filter model, may be automatically queried before prompting an operator to make a decision if the local model is unknown, or the operator of the locally installed monitoring system may manually query the global model 314.

[0434] Figure 28 Another embodiment of the security monitoring system of the present invention is schematically depicted, wherein the feedback loop 312 includes not only Figures 2a to 2d , and is extended to also include a global classifier, for example to automatically provide update information for the global classification model 314.

[0435] Since some local states are only locally relevant, for example because different monitoring sites may have different locally defined rules or workflows which may take into account specific access plans and restriction plans for human workers, different hazard areas, changing environmental conditions of the site, and changing site topology, global updates of critical states may optionally be mediated by the update manager 315.

[0436] Hence, special cases of local relevance as well as globally relevant learning steps are considered to provide an improved monitoring system which allows for a more general and robust monitoring and alarming scheme, in particular wherein false alarms are reduced and only increasingly relevant alarms draw the operator's attention.

[0437] Figure 29 Another embodiment of the security monitoring system of the present invention is schematically shown, wherein the state detector 316 is updated in the same feedback loop 312 as the feedback loop 312 used to train the state filter 303, for example by deriving labels (such as "open door", "person", "intruder", etc.) from the operator's decision to mark the state as critical or non-critical.

[0438] Furthermore, similar to the initialization and updating of local state filter 303 and / or global state filter 317, global state detection model 318, which is an example of local state detector 316, can be used to initialize new local state detection algorithms and models to analyze monitoring data 319. Thus, state detector 316 also improves over time, and the local detection model benefits from operator feedback at other monitoring sites.

[0439] Figures 30 to 36b The invention relates to an automatic monitoring system comprising a plurality of monitoring sensors in combination.

[0440] In the following, without loss of generality, the present invention is described using the human detection task as a use case. Figure 30 The specific case of a spatial context model of a warehouse is shown as an example of a part of a building 3 that has to be protected according to the invention. In the image shown, several areas 90, 91, 92, 93 are highlighted in the image.

[0441] For example, a top area 90 is represented by an overlapping matrix of dots. This top area 90 can be ignored by all detectors in this example, since it is very unlikely that a possible intruder will be found there. This knowledge can help to speed up the automatic detection and / or classification task by reducing the size of the search area. By excluding this top area 90, possible false alarms in this area can also be avoided. According to the present invention, masking of this top area 90 can be obtained by different techniques. For example, a manual definition including a human operator can be obtained, e.g. during an initialization and / or commissioning phase. In another more automated example, by identifying a door in the background as a reference, etc., a ground level estimate can be used to determine the floor of the warehouse 3, and based on this information and knowledge about the height of typical people, any area 90 above this height can be configured to be at least partially ignored in the detection and / or classification of people in the warehouse 5.

[0442] Another example according to the invention comprises at least partially learning the top region 90 based on data from monitoring devices in the warehouse 3. Such learning can comprise running the detectors and / or classifiers of one, more or all modalities for a period of time, in particular during regular and / or simulated use of the warehouse 3, for example when staff and / or actors move around the warehouse at different times of the day, with lights turned on or off, etc., to generate training data. The learning system can then detect and / or classify that nothing, in particular no people, was detected and / or classified in the top region 90 during the entire time. Optionally, the results of such automatic learning can be provided to the operator at least once in order to confirm or adjust the top region 90. In such adaptation, for example, false detections and / or classifications, such as those caused by simple false detections, reflections, boxes of mannequins on shelves, etc., can be removed or corrected.

[0443] Some or all of the areas 90, 91, 92, 93 shown in this example. Training data can be collected, for example, with people standing or walking around the facility 3 at different locations and under different conditions (e.g., lights on / off), wherein the positions of the people can be annotated in the training frames, either manually or, preferably, at least in part, by automatically classifying the people as people. As mentioned, this particularly occurs in areas 91, 92, and 93. In another embodiment, one or more people in the warehouse 3 can also be at least partially modeled or simulated in order to automatically generate synthetic training data. This can be achieved, for example, by means of a 3D model of the warehouse 5 or by augmenting an image of an empty warehouse with images of one or more people, in particular wherein multiple options, people, poses, lighting, etc. are automatically synthesized.

[0444] According to this aspect of the invention, based on this training data, the best combination of classifiers and / or detectors is learned. In particular, learning is performed for each setting (i.e., for each location like indoors versus outdoors, a first warehouse versus a second warehouse, a first view and a second view of the same warehouse, etc.), for each environmental setting (i.e., for each time of day, weather condition, etc.) and / or for each given scene (i.e., an area in pixel coordinates). This best combination produced by the learned classifiers and / or detectors can be hierarchical or parallel or a combination thereof. This step according to the invention can be embodied in that not the actual detectors and / or classifiers for the person or object itself are learned, but rather the combination that performs best when a specific context or environmental condition is applied.

[0445] For example, in Figure 30 In the example of Warehouse 3 in Figure 1, RGB images may be more suitable for detecting people at medium to long distances from the sensor, as shown in Region 1 91. Depth, on the other hand, would be a better modality for detecting people within the range of its depth sensor (i.e., typically several meters away), as shown in Region 2 92. Globally, infrared (IR) can be a very discriminative factor for detecting people in Warehouse 3, especially in complete darkness. However, particularly in winter, areas with heat sources (such as Region 3 93 with a radiator shown) may generate a large number of phantom detections due to IR detectors. In Region 3 93, IR detection can be largely ignored, while other modalities (such as point clouds or visual images) can take over. According to the present invention, these aforementioned characteristics do not need to be manually coded for this specific warehouse location, but rather are machine-learned by an artificial intelligence unit. Because these characteristics are reflected in the machine-learning computational entity, they may be implemented in a form that is not directly accessible to humans. Although the description here is presented as a human-understandable logical flow for illustrative purposes, the actual implementation of the optimal combination according to the present invention will likely be implemented at a more abstract level within the machine-learning artificial intelligence system.

[0446] Figures 31a to 31c Different monitoring sensor modalities (in each row) for different regions 1-3 from above (in different columns) are shown. Figure 31a The top region 3 in the IR image modality is shown, where the radiator is bright. Detecting people near a radiator based on such a thermal image tends to be very error-prone or even impossible, requiring correspondingly very specific detection and / or classification methods that significantly deviate from methods used in other regions without such a heat source. The assessability of this pattern in region 3 93 can also depend, among other things, on whether the radiator is actually in use, for example, on the season or other time of day information, on the power status of the heating system, and / or on the indoor and / or outdoor temperature, or other actual contextual information about the location in question.

[0447] In the middle of the column, a visual image of area 3 93 is shown, such as an RGB image from a visual surveillance camera. In sufficient lighting conditions, a machine-learned detector and / or classifier for detecting and / or classifying people can be used. Therefore, the evaluation of this modality can depend on contextual information about the lighting conditions, such as from day / night information, information about the on / off status of lights in the area, lighting sensors, etc. In certain embodiments, contextual information about the lighting conditions can also be obtained from the digital image by the surveillance camera itself, where brightness and / or contrast levels can be obtained. In certain embodiments, the contextual information can also provide preference weightings for different machine-learned detection and / or classification attempts. For example, in darkness, it would be reasonable to train a classifier specifically for detecting people with flashlights, while in bright conditions, it would be reasonable to train a classifier specifically for detecting people in illuminated environments.

[0448] At the bottom, a depth image is shown, which works for the distance of area 3, but may have accuracy drawbacks when its depth camera operates based on IR radiation. Therefore, particularly in the case of heated radiators as described above, such detection is often disadvantageous because it may lead to poor data, false positives, etc. In one embodiment, such contextual information can also be obtained, for example, from the intensity information in the IR image shown at the top, as discussed with respect to the IR image modality. Furthermore, in certain embodiments, based on this contextual information, the specific clusters of classifiers used to detect people (learned by the evaluation unit for the range image) can be weighted differently in the detection.

[0449] According to this aspect of the invention, a machine learning system based on training data that includes and classifies contextual information can learn the above aspects. Therefore, there is no need to manually hard-code all of the above context dependencies, but rather machine learning based on training data. Specifically, the example of this embodiment shown here has learned context-based segmentation based on spatial context (e.g., areas 1-3 and the top area), lighting context (e.g., lights on / off, day / night), and thermal context (e.g., winter / summer, heating on / off). The machine learning system can provide a basic framework for learning, and the actual functions, thresholds, etc. are essentially machine-learned, rather than purely hand-coded by human programmers. This allows, for example, flexible adaptation to different environments and complex systems while keeping the system easy to manage, adapt, and supervise.

[0450] Figure 31bThe same situation as in Area 1 above is shown. The top IR image works well for detection due to the absence of a radiator. In well-lit environments, the middle visible image detection also works well, as described above. In contrast, the bottom range image does not work well, especially since the area indicated by the dashed line is outside the range of the depth camera, which would result in no, or at least no, valuable information for the security system. In the context of this Area 1, since training revealed unreasonable data from the depth modality in this Area 1, the depth modality will be learned to be at least substantially omitted. For example, the size and / or shape of the indicated Area 1 can actually be learned based on the training data, as defined to establish the context of an area where the depth modality would not provide reasonable results and therefore would not need to be evaluated (saving computing power and sensing time) or would at least be considered in the overall monitoring with a lower weight than other modalities.

[0451] Figure 31c The situation in Region 2 is shown. In this region, all three exemplary modalities shown by the top, middle, and bottom images are found to provide detection data, wherein the system learns that in the spatial context of this Region 2, data from all modalities are evaluated substantially equally. However, the system also learns from the training data that other contextual information exists, such as the lighting context, where the intermediate visual image modality is only applied when there is sufficient light, respectively, only the context-specific cluster of classifiers for detecting suspicious or potentially dangerous persons / objects is applied, said cluster being specifically trained in that context, while at least substantially omitting clusters trained in another context (such as the discussed person with a flashlight in the dark vs. a person with rich texture in the light).

[0452] The above is merely an exemplary embodiment of this aspect of the invention, wherein the machine learning of the automated monitoring system is provided with contextual information contained in the training data, particularly (but not necessarily) supervised machine learning. Thus, the system learns the best combination (or selection) of different modalities to apply in a specific context.

[0453] For example, the optimal combination may be illustrated as a hierarchical tree structure. Figure 32 An example of this structure is shown for different spatial and ambient (lighting) environments 80a / 80b. Specifically shown is an example of a hierarchical model learned in different contexts. The example illustrated is the use of spatial contexts (91, 92, 93) from top to bottom, which learn different models for the different image regions 1-3 mentioned earlier. Also shown is the ambient context from left to right, which reflects the lighting context, such as lights on 80a versus lights off 80b. Note that instead of a single tree, a collection of trees or a forest can also be used for each context.

[0454] The tree structure shown is self-explanatory. For example, for region 1 91, if light 80b is turned off, sensor 40b disables RGB information, and RGB modality 81b is likely useless, and the learned model will likely favor infrared modality 81a of IR sensor 40a. Optionally, and not shown here, despite completely omitting information from the visual camera, specialized visual detectors may be present, particularly those trained for flashlight detection, but perhaps not trained for general person detection. On the other hand, the system may also learn that there are constant detections in infrared modality 81a in region 3 93, even when the training data does not include any annotations for people there—thus achieving no valuable surveillance detection. Another context not shown would be possible machine learning discoveries, particularly when the outside temperature is low and / or the heating system is actually activated. In the latter case, such as at night and in winter, the system may learn that depth image 40c (or point cloud 40d) is the optimal modality for the detection task in the context of region 3 93 and in the context of cold nights.

[0455] In a parallel structure, the weighted contribution from each modality is evaluated, where the weights are learned for each context. Figure 33 An example of such an architecture is shown in which training data is used to tune a multimodal fusion algorithm (also referred to as an information filter 83). Specifically, an example of a weighted parallel model for sensor fusion is shown. The weights 82a, 82b, 82c, 82d for the different modalities 40a, 40b, 40c, 40d of the monitoring system depend on the context and are learned from the training data. In this simple example, there is a weight factor W for the infrared detector 40a. IR 82a, weight factor W for vision camera 40b RGB 82b, weight factor W for depth image detector 40c D 82c and the weight factor W for the point cloud detector 40d PC 82d, these detectors are combined in an information filter 83, which in this example results in a binary decision 81 of 1 or 0 reflecting whether a security alert is issued or not.

[0456] Among them, the weight factor W IR 、W RGB 、W D and W PC Rather than being learned as fixed constants, the learning relies on context variables or vectors, where the context can be obtained from the detectors of the modalities themselves, their cross-combinations, and / or based on auxiliary context information or sensors, such as the state of the thermostat, heating system, electro-optical system, ambient indoor and / or outdoor light sensors, etc.

[0457] For example, the brightness, contrast, etc. of a picture from an RGB camera can be numerically evaluated and trained so that, for example, in the case of low brightness or low contrast, the weight factor W RGB will be low (as a self-contained context), while for example the weighting factor W IR and / or W D will be dominant (as a cross-combination of contexts). In practice, the brightness, contrast, etc. evaluated may not actually be so explicitly pre-programmed, but may be machine-learned from the training data, directly learning those picture properties and / or indirectly, e.g. because the confidence level of the detector and / or classifier will be lower on dark or low-contrast pictures.

[0458] In another example, additionally and / or exclusively, external context of the on or off state of the heating system or values ​​from an external temperature sensor may be included in the training data set of IR weighting factors to be applied to the IR modality in zone 3 93 .

[0459] Through machine learning based on this training data, the model can, for example, learn that in area 1 91, when the light is turned on, the weights W of the RGB modalities RGB Should exceed the weight W of the depth modality D etc.

[0460] In other words, one of the main contributions of this aspect of the invention can be described as a context-adaptive model for human and / or object detection in automated surveillance systems. It utilizes training data and machine learning techniques, where the proposed algorithm learns the contexts, such as spatiotemporal segments or clusters, in which various types of detectors using different modalities are most suitable. In particular, the optimal combination of these detectors (e.g., hierarchical, parallel, or other structures) can be automatically learned based on the data extracted for each of the contexts.

[0461] Figure 34a An example is illustrated with a bird's-eye view floor plan of a warehouse 3, wherein an autonomous surveillance system 43 patrols the aisles between the warehouse shelves. The surveillance robot 43 is equipped with at least a visual RGB camera and an IR camera.

[0462] Figure 34b An example of a visual picture taken by an RGB camera in the direction of the arrow during a night patrol is shown. It can be assumed that it depicts an intruder 65 in warehouse 3, but at night, the image has very low brightness and contrast, so in particular, the automatic detector and / or classification will (if any) find a person in the image with very low confidence, which may be too low to sound an alarm because it may not be a person but some random object in warehouse 3.

[0463] Figure 34c An example of an IR image captured by an IR camera in the direction of the arrow during a night patrol is shown. In this example, an intruder 65 is shown and thus more clearly detected. Figure 34b The state in the visible image of the human body, especially with more contrast with the environment. An automatic detector and / or classifier trained for human detection and / or classification will be able to find the intruder in the IR image 40a, which may have a high confidence level, which will be sufficient to classify it as "critical" and issue a security alarm or trigger any other security event. In particular, when the IR detection and / or classification is substantially high confidence with the RGB detection ( Figure 34b ) and / or low confidence matches of the classification, for example, can be obtained by using information filters 83 or Figure 32 According to this aspect of the invention, the nighttime environment, which may be obtained, for example, from a clock, from sensors 42a, 42b, from the RGB image itself, etc., is trained based on training data comprising information about the environment.

[0464] exist Figure 35a In Figure 2, we show an example of an image captured from an RGB surveillance camera inside a warehouse at night. Due to the low light conditions, the image is noisy and contains almost no color information.

[0465] exist Figure 35b In the figure, there are examples of images taken from an IR surveillance camera in the same warehouse and at the same time. Figure 35a In contrast to the multi-megapixel resolution of IR image sensors, the resolution is lower.

[0466] In one embodiment, the various sensors used to record multiple modalities can all be integrated into the same sensor device. In another embodiment, the various sensors can be grouped into one or more independent devices. The positions of all or a subset of these sensors (particularly within a specific sensor group) can be fixed relative to each other and do not change over time. For example, an imaging 3D scanner device (such as BLK360) can include color and infrared sensors and a laser scanner in a single device (or unit or sensor group). Many of these sensors implement a specific spatial reference for their sensing capabilities, such as the spatial reference frame of the camera view, the acoustic reference base of the microphone array, the spatial reference frame of the ranging camera or laser scanner, etc.

[0467] Thus, according to one aspect of the invention, each of these sensors can be calibrated relative to a device coordinate system or a device group coordinate system, which device group coordinate system can optionally also be spatially referenced to a global coordinate system such as a room, building or geographic coordinates. By such a calibration, the sensed state data can be matched across multiple modalities, for example with an intrinsic transformation of the sensing from the different sensors to a common coordinate system. For example, such a calibration can be established based on extrinsic and intrinsic calibration parameters or a geometric transformation from one sensor image coordinate system to the image coordinate system of any other sensor of the setup. In other words, for example, a projection of pixel information recorded by two or more different sensors to a pixel-aligned common image reference frame can be established thereby, in particular also across different resolutions and / or modalities.

[0468] exist Figure 36a In the example shown, an image is captured at night from an RGB surveillance camera at another location in the warehouse, perhaps by a mobile surveillance robot patrolling the warehouse or automatically commanded to that location due to unusual nighttime noises detected in that part of the building. Besides the fact that the scene shown may be unusual, the RGB image does not contain much valuable information. Only some glare and some lens flare caused by the glare are visible, but no information is available that would be much more useful for further classifying the potential safety situation.

[0469] In addition to the RGB surveillance camera, the mobile surveillance robot is also equipped with an infrared camera or thermal imager placed directly below the RGB camera. Each camera has a separate objective lens arranged side by side, thus having a different field of view and viewpoint, and also has a sensor with a different pixel resolution.

[0470] exist Figure 36b In the example of , a view of an infrared camera is shown. According to this aspect of the invention, at least one of the camera images 36a and 36b is provided to a computing unit configured for image processing. The computing unit establishes a virtual transformation of at least one image to establish a basic pixel correspondence between the two images from different cameras, so that after processing, the items in the resulting image are preferably pixel accurate at the same location in both images. In an advanced embodiment, the computing unit can be configured to obtain a four-channel image (the four-channel image includes three RGB channels from a first RGB camera and an IR channel from an IR camera), or to obtain a single image or data set with a single reference frame for all four channels. Therefore, the method can be extended to other modalities and / or more than four channels in the same image, for example by including an additional depth channel (such as from a range camera or laser scanner), including spatial audio information mapped to the 2D image, etc.

[0471] For example, in embodiments of the present invention, a surveillance device can be built that provides 5-channel image or video information, comprising a 2D pixel array of red, green, and blue channels from a camera, plus a fourth IR channel comprising a 2D pixel array from an IR camera, plus a fifth depth channel pixel array comprising range information from a range camera, the range of which is mapped to the intensity information of the image. According to this aspect of the invention, all three modalities (image, IR, and depth) are mapped to a single 5-channel image or dataset, preferably mapped in such a way that the modalities are combined in such a way that the pixels of each of these arrays spatially correspond to the pixels of the other arrays. According to the present invention, such 4, 5, or more channel images can then be used as datasets for machine learning as discussed herein. By utilizing this multimodal four or more channels of spatially aligned data, machine learning results in multimodal detection and / or classification can be improved, for example, because the interdependence of the modalities will be implicitly present in the training and application data (in the form of real-world data and / or synthetically rendered artificial data). In particular, this enables the use of the same methods, algorithms and artificial intelligence structures that are already well established in the field of image processing, but also on other modalities than plain RGB images.

[0472] In an example embodiment, instead of an RGB representation, the visual image can be captured or converted to the well-known HSV (Hue, Saturation, Lightness) or HSL (Hue, Saturation, Lightness) representation. The saturation value can be omitted primarily because, in current surveillance applications, particularly at night, it often does not contain much valuable information. The saturation channel can be replaced or substituted, for example, by a distance channel from a laser scanner or range camera, thereby producing a three-channel HDV (Hue Distance Value) image. In another embodiment, the saturation value can also be replaced by an IR channel or another modality. As discussed previously, this preferably achieves pixel-accurate matching of distance and image information by correspondingly transforming one or both of these elements. Such HDV images can be processed, for example, in a manner similar to, or even identical to, conventional HSV or RGB images in machine learning, evaluation, and artificial intelligence systems. According to this aspect of the invention, surveillance applications can achieve improved detection results by using HDV images instead of HSV or RGB images. Furthermore, data volume can be reduced, as only three channels are required instead of four.

[0473] Figures 37a to 43b Related to surveillance systems with machine learning detectors and / or classifiers.

[0474] Figure 37aAn example of a part of a building 3 that must be protected according to the invention using a surveillance robot 43 is shown. The main building part 50 shown here is mainly rooms, to which doors 51, 51a, 51b (as examples of building elements) lead, which doors must be closed at night.

[0475] exist Figure 37b Shown in Figure 37a The subsection labeled shows an example of a mobile surveillance robot 43, which includes camera 40b and a fixed (but potentially tiltable and / or motorized) surveillance camera 40a. In practical implementations, one or both of these surveillance devices may be present. By observing the corridor, security threats can be automatically detected and classified. In this example, the door is specifically inspected.

[0476] Figure 38a An example of a picture that camera 40a may take is shown, showing door 51b in a desired closed and normal state, particularly at night.

[0477] Figure 38b An example of a photograph that may be taken by camera 40b is shown, showing door 51a in an abnormal partially open state and a potential safety event that would be detected and classified by the surveillance system as a potentially critical condition.

[0478] Figure 38c Another example of a photo that camera 40b can capture is shown, showing door 51a in a normal state. However, an obstruction 60 remains on the floor next to door 51a, which is unusual and a potential security event to be detected and classified by the monitoring system. According to this aspect of the present invention, the detection of obstruction 60 can be classified as simply a common laptop bag 60, forgotten by a known worker. The system can automatically assess that bag 60, in its current location, does not qualify as a dangerous security alert, but may, for example, only generate a low-level security alert or a simple log entry. In another embodiment, this assessment can be programmed or trained to classify differently, such as in a public place where unattended luggage could be used in a bomb attack, or if bag 60 is located on an escape route that must remain unobstructed, or in a location where such bag 60 is unlikely to be forgotten (e.g., hanging from a ceiling). However, in this image example, the left fire door 51b is shown open. Once this event is detected and classified as causing a critical situation, further inspection and corrective action may be initiated.

[0479] As described above, the present invention can establish detection and / or classification of such safety conditions through an artificial intelligence system that has been trained to find and identify such conditions in data from one or more sensors of a monitoring system, which includes devices such as 40a, 43, etc.

[0480] Specifically, machine learning methods can be developed, specifically including classifiers and / or pre-trained detectors. In many embodiments, visual or visualized sensor data such as RGB images, depth images, IR images, etc. are used as sensors, although the aspects discussed herein are also applicable to any other surveillance sensor data information. For example, image content can be represented in a vector space, such as Fisher vectors (FV) or vectors of locally aggregated descriptors (VLAD), and classified using various classification techniques, such as support vector machines, decision trees, gradient boosted trees, random forests, neural networks (including deep learning methods such as convolutional neural networks), and various example-based techniques, such as k-nearest neighbors, US2015 / 0178383, US2017 / 0185872, and US2017 / 0169313. To detect and / or localize objects in larger scenes, regional variants of the proposed methods can be used, such as sliding windows / shapes, R-CNN variants, semantic segmentation-based techniques, etc. Additional modalities such as thermal images and especially depth images can be used directly in object detection and recognition, such as sliding shapes or as additional image channels.

[0481] As described elsewhere, such a system requires training data, preferably annotated or labeled with meta-information describing the content of the training data, in order to teach the system. According to the present invention, such meta-information may include information reflecting the basic categories of items in the sensor data, such as doors or bags as categories for classification, but may also include sub-categories and other meta-information, such as the specific type, size, color of the bag, up to information about physical constraints that are common for such bags, such as being placed on the floor or on a chair or table, being carried by a person, etc. That meta-information may also be learned during the training phase, for example in a supervised learning, where the meta-information is used as a basis for obtaining supervision information. In other embodiments of the present invention discussed elsewhere, semi-supervised or unsupervised learning methods may also be used, in particular, a combination of unsupervised, semi-supervised and / or supervised learning methods for different machine learning aspects in the monitoring system may also be used, for example, depending on whether data and / or algorithms for supervision are available for a certain aspect.

[0482] The detected real-world objects can, for example, be classified as specific objects, be classified within one or more specific object classes and / or conform to or include other object-specific attributes and / or object-related meta-information. For example, such automatic detection and classification can be performed by a computing unit or processor including a detector and / or a classifier to which the digital real-world image from the camera is provided. Based on this information, an algorithm (e.g., a person detection algorithm, etc.) can be applied, and it can detect specific states, such as the presence of a person, the opening of a door, etc.

[0483] In conjunction with one or more of the aspects presented herein (but also considered as specific inventions in their own right), this aspect of the invention is directed to improving the generalization capabilities of automated decision-making processes in surveillance systems. The invention relates to enhancing the training set for such automated decision-making by synthetically generating or synthetically enhancing training data, for example implemented as a neural network.

[0484] Besides the possibility that training states and scenarios cannot be recreated with reasonable effort, such synthetic data can also include the additional advantage that once the generating function is implemented, any number of samples can be generated based on it. The synthetic samples can then be used to prepare the decision maker, especially together with real data samples.

[0485] Another major advantage is that it is much simpler to provide corresponding meta-information associated with each piece of training data in an at least partially synthetically generated approach. While in the case of training systems based on real-world images, an operator usually has to manually classify the training data, e.g., define what the data actually shows, whether it is normal or abnormal, whether certain restrictions should be imposed, etc., in the case of at least partially synthetically generated training data, this classification and / or labeling of the data can also be automatic, as is clear from the model that generated the data, as shown in the current data example, which is preferably automatically obtained from the model, at least for a large set of synthetically generated data.

[0486] For example, given an application where the task is to determine whether something is normal (e.g., a security door closed at night) or abnormal (a security door open), the expected distribution of training samples captured in a real-world environment is typically shifted towards the normal case (of a closed door). Collecting samples of abnormal cases can be a significant effort, as the installation must be changed in various ways to reflect a reasonable number of different abnormal cases that could actually occur. For particularly critical or high-risk applications in high-security areas (i.e., prisons, military areas, nuclear power plants, etc.), the collection of even a few abnormal data samples may be excessively risky or even impossible. However, especially in such applications, it is very desirable to successfully detect those events (even if such data could not actually be observed during training) for which at least a reasonable percentage of the different training samples collected are not present.

[0487] In contrast, synthetic samples according to this aspect of the present invention can be freely generated in various implementations of abnormal states to assist decision makers, automatic state detection, or criticality classification devices in making their decisions. For example, such synthetic generation of virtual samples allows the generation of the same amount of training data for abnormal states of security events as for normal states.

[0488] Task-specific synthetic data, for example, generated virtually from numerical representations of images or other data from computer models, not only has the advantage of allowing the synthetic generation of difficult-to-obtain data, but also has the advantage that the data collection effort is essentially constant relative to the desired number of samples. A common prior art approach to collecting data is to travel to the object in question and extensively record data in different object states, such as with varying lighting, with partial obstructions, and so on, particularly from different viewpoints. The amount of data thus created is directly proportional to the time spent on recording and the effort expended in modifying the scene to reflect the different object states.

[0489] This is very different from the synthetically created data according to this aspect of the invention. One must diligently create the generating functions, models, parameter modification strategies, and so on. However, once this is done, any number of data samples (both abnormal and normal) can be automatically generated for training without requiring additional manual effort. For example, a high-performance computer or server can essentially establish this training data generation in a largely unattended manner. This is particularly useful in transfer learning methods.

[0490] Below, various use cases for the present invention are described using the example of an autonomous security patrol robot discussed herein, but this aspect of the invention is applicable not only to mobile units but also to fixed installations. The first use case is anomaly detection for high-risk applications. Consider the task of monitoring a warehouse. An autonomous security agent takes a predefined route and reports status or anomalies along the way. The route can, for example, be a predetermined patrol route or a route automatically learned by the security agent from training data (e.g., with the goal of covering all warehouses or at least potentially critical building components such as doors, windows, etc.) and / or can be event-driven based on anomalies detected from other sensors, or even a primarily random path learned by the autonomous navigation unit. In this case, examples of anomalies could be a door or window that is open but should be closed, a person that should not be there, or a misplaced item (i.e., a box blocking a hallway), etc. In order to accurately detect any of these anomalies with an automated system, a large amount of data is required for each different subtask (person, door, or misplaced item detection) in as many variations and combinations as possible, in order to have a reasonable diversity in the training data.

[0491] The creative aspect of at least partially synthetic generation of data samples can bring advantages over mere manual field recording. Figure 3 As shown, the representation of the acquired object in such a synthetic model (e.g., a digital 3D model of the object on a computer) can be changed and modified as needed. In particular, an example of a digital 3D model of a door is shown, for example, represented in CAD software. For an illustrative example of detection and classification using visual camera images of a surveillance system, the digital 3D model can be digitally rendered by a computing unit to synthetically generate a plurality of artificial images of the door 51 as training data, for example, Figure 39b shown.

[0492] In the generation of training data, there are many parameters that can be changed, such as viewpoint 62 and / or lighting 63. Thus, a whole series of virtual or synthetic training data can be synthesized, in many cases reflecting doors 51 as they can actually appear in the camera image. In addition to changes in viewpoint, light and shadows, for example, parameters of the object itself can also be modified, such as its opening stage, material, color, size, proportions. All or at least most of this can be done automatically by the computing unit in a basically unattended workflow. Thus, as described above, not only can individual synthetic training data be automatically generated, but the training data can also be automatically processed by automatically generating metadata indicating the content of each training data, which metadata can then be used for classification of each training data or for training specific aspects of those reflected in the metadata. For example, the digitally rendered images 64a, 64b, 64c to 64n reflect some examples of different variations of the 3D model of the door 51 described above. The computing system can generate thousands and more of such training data items, which are used according to this aspect of the invention to train the security system according to the invention. For example, by using Figure 3 The system trained with the virtual training data shown is used to automatically evaluate Figure 2a 、 Figure 2b or Figure 2c The computing system can generate not only static scenes, but also sequences or scenes of safety-related processes of events to be detected and / or classified.

[0493] Doors 51 (especially critical doors) can be modeled as open or ajar and recorded virtually from many different viewpoints without having to risk actually physically opening the real door. This synthetic generation can also help overcome the bias of the distribution of acquired real-world data samples towards the normally closed state. Since the anomaly of an unclosed door is very rare in the real world, it can be expected that real-world data acquisition will only result in very few samples (unless substantially more effort is invested to manually create such anomalies with a higher abundance than naturally occurring). However, for synthetic data, the distribution can be set arbitrarily as needed to achieve a well-trained classifier and / or detector.

[0494] In another example of an embodiment, consider a surveillance robot equipped with more than one of the various surveillance sensors discussed herein, particularly, for example, an RGB camera and a depth sensor. Such a multimodal detection system typically uses information from each available sensor and combines or integrates the data to produce more informed decisions than would be possible if only a single sensor were considered. However, the multimodal system must first learn how to interpret the (sometimes even contradictory) information from those different sensors. The correct interpretation must be derived from the collected data. To prevent overfitting during the learning process, additional data can be helpful in many embodiments to achieve more realistic and robust results. For example, additional data (such as noise, interference, etc.) that is not directly valuable and is used to train the classifier and / or detector in the direction of its actual intended use can also be randomly appended to the training data. Thus, the classifier and / or detector can be tuned to more realistic readout mechanisms of actual sensors available in the real world, while also addressing potential flaws, biases, and failures of real-world sensors. Such additional data can be obtained from real-world sensors and / or can be simulated. Synthetic data creation according to this aspect of the present invention can be a useful tool for such scenarios, where multimodal input data is required in large quantities. Furthermore, in this synthetic data creation, the configuration and calibration of the sensors with respect to each other can be freely set and also changed if necessary.

[0495] In the example of door status detection in the exemplary warehouse environment described above, the present invention can be implemented as follows using RGB and depth sensor configurations. In 3D modeling software, the environment of the warehouse can be created using the desired monitored object included as a virtual model of the door. The scripting mechanism thus provided can be configured to provide the ability to procedurally change the appearance of the object in question. This can include, for example, visual appearance due to material properties (i.e., metal surface, wood, or plastic), lighting configuration, viewpoint and frustum of the sensor used to capture the warehouse scene. On this basis, physically correct materials can be applied when creating photo-realistic rendered images. For example, ray tracing can be used to derive the RGB channels and depth of the scene from the model to create data similar to that captured by a depth sensor. Such depth images can be obtained as real-world pictures using measuring or metrology instruments, such as by using RIM cameras, laser scanners, structured light scanners, stereo imaging units, SLAM evaluations, etc. For training, such depth images can also be reproduced from virtual digital 3D models by known methods, similar to the reproduction of 2D images already described. In particular, the present invention can work with a combination of 2D and / or 3D pictures and images to provide information for automatic detection and / or classification. Another option for such 3D information is to use synthetically generated point cloud data, either directly or by obtaining its depth image. Optionally, artificial noise can also be added to one or more sensors. For example, for a depth sensor, the usable range can be artificially defined using random noise that is simulated and added outside of this range, which can even be determined with an appropriate distribution as needed. The synthetic data created in this way reflects all the variations that are encoded in an appropriate and definable amount, and can thus produce a more robust classifier after training than would be achievable through pure real-world training.

[0496] The same principle of the method can also be extended to other modalities and sensor data. For example, data from an IR camera sensor can be synthetically generated in a similar way by giving a person a different texture appearance based on simulated body temperature, and thereby Figure 40a The mobile surveillance unit 43 shown in the plan view of FIG generates a thermal image 66 of an intruder 65 in the factory hall 3 virtually. In this case, for example, different clothes can be simulated to absorb some simulated body heat, thereby generating reasonable data samples that can be used for training. For example, Figure 40b shows that the calculation unit is based on Figure 40a, a virtual synthetically generated IR image of a person digitally rendered using a digital model shown in . According to the present invention, multiple such synthetically generated IR images 66 can be generated based on this by varying the lighting, pose, person, clothing, environment, viewpoint, etc. in the digital model, wherein the variations are preferably performed automatically by a computing unit, for example, based on a subroutine script. This additional information can be represented in the picture and / or image, for example, in the form of additional "channels" that are considered in detection and classification in addition to, for example, the red / green / blue (RGB), hue / saturation / value (HSV), hue / saturation / luminance (HSL), or hue / saturation / intensity (HSI) channels of the visible image or image. Automatic detectors and / or classifiers trained on such synthetically generated IR images 66 can then detect and / or classify intruders with high confidence.

[0497] In another specific embodiment according to the invention or in combination with the aforementioned method, there is direct training of detectors and / or classifiers based on the digital 3D model itself, without explicitly rendering images and providing those rendered images as training resources. For example, the 3D model itself (e.g. in the form of CAD data, point cloud data, mesh data, etc.) is fed as training input. As an example, a vector representation of an image can be approximated directly from the 3D model or its processed representation (e.g. a voxel grid). Another possibility is the use of a technique of feature predictors that utilizes defined perspective and lighting descriptions to predict a subset of features using local geometry and material information. Moreover, similar techniques can be used to enhance representations of already existing image content to, for example, appropriately simulate the noise of a process or simulate objects observed in different environments. For this approach, deep learning methods can be used in particular, for example including neural networks. For example, when training such a neural network, the first hidden layers can be activated directly from the 3D model and the responses expected from real images of the object are created. In field classification of real-world data, these first layers are replaced by one or more layers based on real-world images. For example, a combination of synthetic 3D models and real-world image data can be used to fine-tune the connections between new layers and the old network.

[0498] like Figure 41Another example of an embodiment of the method according to this aspect of the present invention, shown in the plan view of a location 3 in FIG, is object surveillance using various installed sensors (e.g., surveillance cameras 40a, 40b, 40c). Each camera 40a, 40b, 40c may only observe a limited (and sometimes even fixed) portion of the overall location, property, building, structure, or object covered by the surveillance. For example, using suitable automatic detectors and / or classifiers for people, each of these cameras 40a, 40b, 40c can create a local profile for each person walking through its field of view. In combination, such a system including multiple cameras 40a, 40b, 40c can automatically track, obtain, and analyze each person's trajectory 67a, 67b, 67c, preferably also classifying the person as a specific person or belonging to a specific category or class of people. Based on these trajectories 67a, 67b, 67c, common behaviors of people can be automatically extracted and / or learned by an artificial intelligence system. In addition to spatial and / or image information, such methods may also include temporal information in the training data. This is also referred to as synthetically generated scenes or sequences, which can be used as training data for detectors and / or classifiers in a surveillance system according to the present invention.

[0499] For example, normal trajectory 67a always starts at the entrance door of a building and continues to a person's office, then arrives at a coffee shop, a lounge or another office and returns to one's office by trajectory 67b. Later, as shown in trajectory 67a, the building leaves through the main entrance door again.

[0500] In this case, the natural constraints prevent unrealistic trajectories 67c from occurring, such as the trajectory seen in the second floor where stairs or elevators are not used. An automated decision maker or classifier can then analyze the trajectories and detect whether they are normal or abnormal, classify them as critical or not critical, and decide to trigger some action if necessary.

[0501] However, many trajectories may be incomplete, due to possible sensor failures or simply due to occlusions that prevent detection of people, as in the example shown, where there is no camera at the staircase. In order for the decision system to determine whether a partially observed trajectory is normal or anomalous, a large set of labeled data is required, for example using supervised machine learning methods. In this embodiment as well, synthetic data according to this aspect of the invention can be applied to advantage, particularly with a view to automating not only detection and classification, but also the training of the detection and classification units used thereby.

[0502] First, virtual scenarios such as those described above can be easily created and simulated through synthetic modeling and, optionally, sampling plausible trajectories from real-world observations. Sensor failures, partial blockages, and the like can be naturally incorporated by omitting portions of the sampled trajectories from the modeling. In addition to normal scenarios, there can also be abnormal scenarios derived from simulations that include anomalies, such as trajectory 67c of a possible intruder. For example, such abnormal constellations can be generated by carefully disabling some of the natural constraints in the model. This can include things like the camera view of simulated trajectory 67c, which picks up a person on the second floor who did not pass through the main door because they entered the building through a window. Thus, before such a critical situation occurs, the automated situation detection system can pre-learn the resulting characteristics from the surveillance data that constitutes abnormal trajectory 67, allowing, for example, a supervised classifier to specifically learn to classify and / or detect such situations. This can include transfer learning, but according to the present invention, virtual sensor data can also be synthetically generated from a digital model to serve as training data customized for the actual location.

[0503] For example, another embodiment of such a state pattern or event sequence at an office location that can be synthetically generated from a virtual model may include the following events: an "open door" event, a "turn on lights" state, and a "turn on PC" event. In the morning or daytime, this simulated state pattern is learned by the machine as normal. At night, this may be quite rare, but it may not be important considering that someone is working overtime, etc. However, for example, the above sequence without the "turn on lights" event is not critical during the day, but at night, it can be learned by the machine as a critical state because it is likely to be a thief stealing data.

[0504] In other words, in one embodiment, it is not the state itself that can be considered critical or not, but rather a state comprising multiple detections (and optionally also contextual data) can be defined to be classified as critical. Optionally, it can also be trained to automatically trigger a certain action and / or which action is sufficient for a certain critical state. Detecting an anomaly only means that there is a significant deviation from the normal data in the sensed data, but it does not necessarily mean that this is critical in this context, thus establishing a classification of the entire state.

[0505] In state detection, specific states are detected in sensed survey data (such as in images or video streams) based on the application of algorithms, such as the state of human presence through human detection algorithms, the state of door opening (for example, through a camera or through a door switch, etc.).

[0506] Based on the combination of one or more states, a classification as critical or non-critical can be established, but in many cases this also depends on contextual data, such as location in the building, time of day, floor level, etc. If the state is critical, action must be taken, such as calling the police or fire department. This action can be initiated automatically by the system or after confirmation by an operator working in the central security company. Non-critical states can be ignored, but can nevertheless be recorded in the surveillance storage system, for example, in the form of state information, actual images of the state and its classification, date, time, etc. This log data can be used, for example, to fine-tune the detection and / or classification algorithms learned according to the present invention.

[0507] As another example, according to this aspect of the invention, the state "open window in room 03 (first floor)" and the state "door of room 15 open" at 3:15 am can be artificially generated, but there is no state "light in room 15 on", and the resulting state can train a classifier to be classified as a critical state requiring action, such as indicating an intruder.

[0508] An example of a non-critical state could be a sequence where a person enters the building through the main entrance (door open state - door 01), walks to the office door, opens the office door (approximately 23 seconds after door 01, door open state - door 12), turns on the light (approximately 2 seconds after door 13, power on event), walks to the desk, turns on the PC (approximately 10 seconds after light on, power on event), etc.

[0509] In addition to or as an alternative to raw monitoring sensor data, these event patterns can also be generated synthetically or simulated, for example, manually or based on an automated process. Manual means that a person performs a prescribed process (opening a door, turning on a light at a certain time of day, etc.), and then captures and labels the resulting state pattern, for example, as a non-critical state. Semi-automatic modification of the program to defined variables or environmental conditions can be performed automatically to achieve greater variability in the training data. Automated processes for synthesizing state patterns are also possible, for example, based on rule-based systems, expert knowledge, expert systems, etc.

[0510] According to the present invention, for example, a rule-based expert system can synthetically simulate these states in order to generate virtual training data for training. In the evaluation of real-world data while monitoring the security of a building, events such as "turn on the light" or "turn on the PC" can be obtained directly from the physical device (for example, in a smart home environment, through log files from a PC, through a server or network, etc.). Alternatively, such real-world data can also be obtained indirectly, for example by monitoring the power consumption at the mains power supply (etc.) and automatically detecting change events in the electrical power consumption processed by the lights and / or computers. For example, power-on events can be detected in readings from a central power meter that measures power consumption in the entire building, on a floor, or by means of power sensors for rooms, sockets, etc.

[0511] Optionally, such monitoring of power consumption etc. may include automatic identification of (potential) causes of specific changes in power consumption, which may also be machine learned based on real and / or virtual training data to establish automatic detection and preferably classification of such changes.

[0512] In some common implementations of this aspect of the invention, machine learning may be denoted by the term "supervised learning" because the 3D model contains or is linked to metadata that can be used as supervisory information (e.g., defining the categories of synthetic objects to be learned for recognition or classification and / or the positions or bounding boxes of synthetic objects for their detection, etc.). In many implementations, the invention can utilize traditional machine learning methods (in the context of so-called "shallow learning"), in particular, for example, random forests, support vector machines, because there are predefined features, descriptors, labels and / or feature vectors, preferably defined as meta-information together with a digitally rendered learning source, in particular in an automatic manner. Other implementations may at least partially implement so-called "deep learning" methods, generative adversary networks (GANs) or other artificial neural networks (ANNs), etc.

[0513] Artificial training data generation can be implemented on a computing system, or at the most real-world level, by actually simulating the rendering of the corresponding camera images as training data obtained from the digital 3D model of the venue 5, but can also be implemented at an abstract level. Similarly, by simulating abstract trajectories within the floor plan, the detector and / or classifier that is trained to work on the person trajectories 67 obtained (also at least partially) from another detector and / or classifier configured to detect people obtains such trajectories 67 for each person and tracks the number and the entry and exit of people in the venue. In other words, embodiments of the system are trained using abstract trajectories rather than camera views.

[0514] Applicable to Figure 41The example of FIG. 5 illustrates such monitoring of an office building, wherein exemplary trajectories 67a, 67b, and 67c of a person moving between rooms are shown. The person resulting in dotted trajectory 67a enters the building through gate 68 before heading to their office. The same person then visits another room along 67b. The person resulting in the unobserved dotted trajectory 67c enters the building through large opening 68. This person may be a potential intruder, who made his or her way through window 69. The classifier trained as described above is configured to recognize such situations and automatically generate specific alerts or actions.

[0515] Another exemplary embodiment used to illustrate this aspect of the invention is the synthetic processing of normality examples of a training dataset to also include anomalies, where the training dataset can be a real-world dataset or a synthetically generated dataset, as well as a combination thereof. This can, for example, include person detection, where images of real-world or simulated people are automatically altered in a way that at least partially obscures the person's face or body, as might be the case for an intruder wearing a mask. By automatically varying the amount and manner in which the training data is perturbed, the robustness of the resulting classifier can be improved and / or undesirable overtraining can be avoided.

[0516] For example, Figure 42 As shown, a potentially safe state of damaged window 71b is shown, as it may be caused by a thief. For training of detectors and / or classifications of such states, it is not reasonable to physically destroy windows at a building just to obtain training data images to allow the automatic detection unit 43 to learn such states.

[0517] This embodiment can learn not only based on visual images of broken windows, but can also alternatively or additionally learn based on thermal images 66b, for example, due to temperature variations caused by hot (or cold) air passing through the hole in the window. Such thermal images can also be applied to doors or windows that are not completely closed, as these will regularly result in the ventilation of air with different temperatures, which can be visualized in thermal images. According to this aspect of the invention, such thermal training images can be trained based on a simulation model that easily allows for the simulation of a wide range and combination of interior and exterior temperatures, which are nearly impossible to capture by capturing real-world thermal images, or which would require extended capture over many seasons. According to this aspect of the invention, visual and / or thermal images of the broken window are artificially generated, for example, digitally reproduced or augmented onto real-world images. Such visual and / or thermal images can be generated with a wide range of parameters for different environmental conditions (e.g., lighting, color, texture, glass type, etc.) without significant effort. These virtually generated or augmented images are then used to train detection and / or classification units, for example, in a supervised learning approach. The resulting detector and / or classifier can then be loaded by a surveillance detector, which is thereby configured to detect and / or classify such security events based on real-world photos taken by an automated building or property surveillance detector that has been specifically trained on the virtually generated visual (or thermal) appearance of such events. In another example, rather than training on a specific window 71a, the general aspect of broken glass can be trained on a broad range of virtually generated training data including broken glass in many different variations. Thus, for example, broken glass in door 51 can also be automatically detected and classified, and can also be trained that such broken glass typically indicates a condition that must be indicated to security and / or service personnel, and that such a condition, when occurring at night, would be classified as critical and require triggering action.

[0518] like Figure 43a As shown, in one possible embodiment, synthetic training data can be generated by combining multiple virtual objects, preferably in multiple different combinations, where the virtual objects themselves vary, under different lighting conditions, with intentionally introduced disturbances, etc. For example, there can be one or more backgrounds 70r, such as 3D models or 2D images, specifically pictures of the specific location where the security system will be installed. Those backgrounds can also be varied, for example by simulating daytime and nighttime. The figure also shows disturbances 70f, such as grass growing in front of a building. In this example, there is also a person 70p.

[0519] exist Figure 43b, two examples of synthetically generated (rendered) training images according to this aspect of the present invention are shown. The upper image 64n shows a normal example that is labeled as normal or non-critical, and for which no alarm should be issued. There is a foreground 70f, a background of a building 70r to be measured and the sky, as well as a person 70p. When the person 70p is modeled as being on a regular path to the building 70r, such a rendering is trained to be normal. Therefore, the rendering can be reproduced using the person 70p in different locations, using different people 70p, during the day and at night, and so on.

[0520] In the lower image 64n, the synthesis is based on the same subject, but a person 70p is attempting to climb up a building and enter the building through a window. This is trained as a high-level critical safety situation requiring immediate action.

[0521] In other words, embodiments of the present invention can be described by building a synthetic pre-trained detector and / or classifier by obtaining multiple numerical renderings from a 2D model and / or a 3D model and feeding these renderings as training resources to the classifier and / or detector for supervised learning. The renderings may include at least one object of interest to be trained, preferably embedded in a virtually generated real environment, in particular from multiple different views and / or different lighting conditions and / or environmental conditions. Thus, the general classifier and / or detector is trained using virtual information.

[0522] In an optional further stage of this aspect of the invention, such a general classifier and / or detector can be additionally post-trained using real-world images. This can be done not only during the initial training phase, but also during field use of the classifier and / or detector. For example, real-world images to which the detector and / or classifier is applied can be used as an additional training resource to enhance the detector and / or classifier, for example to improve its real-world success rate, particularly using real-world images of conditions being detected and / or classified as critical. This real-world information is particularly advantageous because there is often feedback available to the user or operator that confirms or corrects the results of the automatic detector and / or classifier applied to the real-world images. This feedback can be used as metadata and classification information for learning and building upon improvements in the learning of previously learned detectors, classifiers, or artificial ...

Claims

1. A monitoring system (103) for monitoring a facility (102), the monitoring system (103) comprising: A monitoring robot (100) comprising a main body (100b), a drive system (100a), and a motion controller for controlling the motion of the monitoring robot (100). The monitoring system (103) further includes: at least a first monitoring sensor (110) designed to acquire first monitoring data of at least one facility component (105-108, 150-152) of the facility (102), a state detector (100c) configured to detect at least one state (114) associated with the facility component (105-108, 150-152) based on the first monitoring data, It is characterized by The state detector (100c) is configured to: Note the state ambiguity (115) of the state (114), which is between two or more possible states and / or is a result of the object condition being previously unknown to the surveillance robot and / or a result of unfavorable conditions for survey data collection, and Upon noticing the state ambiguity, the action controller triggers an action (117, 120, 130) of the monitoring robot (100), the triggered action (117, 120, 130) being adapted to generate state verification information about the facility component (105-108, 150-152), the state verification information being adapted to verify or falsify the detected state and resolve the state ambiguity, and taking into account the state verification information to resolve the state ambiguity, The triggered action (117, 120, 130) includes acquiring, by the monitoring robot (100), second monitoring data of the facility components (105-108, 150-152), and The triggered action (117, 120, 130) includes changing an acquisition position and / or direction so that acquisition of the second monitoring data is performed using at least a second acquisition position and / or direction (P20) different from the first acquisition position and / or direction (P10) for acquiring the monitoring data.

2. The monitoring system (103) according to claim 1, It is characterized by Note that the state blurring (115) of the state (114) is performed by comparing it with a predetermined blur threshold.

3. The monitoring system (103) according to claim 1, It is characterized by The state detector (100c) is further configured to plan actions (117, 120, 130) to be triggered such that the actions (117, 120, 130) are optimized with respect to the generation of verification information.

4. The monitoring system (103) according to claim 1, It is characterized by The state detector (100c) is configured to determine the second acquisition position and / or direction (P20) in such a way that the second acquisition position and / or direction (P20) is optimized with respect to the generation of state verification information.

5. The monitoring system (103) according to claim 4, It is characterized by The state detector (100c) provides a correlation map relating state ambiguity to acquired positions and / or directions (P10, P20), wherein The relevant mapping is based on a defined standard representing the state (114), and / or The correlation map is established by machine learning, and / or The correlation map includes optimal acquisition positions and / or directions (P10, P20) for a plurality of detectable states (114).

6. The monitoring system (103) according to claim 1, It is characterized by The monitoring system (103) includes at least a second monitoring sensor (111), and The state detector (100c) is configured to determine which monitoring sensor (110, 111) is used to acquire the second monitoring data, so that the generation of the second monitoring state verification information is optimized, and / or The first monitoring sensor (110) and the second monitoring sensor (111) are different types of sensors, wherein the second monitoring sensor (111) is adapted to acquire monitoring data with a higher resolution than the first monitoring sensor (110).

7. The monitoring system (103) according to claim 1, It is characterized by The triggered action (117, 120, 130) is to obtain data about the facility components (105-108, 150-152), wherein the data is third monitoring data of a third monitoring sensor (104) of the monitoring system (103), the third monitoring sensor (104) not being part of the monitoring robot (100), wherein the triggered action comprises triggering acquisition of the third monitoring data, and / or Stored in the database of the monitoring system (103).

8. The monitoring system (103) according to claim 1, It is characterized by The monitoring robot (100) is an unmanned ground vehicle (UGV).

9. The monitoring system (103) according to claim 8, It is characterized by The unmanned ground vehicle (UGV) includes a drone (UAV) as a subunit, wherein the drone is detachable from the main body (100b) and has monitoring sensors (110, 111).

10. The monitoring system (103) according to claim 1, It is characterized by The first monitoring sensor (110) includes at least one of the following: camera, microphone, RIM-Camera, laser scanners, LIDAR, radar, Motion detector, and / or Radiometer.

11. A monitoring system (103) for monitoring a facility (102), the monitoring system (103) comprising: A monitoring robot (100) comprising a main body (100b), a drive system (100a), and a motion controller for controlling the motion of the monitoring robot (100). The monitoring system (103) further includes: at least a first monitoring sensor (110) designed to acquire first monitoring data of at least one facility component (105-108, 150-152) of the facility (102), a state detector (100c) configured to detect at least one state (114) associated with the facility component (105-108, 150-152) based on the first monitoring data, It is characterized by The state detector (100c) is configured to: Note the state ambiguity (115) of the state (114), which is between two or more possible states and / or is a result of the object condition being previously unknown to the surveillance robot and / or a result of unfavorable conditions for survey data collection, and Upon noticing the state ambiguity, the action controller triggers an action (117, 120, 130) of the monitoring robot (100), the triggered action (117, 120, 130) being adapted to generate state verification information about the facility component (105-108, 150-152), the state verification information being adapted to verify or falsify the detected state and resolve the state ambiguity, and taking into account the state verification information to resolve the state ambiguity, The triggered actions (117, 120, 130) include interactions of the monitoring robot (100) with the facility components (105-108, 150-152), The interaction with the facility component (105-108, 150-152) includes at least one of the following: tactile contact, and Material is applied to the facility components (105-108, 150-152).

12. The monitoring system (103) according to claim 11, It is characterized by The state detector (100c) is configured to determine the interaction from at least two possible interactions in such a way that the interaction is optimized with respect to the generation of state verification information.

13. The monitoring system (103) according to claim 11, It is characterized by Note that the state blurring (115) of the state (114) is performed by comparing it with a predetermined blur threshold.

14. The monitoring system (103) according to claim 11, It is characterized by The state detector (100c) is further configured to plan actions (117, 120, 130) to be triggered such that the actions (117, 120, 130) are optimized with respect to the generation of verification information.

15. The monitoring system (103) according to claim 11, It is characterized by The tactile contact is made for the purpose of moving the facility component (105-108, 150-152) and / or for the purpose of acquiring tactile sensor data regarding the facility component (105-108, 150-152), and The applied material is a liquid and / or a paint.

16. The monitoring system (103) according to claim 11, It is characterized by The state detector (100c) provides a correlation map relating state ambiguity to interaction position and / or direction, wherein The relevant mapping is based on a defined standard representing the state (114), and / or The correlation map is established by machine learning, and / or The correlation map includes optimal interaction positions and / or directions for a plurality of detectable states (114).

17. The monitoring system (103) according to claim 11, It is characterized by The interaction is based on a state-based Markov model (114).

18. The monitoring system (103) according to claim 11, It is characterized by The triggered action (117, 120, 130) is to obtain data about the facility components (105-108, 150-152), wherein the data is third monitoring data of a third monitoring sensor (104) of the monitoring system (103), the third monitoring sensor (104) not being part of the monitoring robot (100), wherein the triggered action comprises triggering acquisition of the third monitoring data, and / or Stored in the database of the monitoring system (103).

19. The monitoring system (103) according to claim 11, It is characterized by The monitoring robot (100) is an unmanned ground vehicle (UGV).

20. The monitoring system (103) according to claim 19, It is characterized by The unmanned ground vehicle (UGV) includes a drone (UAV) as a subunit, wherein the drone is detachable from the main body (100b) and has monitoring sensors (110, 111).

21. The monitoring system (103) according to claim 11, It is characterized by The first monitoring sensor (110) includes at least one of the following: camera, microphone, RIM-Camera, laser scanners, LIDAR, radar, Motion detector, and / or Radiometer.

22. A monitoring method adapted for use in a facility monitoring system (103), the facility monitoring system (103) comprising at least a first monitoring sensor (110) and a mobile monitoring robot (100), the method comprising the following steps: acquiring first monitoring data of facility components (105-108, 150-152) of a facility (102) using the first monitoring sensor (110), detecting at least one state (114) associated with the facility component (105-108, 150-152) based on the first monitoring data, Note that the state is ambiguous (115) between two or more possible states and / or is a result of an object condition previously unknown to the mobile surveillance robot and / or a result of unfavorable conditions for survey data collection, triggering an action (117, 120, 130) of the mobile surveillance robot (100) upon noticing a state ambiguity, wherein the action (117, 120, 130) is adapted to generate state verification information about the facility component (105-108, 150-152), the state verification information being adapted to verify or falsify the detected state and resolve the state ambiguity, and Resolving the state ambiguity based on the state verification information, The triggered action (117, 120, 130) includes acquiring, by the mobile monitoring robot (100), second monitoring data of the facility components (105-108, 150-152), and The triggered action (117, 120, 130) includes changing an acquisition position and / or direction so that acquisition of the second monitoring data is performed using at least a second acquisition position and / or direction (P20) different from the first acquisition position and / or direction (P10) for acquiring the monitoring data.

23. The monitoring method according to claim 22, It is characterized by The step of blurring the state of the noted state (115) is performed by comparing it with a predetermined blur threshold.

24. A monitoring method adapted for use in a facility monitoring system (103), the facility monitoring system (103) comprising at least a first monitoring sensor (110) and a mobile monitoring robot (100), the method comprising the following steps: acquiring first monitoring data of facility components (105-108, 150-152) of a facility (102) using the first monitoring sensor (110), detecting at least one state (114) associated with the facility component (105-108, 150-152) based on the first monitoring data, Note that the state is ambiguous (115) between two or more possible states and / or is a result of an object condition previously unknown to the mobile surveillance robot and / or a result of unfavorable conditions for survey data collection, triggering an action (117, 120, 130) of the mobile surveillance robot (100) upon noticing a state ambiguity, wherein the action (117, 120, 130) is adapted to generate state verification information about the facility component (105-108, 150-152), the state verification information being adapted to verify or falsify the detected state and resolve the state ambiguity, and Resolving the state ambiguity based on the state verification information, The triggered actions (117, 120, 130) include interactions of the mobile surveillance robot (100) with the facility components (105-108, 150-152), The interaction with the facility component (105-108, 150-152) includes at least one of the following: tactile contact, and Material is applied to the facility components (105-108, 150-152).

25. The monitoring method according to claim 24, It is characterized by The step of detecting at least one state (114) associated with the facility component (105-108, 150-152) based on the monitoring data comprises determining the interaction from at least two possible interactions in such a way that the interaction is optimized with respect to the generation of state verification information.

26. The monitoring method according to claim 24, It is characterized by The step of blurring the state of the noted state (115) is performed by comparing it with a predetermined blur threshold.

27. A computer program product comprising a program code stored on a machine-readable medium and having computer-executable instructions for executing the method according to claim 22 when run on a central computing unit of a monitoring system (103) according to any one of claims 1 to 10.

28. A computer program product comprising a program code stored on a machine-readable medium and having computer-executable instructions for executing the method according to claim 24 when run on a central computing unit of a monitoring system (103) according to any one of claims 11 to 21.

Citation Information

Patent Citations

  • Information processing method, device and system

    EP3156898A1

  • Building-specific anomalous event detection and alerting system

    GB2546486A

  • Fusion technology-based security method and security system thereof

    KR101125233B1

  • System and method for premises monitoring and control using self-learning detection devices

    US20090027196A1

  • Monitoring and security devices comprising multiple sensors

    US20140320312A1