Context information for a surveillance camera
Patent Information
- Application Number
- EP2024703280
- Authority / Receiving Office
- EP · EP
- Patent Type
- Applications
- Current Assignee / Owner
- Priority Date
- 2023-02-07
- Filing Date
- 2024-01-29
- Publication Date
- 2025-12-17
AI Technical Summary
Surveillance cameras face challenges in detecting objects under poor visibility conditions such as reflections, low illumination, and image noise, leading to inaccurate or missed detections due to the limitations of traditional computer vision methods.
The integration of context information, such as background images, semantic segmentation, and calibration data, into the neural network input of surveillance cameras to enhance object detection accuracy by providing additional knowledge about the scene, reducing the reliance on current image quality.
This approach improves the detection rate and reliability of object detection in surveillance cameras by leveraging contextual knowledge, even under adverse conditions, with minimal performance overhead and without requiring complex preprocessing or additional resources.
Smart Images

Figure EP2024052039_15082024_PF_FP
Abstract
Description
[0001] Description
[0002] title
[0003] Context information for a surveillance camera
[0004] The invention relates to the detection of objects in a surveillance area using a camera.
[0005] State of the art
[0006] DE 10 2019 207 711 A1 discloses a machine learning system configured to detect smoke within a plurality of consecutively acquired images. The machine learning system comprises a convolutional recurrent neural network. Furthermore, a method for detecting smoke using this machine learning system is known.
[0007] Disclosure of the invention
[0008] The object of the invention is to propose improvements with regard to the detection of objects in a surveillance area by means of a camera.
[0009] The object is achieved by a camera arrangement according to patent claim 1. Preferred or advantageous embodiments of the invention and other categories of invention emerge from the further claims, the following description and the attached figures.
[0010] The camera arrangement serves to generate parameters. The parameters, in turn, serve to detect objects in a surveillance area. The camera arrangement contains a camera. This camera is intended to be mounted or is mounted in a known mounting position relative to the surveillance area. "Intended" means that the invention is based on a camera mounted in this way. It is therefore assumed that the camera is or is aimed at an environmental area in order to be able to create images of this area and is configured for such use there. In other words, a relevant environmental area and an orientation of the camera towards this area are assumed to be given – at least during camera operation.
[0011] The camera is configured to generate or be able to generate images of the surveillance area in this mounting position. In other words, in the mounting position, i.e., in a mounted state, the camera is aligned with the surveillance area so that it can capture images of it.
[0012] The camera is fixed in its mounting position, making it a so-called static camera. "Fixed" means that it is installed at a fixed location relative to the surveillance area. The camera can be fixedly aligned, i.e., it can be directed in a fixed, fixed direction of view. However, a fixed installation can also be understood here as a swiveling and / or zoomable camera (so-called PTZ camera, pan / tilt / zoom). This fixed installation, however, must be distinguished from, for example, a mobile camera mounted on a vehicle; such a camera will not be considered here. The background to this is that only with a fixed-mounted camera can the surrounding area be assumed to be "known" during its operation, in the sense that definitively relevant contextual information about it can be determined, e.g., a known recording / image of the surrounding area under known good visibility and lighting conditions.
[0013] The surveillance area could be a landscape, a company premises, or the interior of a building. Specifically, for example, an intersection, which includes both a roadway and its surroundings (houses, lawns, traffic islands, traffic lights, etc.). All of this is then depicted in the images.
[0014] Monitoring the area involves discovering or detecting any objects in the surveillance area that are not originally part of the surveillance area. These objects are therefore located in or on the surveillance area. For example, if the surveillance area is the intersection mentioned above, objects include people or vehicles that may be located in or moving through the surveillance area.
[0015] The camera array contains a neural network. The network has an input for the images generated by the camera and an output. The network is essentially configured to process the images fed into the input (one or more) into at least one characteristic. This characteristic is correlated with the desired object detection. The network is also configured to provide the determined characteristics at the output.
[0016] The characteristic is, for example, a yes / no value as to whether an object is in the surveillance area, or it is the position coordinates of a potentially detected object in the camera image, or a classification of what the object could be (for example, "car", "truck", "pedestrian", "bicycle", a vehicle's license plate number, etc.).
[0017] The camera arrangement contains a context module. This module is configured to provide at least one piece of context information correlated with the surveillance area. The "context information" is explained in more detail below.
[0018] The camera arrangement is configured to feed not only the images but also the context information into the network input. The network is configured to process the context information fed into the input, together with the images fed into the input, into at least one parameter correlated with the detection of objects and to provide the parameters at the output. The camera arrangement thus provides the parameters and thus allows the parameter to be further used for the actual detection of objects in the surveillance area. If necessary, the actual detection itself can also take place in the camera arrangement. For this purpose, the camera arrangement then contains, in particular, an evaluation unit. This unit is then configured to detect any objects in the surveillance area based on the parameters.The camera arrangement is then designed to detect objects in the surveillance area and thus functionally expands the actual detection capability compared to the first-mentioned camera arrangement (generating parameters). Preferably, the evaluation unit is configured to trigger an alarm upon detection of an object, for example, by activating a siren and / or a flashing light. In one variant, the evaluation unit is configured to transmit the detected object and / or the resulting alarm to a control center via wired and / or wireless signals.
[0019] Such a camera arrangement is occasionally referred to simply as a "camera", for example if it is designed as a structural unit containing the actual camera and the neural network as well as the context module and, if applicable, the evaluation unit.
[0020] According to the invention, context information is generated and incorporated for or into a surveillance camera or into a corresponding camera arrangement.
[0021] The invention is based on the realization that deep learning and machine learning methods are becoming increasingly used in surveillance cameras and connected cloud services. Similar to classic methods in computer vision, the performance of these methods suffers from poor quality of the images generated by the camera, which can be caused, for example, by poor lighting, weather-related poor visibility conditions, or disruptive effects such as backlighting. To counteract this, it is proposed here to provide the neural network with additional information in the form of context information, e.g., additional channels, during inference (i.e., at runtime). For example, in addition to a current image, a background image taken under good visibility conditions could be provided as context information.Calibration information, such as distances (from components of the surrounding area to the camera) and ground plane normals (of the surrounding area), could also be provided as context information in the form of a 2D or 3D array. Finally, results (output variables) from another, perhaps larger, neural network, which, for example, received the background image (or entire image sequences) as input, could also be used as context information.
[0022] The main advantage is that the neural network receives background information that it does not have to derive from the current image (the context information therefore represents contextual knowledge of the current camera image). This background information can be determined in advance and stored on the camera, so that only a low overhead arises during runtime, e.g., through the provision of additional input channels (part of the input for feeding in the context information). In a variation, the background information (context information, sensor information from additional sensors, etc.) can also be preprocessed by a neural network, for example, reducing the number of channels (condensing it, e.g., into a context value).
[0023] In a preferred embodiment, the input contains at least one, in particular several, channel inputs. Alternatively, the input consists exclusively of one, in particular several, channel inputs. The camera arrangement is then configured to provide at least one of the images and / or at least one piece of context information to the neural network (the first and / or second, see below) in the form of at least one channel at the channel inputs. In particular, exactly one channel is provided to exactly one channel input. The channel inputs for the context information can, in particular, be additional channel inputs to those via which the images are fed into the network. "Channel inputs" and "channels" are to be understood as follows: For example, a color image generated by the camera is an RGB image (red-green-blue). The image is provided in the form of three channels orthree partial images, namely a red image, a green image, and a blue image. Each of these partial images has identical dimensions (same number and arrangement of pixels). The pixels in the red image are the red values of a respective pixel; those in the green image are the green values; and those in the blue image are the blue values of a respective color triplet per color pixel. The context information is then also provided, for example, as an additional (fourth) channel, i.e., as a data structure with the same "pixel dimensions." In other words, for each image pixel, in addition to the three values for the red, green, and blue components, a fourth value is provided as context information. This type of channel processing is particularly advantageous in a neural network.
[0024] In a preferred embodiment, at least part of the context information is stored in the camera arrangement, at least during operation. For this purpose, the camera arrangement, in particular, has a memory. In particular, the context information is already determined before the image is captured by the camera and the corresponding current images are processed. Therefore, no effort is required in the camera arrangement to generate or determine the context information, i.e., no resources need to be reserved; a simple memory, for example, is sufficient. In the camera arrangement, there is then only a small performance overhead compared to one without the processing of context information. For example, only a "larger" input needs to be provided, for example, additional channels for the context information, in order to be able to feed this into the neural network alongside the images.
[0025] In a preferred embodiment, at least one piece of context information is a background image of the surveillance area. This corresponds in particular to the image of the (pure) surveillance area, i.e., without any objects, in particular a camera image taken under "good" visibility conditions, for example, in daylight, without stray light sources, without rain / snow, and with a dry and clean camera lens. This provides the neural network with, for example, an "optimal / good target image" of the surveillance area.
[0026] Alternatively or additionally, at least one piece of context information is a semantic segmentation of the surveillance area. This provides, for example, information about which image areas correspond to a street, a building, a tree, a meadow, etc. This can also improve the performance of the neural network.
[0027] In a preferred embodiment, at least one piece of context information is calibration information for the camera with respect to the surveillance area. Such calibration information is, for example, a so-called depth image, which - in the mounting position of the camera - provides, for example, for each pixel a distance between the camera and a point in the surveillance area depicted in the pixel. However, corresponding calibration information can also be, for example, a respective ground plane normal of the surveillance area. Here, for example, the normal direction of the point in the surveillance area depicted in the pixel is specified for each pixel. This can also significantly increase the performance of the neural network.
[0028] In a preferred embodiment, the neural network described so far is a first neural network. The camera arrangement then contains a further, second neural network. This is configured to process at least one piece of context information into at least one context value. The camera arrangement is then further configured to feed at least one piece of context information into the second network and (after processing) to feed the context value (and thus the context information preprocessed into the context value) into the input of the first network. In other words, the context information is preprocessed by the second network into a context value and then fed (compressed / condensed) as a context value into the first network. The second network generally also has an input and output, which will not be explained in detail here.For example, several additional context channels (channels with respective context information) can be reduced to a single channel (context value) through preprocessing in the second neural network. The first neural network then only needs to process the single context channel (context value). This allows, for example, the processing load in the first neural network to be reduced while maintaining the same amount of processed context information.
[0029] In a preferred embodiment, the camera arrangement is configured to store at least one piece of registration information, optionally also to generate it beforehand. The camera arrangement is then also configured to consider the registration information with regard to the processing in the camera arrangement, in particular in the (first or possibly also second) neural network and all other steps of processing context information and / or images. The registration information represents a spatial / local relationship between the images recorded or to be recorded by the camera and the context information. The registration information thus describes the local or spatial relative relationships between images and context information.Here, for example, movements of the camera, whether intentional or unintentional (maladjustment due to aging / wind pressure or readjustment during service work or regular movement / image changes during operation, e.g. due to PTZ), can be taken into account in order to be able to assign the context information and the images to the correct location.
[0030] In a preferred embodiment, at least one of the pieces of context information is one that is generated by processing at least one image taken by the camera in the mounting position. In particular with regard to the above (registration information), it is thus ensured that the context information and the image, and thus also the camera, are fully registered with one another, i.e. that they are spatially correct. Here, it only needs to be ensured that the camera has not changed in its relative position to the surveillance area when later images are taken. A known change in the actual or current relative position (PTZ camera) can then be assumed to be known and taken into account. In a preferred embodiment, at least one of the pieces of context information is based on at least one piece of sensor information from at least one sensor. The sensor is different from the camera and, in particular, does not belong to the camera arrangement.The sensor then represents an external sensor and provides at least a portion of context information for the camera arrangement. The camera arrangement then has, in particular, an interface configured to receive the sensor information and / or the corresponding context information, particularly from outside the camera arrangement. For example, expensive and complex sensor technology / preprocessing can be used to determine "external" context information in order to set up the camera arrangement or camera once. The camera arrangement itself can thus be kept simple and inexpensive. This is particularly useful in combination with the storage of pre-generated context information in the camera arrangement.
[0031] The object of the invention is also achieved by a method according to claim 10. This preferably serves to operate the camera arrangement according to the invention as described above.
[0032] The invention thus relates to a computer-implemented method for generating parameters for detecting objects in a surveillance area, wherein images of the surveillance area generated by a camera are provided, wherein at least one piece of context information correlated with the surveillance area is provided, wherein the neural network processes the context information fed into an input of the neural network together with the images fed into the input to form at least one parameter correlated with the detection of objects and provides the parameter at the output of the neural network.
[0033] Preferably, the neural network is a neural network trained with images and context information correlated with the images.
[0034] In this method, the camera is mounted in its intended mounting position relative to the surveillance area. During operation, the camera generates images of the surveillance area in its mounted position. The images are fed into the input of the neural network. The context module provides the context information correlated with the surveillance area. In addition to the images, the context information is also fed into the input of the neural network. The neural network processes the context information together with the images to produce the characteristic value(s) and provides the characteristic value(s) at the output.
[0035] The method and at least some of its possible embodiments as well as the respective advantages have already been explained in connection with the camera arrangement according to the invention.
[0036] The invention further relates to a computer program configured to execute all steps of the described method. The computer program preferably runs on a control unit or a computer with a microprocessor and a memory. Furthermore, the invention relates to a machine-readable storage medium, in particular a non-volatile machine-readable storage medium, on which the computer program is stored.
[0037] The invention is based on the following findings, observations, and considerations and also includes the following preferred embodiments. These embodiments are sometimes referred to as "the invention" for simplicity. The embodiments may also contain parts or combinations of the above-mentioned embodiments or correspond to them and / or may also include previously unmentioned embodiments.
[0038] The invention aims to improve the performance, e.g. detection rate, of a neural network that processes data from a static surveillance camera by incorporating context information and generating and incorporating it.
[0039] The invention is based on the following basic idea: Poor visibility conditions in the broadest sense are a major problem for surveillance cameras (and also moving cameras). These include, for example, reflections on wet roads, stray light from water on the lens or fog, low / poor illumination of the scene and the resulting image noise, compression artifacts due to excessive compression of the image material, obscuration of fields of view by interfering objects on the lens or, for example, large vehicles, blurring of the camera image, etc.
[0040] These impairments typically result in objects, such as people or vehicles, not being detected or only being detected with very low reliability. If these objects are to be detected, the threshold could be set low to ensure minimum detector reliability, but this often leads to false detections.
[0041] In all of these cases (and even in cases where no impairment is present), it is therefore advantageous to add additional information or contextual knowledge about the scene (surveillance area). One simple option is to add a depth image. For example, another channel is added to the three channels of the RGB image (and fed into the neural network). It is important that the depth image has the same resolution as the original image. Alternatively, the image (context information) can be (artificially) enlarged or fed into the neural network at a later layer (i.e., via an input that leads to one of the inner layers of the network).
[0042] In general, it is difficult to introduce additional information into a neural network that is not present in a form that has the same or similar resolution as the image (RGB image). One example of this is information about an extrinsic calibration, which consists of, for example, the tilt and roll angle and a camera height. Although these are only three parameters, it can be useful to create an entire tensor (with the same size as the original image) in which this information is stored. This also has the advantage that this information is close to the corresponding image information during processing (locality). In the case of a convolutional neural network (CNN), the input channels are convolved by a kernel that simultaneously considers all channels (i.e. RGB = image and additional information = context information) but only restricts its analysis to a small section of the image.By placing the input channels one after the other, it is ensured that the information relevant for the local image section is present within the kernel.
[0043] The advantage of a fixed surveillance camera over a moving camera, e.g., one mounted on a vehicle, is that the camera's orientation typically does not change (see below). This means that the background information (context information, especially the additional channels) only needs to be generated once.
[0044] For a neural network, this additional information must be available for both training and inference (use on the camera or in the cloud or edge box (microcomputer / controller) with current images).
[0045] A detailed description of possible input channels and obvious alternatives follows.
[0046] A possible additional piece of information (context information) is a background image. This can be, for example, a single daytime image, a temporally filtered image series where moving objects are removed (e.g., by median calculation), or the result of processing a series of images by a neural network.
[0047] Further additional information (context information) could be semantic segmentation, generated, for example, with the help of another neural network. The background image, for example, could serve as input for this. Alternatively, semantic segmentation could also be generated directly by a person (annotation / labeling). Semantic segmentation on the background image can be generated, for example, by "Rene Ranftl, Alexey Bochkovskiy, Vladlen Koltun: Vision Transformers for Dense Prediction, https: / / github.com / isl-org / DPT". Similar to semantic segmentation, a depth image could also be stored as context information. This could, for example, be determined by a neural network (e.g., the depth image can be generated by "Rene Ranftl" (see above), applied to the background image) or, for example, by an additional sensor and registration (see below) during camera installation.
[0048] An extrinsic calibration as context information can, for example, be information about a ground plane. For example, the parameters of a plane normal and distance (i.e., three parameters) could be stored in an array of the size image width x image height x 3. This has the advantage, for example, that it is immediately apparent how an upright person (as an object) should be oriented in the image and how large they should appear in the image.
[0049] Furthermore, registration is a topic of discussion here: Two types of registration are discussed below. First, the registration of background information (additional information / context information) to the current image, and second, the registration of information from an additional sensor (e.g., a depth camera).
[0050] In the simplest case, all images required to generate the additional information are obtained directly from the camera. This has the advantage that no registration of any kind is necessary (except for calibration information). However, if the camera moves over time, e.g., due to wind or deliberate panning of the camera, the background information and the current image no longer directly match. This may be irrelevant, since the background information usually does not change abruptly, and these effects could be simulated (augmented) during training to achieve a corresponding non-susceptibility.
[0051] Alternatively, however, registration can also be performed. For example, corresponding pixels would be calculated from the background image and the current image, and a compensating image transformation would then be performed. This can also be performed, for example, taking into account a known intrinsic value of the camera. Corresponding points in the background information where no information is available after the transformation would have to be marked accordingly (as ignore) or extrapolated. Both can also be simulated / augmented during training to achieve a corresponding non-vulnerability.
[0052] Alternatively, the information from an IMU (Inertial Measurement Unit) can be used.
[0053] If information from an additional sensor is incorporated—for example, if the installer has a depth camera with them—this information must also be registered and transformed against the current image. The procedure can be essentially the same as described above.
[0054] In general, however, it is advantageous if the background information is obtained from the same perspective (e.g. directly next to the camera), so that in the simplest case only a virtual rotation of the data needs to be carried out.
[0055] Regarding the generation of additional channels / background information, it should be said:
[0056] So far, it has been proposed to generate the background information once. However, it can generally also be updated after a certain period of time. This can be done periodically, for example, or after a major change in the scene has been detected. The additional information can be (I) generated on the camera itself. The relevant information would be collected (e.g., images for generating the background image or autocalibration of the scene) and generated by a neural network or other algorithm. This process does not have to occur in real time.
[0057] This has the advantage that, for example, a significantly larger and more complex neural network can be used than is later the case in operation.
[0058] Alternatively, the background information could be generated (II) on the installer's laptop / PC or (III) in the cloud or edge box and then copied to the camera. Processing in the cloud would have the particular advantage that the background information could also be updated if an improved version of its generation becomes available.
[0059] The additional information / background information (synonymous with context information) can be introduced in various ways. The relevant channels can either be simply appended to the input image or to a later layer. Alternatively, preprocessing can be performed using another (second) neural network. In this case, many additional channels are reduced to a few. For this purpose, another (second) neural network is used, which can, for example, be trained together with the main network. This condensation does not need to be performed for all additional channels.
[0060] In addition to the main information, e.g. depth, a corresponding uncertainty can also be stored in the additional channels.
[0061] Possible alternatives / variants of the invention are: A neural network can be, for example, a CNN or transformer-based method for machine learning. The method / arrangement can be used for individual images or image series. It is irrelevant whether the image is RGB, YUV, or even just a grayscale image. The method / arrangement could also be extended to moving cameras, where, for example, the calibration information could be of particular interest. For a PTZ camera, it may be useful to create a background image by stitching several individual images and to provide the corresponding section as context information during operation.
[0062] Many advantages of the method / arrangement have already been described. The main advantage of the variants described here is that valuable background information can be generated in the form of context information and made available to the neural network without significantly increasing the inference effort. In the simplest case, only the kernel depth in the input layer changes, which should have only a minor impact on the overall runtime.
[0063] The main idea of the invention can be illustrated as follows: Traditionally / in practice, the input of a neural network on a surveillance camera consists of a single RGB image (or a series of images or one or more grayscale images), which thus corresponds to three channels. The performance of these networks typically decreases at night or in the rain, for example, people or vehicles are not recognized or are therefore recognized with significantly less certainty. To counteract this, it is proposed here to add context knowledge in the form of additional channels. This context knowledge can, for example, be the result of semantic segmentation that was previously determined on a good weather image using a different neural network that does not have to run on the camera. Likewise, the good weather image itself, or a dedicated background image, the result of a depth estimation or measurement and calibration information can be added.
[0064] To reduce the number / overhead of the additional channels (condense), another neural network could be used to first condense the background information, for example, reducing the number of channels. This can be done for all or just parts of the background information. The additional neural network can be trained together with the original network.
[0065] Further features, effects, and advantages of the invention will become apparent from the following description of a preferred embodiment of the invention and the accompanying figures. Each of these figures shows a schematic diagram:
[0066] Figure 1 shows a surveillance area monitored by a camera arrangement and context information for the camera arrangement,
[0067] Figure 2 shows a camera image of the camera arrangement in poor visibility conditions with objects to be detected,
[0068] Figure 3 further context information for the camera arrangement, Figure 4 a camera arrangement with a second neural network,
[0069] Figure 5 shows the generation of context information from current camera images.
[0070] Figure 1 shows a surveillance area 2, which is only indicated schematically here, in this case a section of a landscape, namely a road intersection 4. This is surrounded by four lawns 6 as well as two houses 8, a tree 10 and two traffic lights 12. The surveillance area 2 is to be monitored or is monitored by a camera arrangement 14, which is also only shown schematically in Figure 1. For this purpose, the camera arrangement 14 is intended to generate parameters 16, which in turn are used to detect objects 18 in the surveillance area 2. All objects shown in Figure 1 in the surveillance area 2 (street, houses, trees, ...) are not objects 18 in this sense, but merely fixed components of the surveillance area 2. Objects 18 are, for example, people or animals or vehicles that move in or through the surveillance area 2. These can be seen in Figure 2 (see below).
[0071] The camera assembly 14 contains a camera 20. This is located in a designated mounting position M in a fixed relative position R to the surveillance area 2, indicated by a double arrow. In the example, the camera 20 is mounted on the roof 24 of the house 8, of which only a portion of the roof 24 is visible from above in Figure 1. The camera 20 views the surveillance area 2 from the house 8 in a fixed orientation, i.e., with an unchanging viewing direction / angle.
[0072] In the mounting position M and during operation, the camera 20 continuously generates images 22 of the surveillance area 2, one after the other, here, for example, one image 22 per second. Due to the mounting of the camera 20 on the house 8, only its roof 24 is visible in the foreground of the recorded images 22. One of the images 22 is indicated by a dashed frame in the surveillance area 2.
[0073] The camera arrangement 14 contains a neural network 26, which has an input 28 for the images 22 and an output 30 for outputting the characteristic variable 16. The network 26 is fundamentally configured to process the images 22 fed into the input 28 into the characteristic variable 16 and to provide the characteristic variable 16 at the output 30.
[0074] The camera arrangement 14 also contains a context module 32. This is configured to provide at least one piece of context information 34 correlated with the surveillance area 2 (reference symbol "34" generally also represents one or more of "34a-g"). The camera arrangement 14 is configured to feed the context information 34 into the input 28 in addition to the images 22. The network 26 is thus ultimately and specifically configured to process not only the images 22 fed into the input, but also these together with the context information 34 fed into the input 28, into the characteristic variable 16. The characteristic variable 16 is correlated with the detection of objects 18 and is provided at the output 30. In the example, it indicates the number and location in the image 22 of detected objects 18.
[0075] In the example, input 28 contains a total of five channel inputs 36a-e. Camera arrangement 14 is configured to provide the images, or each of the images 22, to the network 26 in the form of three channels 38a-c each. It is also configured, in the example, to provide two pieces of context information 34a, b as additional channels 38d, e at channel inputs 36d, e. Channels 38a-c are the red, green, and blue channels of a respective image 22 and are each fed individually into the respective channel inputs 36a-c. Each of the channels 38a-e has the format of a (1024 x 768) pixel image with 768 rows of 1024 pixels each (images 22) or values (context information). The context information 34 is also provided in such an image format, although it does not necessarily have to be / represent an actual "image."
[0076] In Figure 1, two pieces of context information 34a, b are shown in simplified form, representing a large number of these.
[0077] The context information 34a, b is stored permanently in the camera arrangement 14 from the time the camera arrangement 14 is installed in the surveillance area 2, here in a memory 40 of the camera arrangement 14, and was thus created before the camera 20 was put into operation. The image 22 of the surveillance area 2 shown in Figure 1 and indicated by the dashed frame was taken without any objects 18 and under optimal visibility conditions (daylight, no precipitation, clean camera lens, etc.). This image thus forms a background image 42 of the surveillance area 2 and thus represents exemplary context information 34.
[0078] Figure 1 symbolically shows a piece of registration information 62, which contains information about how the images 22 currently generated by the camera 20 and the context information 34a, b are spatially related to one another. The camera arrangement 14 is configured to also take this piece of registration information 62 into account when determining the characteristic variable 16, ie, to assign the relevant context information 34a, b to a respective current image 22 in a spatially correct manner or to take it into account in its processing.
[0079] Figure 1 also shows—indicated purely symbolically here—a sensor 66, different from camera 20, which determines sensor information 68 from the surveillance area 2. This is, in particular, depth information, namely distances between points in the surveillance area 2 and the camera 20 (see below). Alternatively, the sensor 66 determines orientation information as sensor information 68, i.e., a respective normal vector at respective points in the surveillance area 2 (see below). Context information 34 is then obtained from the sensor information 68 or based on it and provided to the camera arrangement 14 for use, in particular stored in its memory 40.
[0080] Figure 2 shows another image 22 taken during operation of the camera arrangement 14 or the camera 20. This image 22 was taken in winter. The roof 24 is therefore no longer clearly visible because it is covered by snow 44. There are two water spots 46 on the lens of the camera 20, which blur or impair the camera image 22 at the corresponding locations. In the surveillance area 2, there are two fog patches 48, indicated by horizontal dashed lines, which also impair the view in the camera image 22. Two objects 18 in the surveillance area 2, here a pedestrian and a vehicle, are therefore difficult to recognize in image 22.
[0081] Figure 2 shows further context information 34b in the form of a semantic segmentation 50, which is indicated in Figure 2 as follows: The roof 24, which is no longer recognizable in the actual image 22, is segmented as a hatched area, and the lawns 6 are also indicated by other hatching. The context information 34b is also fed to the network 26.
[0082] When generating the characteristic variable 16, the neural network 26 is thus informed about known contents in the camera image 22 with helpful information in the form of the context information 34a, b and can thus take over the recognition of the objects 18 or the generation of the characteristic variables 16 in a simplified and improved manner.
[0083] Figure 3 shows further examples of context information 34, here in the form of calibration information 52. In a first example, this information is provided as a depth image, here clarified or illustrated in the form of depth lines 54. Each depth line 54 designates points in the surveillance area 2 depicted at the relevant "image location" of the context information, which are at a specific distance from the camera 20. In Figure 3, three depth lines 54 are indicated by dashed lines: depth lines 54a for distances of 5 m, 54b for 10 m, and 54c for 15 m. In fact, the calibration information 52 is a depth image of the surrounding area 2, in which each "pixel" has the corresponding distance value of the surveillance area 2 from the camera 20.
[0084] As a further alternative for context information 34 in the form of calibration information 52, a normal image is indicated in Fig. 3. This image has values at each "pixel" that describe a ground plane normal 56 (normal vector of the surface) at the corresponding location in the surveillance area 2. The normal image thus describes at each "pixel" which respective vectorial orientation the surface (its normal) has at the corresponding location in the surveillance area 2. For example, it can be seen that in Fig. 3, the two upper lawns 6 each slope downwards towards the intersection 4 as an embankment (normals inclined obliquely to the street), whereas the two lower lawns 16 shown in the image merge flatly (perpendicular normals) into the intersection 4. The inclination / orientation of the roof 24 is also visible. Here, too, the normals are only indicated as examples in a few places.The actual normal image contains the orientation of the monitoring area 2 at each "pixel".
[0085] The context information 34 in the form of the calibration information 52 is determined in particular based on the sensor 66 or its sensor information 68.
[0086] Figure 4 symbolically shows a section of an alternative camera arrangement 14. Here, the previously designated neural network 26 is a first neural network. The feeding of the images 22 into the input 28 of this first neural network 26 is not shown in Figure 4 for the sake of clarity. Only the feeding of the context information 34 is shown. A total of four pieces of context information 34d-g are fed into the neural network 26, each into a respective channel 38d-g. However, preprocessing takes place here before the actual feeding into the neural network 26. For this purpose, the camera arrangement 14 contains a second neural network 58. The four pieces of context information 34 are fed into this second neural network and compressed or condensed into a single piece of context information 34 in the form of a context value 60.Only this is then fed into the first neural network 26 as a single channel 38h and, as usual, processed together with the images 22 to form the characteristic 16. The context information 34 is thus fed into the neural network 26 as a preprocessed context value 60.
[0087] Figure 5 shows one possibility for generating context information 34, namely by processing images 22 actually captured by the camera 20 itself (only one is shown as an example) using a processing unit 64. The latter may be part of the camera arrangement 14, but does not have to be. Figure 5 once again illustrates the structure of an image 20 from the three channels 38a (red), 38b (green), and 38c (blue), as well as their input into the neural network 26.
[0088] In particular, the processing unit 64 is used to obtain, for example, the background image 42 and the semantic segmentation 50 from the images 22. In summary, when the camera arrangement 14 is operated, the camera 20 is first fixedly mounted in the known mounting position M relative to, i.e., in the relative position R, the surveillance area 2. The camera 20 then generates images 22 of the surveillance area 2. The images 22 are fed into the input 28 of the neural network 26. In addition, the context information 34 is also fed into the input 28. The neural network 26 processes all of this into the characteristic variables 16 and makes them available at the output 30.
Claims
Claims 1. Camera arrangement (14) for generating parameters (16) for detecting objects (18) in a surveillance area (2), - with a camera (20) which is intended to be mounted or mounted fixedly relative to the surveillance area (2) in a known mounting position (M), which camera is designed to generate images (22) of the surveillance area (2) in the mounting position (M), - and with a neural network (26) having an input (28) and an output (30), - and with a context module (32) which is designed to provide at least one piece of context information (34a-g) correlated with the monitoring area (2), - wherein the camera arrangement (14) is configured to feed the context information (34a-g) into the input (28) in addition to the images (22), - and the neural network (26) is designed to combine the context information (34a-g) fed into the input (28) together with the images (22) fed into the input (28) into at least one function associated with the detection of objects (18) correlated characteristic (16) and to provide the characteristic (16) at the output (30).
2. Camera arrangement (14) according to claim 1, characterized in that - the input (28) contains at least one channel input (36a-g), - and the camera arrangement (14) is configured to provide at least one of the images (22) and / or at least one of the context information (34d-g) to the network (26) in the form of at least one channel (38a-g) at the channel inputs (36a-g).
3. Camera arrangement (14) according to one of the preceding claims, characterized in that at least part of the context information (34a-g) is stored in the camera arrangement (14) at least during operation of the camera arrangement (14).
4. Camera arrangement (14) according to one of the preceding claims, characterized in that at least one of the context information (34a-g) is a background image (42) of the surveillance area (2) and / or a semantic segmentation (50) of the surveillance area (2).
5. Camera arrangement (14) according to one of the preceding claims, characterized in that at least one of the context information (34a-g) is calibration information (52) of the camera (20) with respect to the surveillance area (2).
6. Camera arrangement (14) according to one of the preceding claims, characterized in that the neural network (26) is a first neural network (26) and the camera arrangement (14) contains a second neural network (58) which is set up to process at least one of the pieces of context information (34a-g) into at least one context value (60), wherein the camera arrangement (14) is set up to feed at least one of the pieces of context information (34a-g) into the second network (58) and, after processing thereof, to feed it into the input (28) of the first network (26) as a context value (60).
7. Camera arrangement (14) according to one of the preceding claims, characterized in that the camera arrangement (14) is designed to - to store and optionally also generate at least one registration information (62) between the images (22) and the context information (34a-g), - and to take into account the registration information (62) with regard to the processing in the camera arrangement (14).
8. Camera arrangement (14) according to one of the preceding claims, characterized in that at least one of the context information (34a-g) is one which is determined by Processing of at least one image (22) taken by the camera (20) in the mounting position (M) 9. Camera arrangement (14) according to one of the preceding claims, characterized in that at least one of the context information (34a-g) is based on at least one sensor information (68) of at least one sensor (66) different from the camera (20).
10. Computer-implemented method for generating parameters (16) for detecting objects (18) in a surveillance area (2), wherein images (22) of the surveillance area (2) generated by a camera (20) are provided, wherein at least one piece of context information (34a-g) correlated with the surveillance area (2) is provided, wherein a neural network (26) processes the context information (34a-g) fed into an input (28) of the neural network (26) together with the images (22) fed into the input (28) to form at least one parameter (16) correlated with the detection of objects (18) and provides the parameter (16) at the output (30) of the neural network (26).
11. The method according to claim 10, characterized in that the neural network (26) is a neural network (26) trained with images (22) and context information (34a-g) correlated with the images (22).
12. Method according to one of claims 10 or 11 for operating a camera arrangement (14) according to one of claims 1 to 9.
13. The method according to claim 12, wherein: - the camera (20) is mounted as intended in the known mounting position (M) firmly relative to the surveillance area (2), - the camera (20) in the mounting position (M) generates images (22) of the surveillance area (2), - the images (22) are fed into the input (28) of the neural network (26), - the context module (32) provides the context information (34a-g) correlated with the monitoring area (2), - in addition to the images (22) also the context information (34a-g) in the input (28) are fed in, - and the neural network (26) processes the context information (34a-g) together with the images (22) to form the parameters (16) and provides the parameters (16) at the output (30).
14. A computer program configured to perform all steps of the method according to any one of claims 10 to 13.
15. Machine-readable storage medium, in particular non-volatile machine-readable storage medium, on which the computer program according to Claim 14 is stored.