Method and image processing device for determining probability value indicating that object captured in stream of image frames belongs to object type

JP2024046748A5Active Publication Date: 2025-05-22AXIS
View PDF 0 Cites 0 Cited by

Patent Information

Application Number
JP2023151778
Authority / Receiving Office
JP · JP
Patent Type
Applications
Current Assignee / Owner
Priority Date
2022-09-23
Filing Date
2023-09-19
Publication Date
2025-05-22
Estimated Expiration
2043-09-19

AI Technical Summary

Technical Problem

Existing surveillance systems struggle to accurately detect and mask partially occluded objects, particularly humans, leading to lower detection scores and potential identification of individuals due to occlusion by other objects.

Method used

A method and device that utilize a multi-camera system to enhance object detection by sharing detection information across cameras, adjusting detection scores based on spatial overlap and probability values from multiple angles, and applying privacy masks to ensure accurate masking and counting of occluded objects.

Benefits of technology

Improves the detection and masking of partially occluded objects by increasing the probability values, ensuring accurate identification and anonymization of individuals in surveillance footage, thereby enhancing privacy and data utility.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure 00000000_0000_ABST
    Figure 00000000_0000_ABST
Patent Text Reader

Abstract

To provide a method and an image processing device for determining a probability value indicating that an object captured in a stream of image frames belongs to an object type.SOLUTION: A method comprises: detecting a first object or a first part of the first object in a first area of an image frame captured by a first camera of a multi-camera system; determining a first probability value indicating a first probability that the detected first object belongs to an object type; detecting a second object or a second part of the second object in a second area captured in a second stream of image frames by a second camera; determining a second probability value indicating a second probability that the detected second object belongs to the object type; and when the second probability value is below a second threshold value and the first probability value is above a first threshold value, determining an updated second probability value by increasing the second probability value.SELECTED DRAWING: Figure 5
Need to check novelty before this filing date? Find Prior Art

Description

[Technical field]

[0001] The embodiments herein relate to a method and an image processing device for determining a probability value indicating that an object captured in a stream of image frames belongs to an object type. A corresponding computing program and a computing program carrier are also disclosed. [Background technology]

[0002] Public surveillance using imaging, especially video imaging, is common in many areas around the world. Areas that may require surveillance are, for example, banks, stores, and other areas that require security, such as schools and government facilities. However, it is illegal to install cameras in many locations without a license / permit. Other areas requiring surveillance are processing, manufacturing, and logistics applications, and video surveillance is primarily used to monitor processing.

[0003] However, there may be a requirement that people cannot be identified from video surveillance. The requirement that people cannot be identified may be contrasted with the requirement that what is happening in the video can be determined. For example, it may be important to perform people counting or queue monitoring on anonymous image data. In practice, there is a trade-off between meeting these two requirements, i.e., unidentifiable video, and extracting large amounts of data for different purposes, such as people counting.

[0004] Several image processing techniques have been described to avoid identifying people while being able to recognize activities. Edge detection / representation, edge enhancement, silhouetted objects, and different kinds of "color blurring" such as color shifting or dilation are examples of such operations. Privacy masking is another image processing technique used in video surveillance to protect the privacy of individuals by hiding parts of an image from view with masked regions.

[0005] Image processing refers to any processing applied to an image. Processing can include the application of various effects, masks, filters, etc. to the image. In this manner, the image can be, for example, sharpened, converted to grayscale, or altered in some way. Images are typically captured by video cameras, still image cameras, etc.

[0006] As mentioned above, one way to avoid identifying people is by masking moving people and objects in the image in real time. Masking in live and recorded video can be done by comparing the live camera view with a set background scene and applying dynamic masking to areas of people and objects that change and are essentially moving. Color masking, which may also be called solid color masking or monochrome masking, where objects are masked by an overlaid solid mask of a specific color, provides privacy protection while allowing the movement to be seen. Mosaic masking, also called pixelation, pixelated privacy masking, or transparent pixelated masking, shows moving objects in a lower resolution, allowing the form to be better distinguished by seeing the color of the object.

[0007] Live and recorded video masking is suitable for remote video monitoring or recording in areas where surveillance is an issue due to privacy rules and regulations. It is ideal for processing, manufacturing, and logistics applications where video surveillance is primarily used to monitor processes. Other potential applications are retail, education, and government facilities.

[0008] Before masking an object, the object may need to be detected as an object to be masked, or in other words, classified as an object to be masked. Patent document 1 discloses a method including obtaining a video of a scene, detecting objects in the scene, determining an object detection probability value indicating the likelihood that the detected object belongs to an object category to be redacted, and redacting the video by obfuscating the detected object that belongs to the object category to be redacted.

[0009] A challenging problem found when dynamic masking is used in surveillance systems is that the detection of a partially occluded person gives a much lower detection score, such as a lower object detection probability value, compared to the detection score obtained for the detection of an unoccluded person. The occlusion may be due to another object in the scene. The lower detection score may result in the partially occluded person not being masked in the captured video stream. For example, if an object (e.g., a person) captured in the video stream is half occluded, the detection score obtained from the object detector may be 67% that the object is a human, whereas if the captured object is unoccluded, the detection score may be 95%. If the threshold detection score for masking humans in the video stream is set to, for example, 80%, the partially occluded person will not be masked, whereas the unoccluded object will be masked. This is problematic because the partially occluded person should also be masked to avoid identification. [Prior art documents] [Patent documents]

[0010] [Patent Document 1] US Patent Publication No. 20180268240 Summary of the Invention [Problem to be solved by the invention]

[0011] Thus, an objective of the embodiments herein may be to obviate some of the problems mentioned above, or at least reduce their impact. Specifically, an objective of the embodiments herein may be how to detect captured objects in a stream of images that are occluded by other objects, and how to classify them according to known object types. For example, an objective of the embodiments herein may be to detect humans in a stream of video images, even though the humans are not fully visible in the stream of images. Once an object is classified as a human, the classification may lead to human masking or human counting.

[0012] Thus, a further objective of embodiments herein may be to de-identify or anonymize people in a stream of image frames, for example by masking the people, while still being able to determine what is happening in the stream of image frames.

[0013] A further objective may be to improve the determination of probability values ​​indicative of whether an object captured in a stream of image frames belongs to an object type, e.g., an object type that is masked, filtered or counted. In other words, a further objective may be to improve the determination of object detection probability values ​​indicative of the likelihood that a detected object belongs to a particular object category. [Means for solving the problem]

[0014] According to one aspect, this object is achieved by a method, executed in a multi-camera system, for determining a probability value indicative of whether an object captured in a stream of image frames belongs to an object type.

[0015] The method comprises detecting a first object or a first portion of a first object in a first area of ​​a scene captured in a first stream of image frames captured by a first camera of a multi-camera system.

[0016] The method further comprises determining a probability that the detected first object or first portion belongs to the object type based on a characteristic of the first object or a first portion of the first object.

[0017] The method further includes detecting a second object or a second portion of the second object in a second area of ​​the scene captured in a second stream of image frames by a second camera of the camera system different from the first camera, the second area at least partially overlapping with the first area.

[0018] The method further includes determining a second probability value indicating a probability that the detected second object or second portion belongs to the object type based on characteristics of the second object or the second portion of the second object.

[0019] The method further includes determining an updated second probability value by increasing the second probability value when the second probability value is below a second threshold and the first probability value is above the first threshold.

[0020] According to another aspect, the above object is achieved by an image processing device arranged to carry out the above method.

[0021] According to a further aspect, this object is achieved by a computer program and a computer program carrier corresponding to the above aspects.

[0022] The embodiments herein use a probability value from another camera that indicates that a first object or first part detected belongs to an object type.

[0023] The second probability value can be compensated when the second object is occluded, since the second probability value is increased when the second probability value is below the second threshold and the first probability value is above the first threshold. By doing so, it is possible to compensate the second probability value for a particular object having a high first probability value from the first camera. This means that even though the second object is partially occluded in the second camera, there is a high probability of detecting that the second object belongs to a particular object type, while the number of second objects that are erroneously detected as the object type to be detected is still low, i.e. the number of false detections is low. Effect of the Invention

[0024] Thus, an advantage of the embodiments herein is that the probability of detecting an object of a particular object type that is partially occluded in a particular camera is increased. Thus, detection of a particular type or category of objects may be improved. Detection of an object may be determined when the probability value indicates that an object captured in a stream of image frames belongs to an object type. For example, when the second probability value indicates that the second object or the second portion belongs to the object type to be detected, the second object is determined to belong to the object type. For example, a high second probability value may indicate that the second object or the second portion belongs to the object type, and a low second probability value may indicate that the second object or the second portion does not belong to the object type. The high and low probability values ​​may be distinguished by one or more thresholds.

[0025] A further advantage is improved masking of certain types of objects. Yet another advantage is improved counting of certain types of objects.

[0026] Hereinafter, an embodiment of the present invention will be described with reference to the drawings. The various aspects, including particular features and advantages of the embodiments disclosed herein, will be readily understood from the following detailed description and the accompanying drawings. [Brief description of the drawings]

[0027] [Figure 1] FIG. 1 illustrates an exemplary embodiment of an image capture device. [Figure 2a] 1 illustrates an example embodiment of a video network system. [Figure 2b] 1 illustrates an exemplary embodiment of a video network system and user equipment. [Diagram 3] FIG. 1 is a schematic block diagram illustrating an exemplary embodiment of an imaging system. [Figure 4a] FIG. 1 is a schematic block diagram illustrating an embodiment of a method in an image processing apparatus. [Figure 4b] FIG. 1 is a schematic block diagram illustrating an embodiment of a method in an image processing apparatus. [Figure 4c] FIG. 1 is a schematic block diagram illustrating an embodiment of a method in an image processing apparatus. [Figure 4d] FIG. 1 is a schematic block diagram illustrating an embodiment of a method in an image processing apparatus. [Diagram 5] 1 is a flow chart illustrating an embodiment of a method in an image processing device. [Figure 6] 1 is a block diagram showing an embodiment of an image processing apparatus; DETAILED DESCRIPTION OF THE PREFERRED EMBODIMENTS

[0028] The details will be explained below.

[0029] The embodiments herein may be implemented in one or more image processing devices. In some embodiments herein, the one or more image processing devices may include or be one or more image capture devices, such as a digital camera. FIG. 1 illustrates various exemplary image capture devices 110. The image capture devices 110 may be or include, for example, a video camera 120, such as a camcorder, a network video recorder, a camera, a security camera or a monitoring camera, a digital camera, a wireless communication device 130, such as a smartphone including an image sensor, or an automobile 140 including an image sensor.

[0030] 2a illustrates an example video network system 250 in which embodiments herein may be implemented. The video network system 250 may include an image capture device, such as a video camera 120, that can capture digital images 201, such as digital video images, and perform image processing thereon. A video server 260 in FIG. 2a may obtain images from the video camera 120, for example, over a network or the like, which is indicated by a bidirectional arrow in FIG. 2a.

[0031] Video server 260 is a computer-based device dedicated to delivering video. Video servers are used in many applications and often have additional features and capabilities to address the needs of a particular application. For example, video servers used in security, surveillance, and inspection applications are typically designed to capture video from one or more cameras and deliver the video over a computer network connection. In video production and broadcast applications, the video server may have the ability to record and play back recorded video and deliver many video streams simultaneously. Today, many video server functions are built into video cameras 120.

[0032] However, in FIG. 2a, video server 260 is connected via video network system 250 to an image capture device, here exemplified by video camera 120. Video server 260 may be further connected to video storage 270 for storage of video images and / or to monitor 280 for display of video images. In some embodiments, video camera 120 is directly connected to video storage 270 and / or monitor 280, as indicated by the direct arrows between these devices in FIG. 2a. In some other embodiments, video camera 120 is connected to video storage 270 and / or monitor 280 via video server 260, as indicated by the arrows between video server 260 and the other devices.

[0033] 2b shows a user device 295 connected to video camera 120 via video network system 250. User device 295 may be, for example, a computer or a mobile phone. User device 295 may, for example, control video camera 120 and / or display video originating from video camera 120. User device 295 may further include the functionality of both monitor 280 and video storage 270.

[0034] To better understand the embodiments herein, an imaging system is first described.

[0035] 3 is a schematic diagram of an imaging system 300, in this case a digital video camera such as video camera 120. The imaging system images a scene onto an image sensor 301. The image sensor 301 may be equipped with a Bayer filter so that different pixels receive radiation in specific wavelength regions in a known pattern. Typically, each pixel of a captured image is represented by one or more values ​​that represent the intensity of the captured light in a certain wavelength band. These values ​​are usually called color components or color channels. The term "image" can refer to an image frame or a video frame that includes information resulting from the image sensor that captured the image.

[0036] After reading the signals of the individual sensor pixels of the image sensor 301, different image processing actions can be performed by the image signal processor 302. The image signal processor 302 can comprise an image processing section 302a, sometimes called an image processing pipeline, and a video post-processing section 302b.

[0037] Typically, for video processing, images are included in a stream of images. Figure 3 shows a first video stream 310 from an image sensor 301. The first image stream 310 may comprise multiple captured image frames, such as a first captured image frame 311 and a second captured image frame 312.

[0038] Image processing may include demosaicing, color correction, noise filtering (to remove spatial and / or temporal noise), distortion correction (e.g., to remove the effects of barrel distortion), global and / or local tone mapping (e.g., to enable imaging of scenes containing a wide range of intensities), transformations (e.g., rectification and rotation), flat-field correction (e.g., to remove the effects of vignetting), application of overlays (e.g., privacy masks, legends), etc. The image signal processor 302 may also be associated with an analytics engine that performs object detection, recognition, alarms, etc.

[0039] The image processor 302a may, for example, perform image stabilization, apply noise filtering, distortion correction, global and / or local tone mapping, transformations, and flat-field correction. The video post-processor 302b may, for example, crop portions of an image, apply overlays, and include an analysis engine.

[0040] Following the image signal processor 302, the image is transferred to an encoder 303 where the information in the image frames is encoded according to an encoding protocol such as H.264. The encoded image frames are then transferred to, for example, a receiving client, here exemplified by a monitor 280, a video server 260, a storage device 270, etc.

[0041] The video encoding process produces a number of values ​​that can be encoded to form a compressed bitstream. These values ​​may include the following: Quantized transform coefficients information to enable the decoder to recreate the prediction Information about the structure of the compressed data and the compression tool used during encoding; - Information about the complete video sequence.

[0042] These values ​​and parameters (syntax elements) are converted into binary code, for example using variable length coding and / or arithmetic coding. Each of these coding methods produces an efficient and compact binary representation of the information, also called a coded bitstream. The coded bitstream can then be stored and / or transmitted.

[0043] In a camera system having multiple cameras, the embodiments of the present specification propose to improve the masking quality in the camera by using detection information from surrounding cameras. For example, the masking quality can be improved by increasing the probability of detecting an object of a certain object type even when the object of the certain object type is partially occluded in one of the cameras, but the number of objects that are erroneously detected as the object type to be detected is still low, i.e., the number of false detections is low.

[0044] The detection information may include a detection score or a detection probability. Different cameras may have different prerequisites for detecting people due to the camera resolution, different camera installation positions, and their processing power, which affects the characteristics of the video network used with the cameras. An example of these characteristics is the size of the image input to the video network. A person who is slightly occluded in the image frame of one camera may be completely visible in the image frame of another camera. A person who has a pixel density that is too low to be detected by one camera may be detectable by another camera where the pixel density of the person is higher.

[0045] The camera system may need to understand the spatial overlap between the cameras. For example, the camera system may be presented with or may determine the overlap of areas captured by the cameras that will share detection information. This may be done via known methods of scene segmentation using information of geolocation, pan and tilt angles, zoom, field of view, and installation height. An image processor, such as a camera, collects detection scores from surrounding cameras for objects in a region captured by two or more cameras, sometimes called the spatial overlap region. The image processor adjusts the detection score accordingly based on the detection score from at least one other camera. The image processing device may apply a privacy mask to the detected object based on the adjusted detection score, which is described in more detail below, e.g., in connection with action 503 of FIG.

[0046] 4a, 4b, 4c, 4d and 5, and further with reference to FIGS. 1, 2a, 2b and 3, exemplary embodiments herein will now be described.

[0047] 4a shows a scene including an object, such as a person, obstructed by another object, such as a wall, and a multi-camera system 400. The multi-camera system 400 comprises at least a first camera 401 and a second camera 402. The first camera 401 and the second camera 402 may be video cameras. The multi-camera system 400 may also comprise a video server 460. The first camera 401 and the second camera 402 may capture the object from different angles.

[0048] In a scenario in which the embodiments may be implemented, the first camera 401 may capture most of the entire person, while the second camera 402 captures only a smaller part of the person, such as the person's head and arms. Each of the first camera 401 and the second camera 402 may be assigned a detection score, e.g., a probability value indicating that the detected object belongs to an object type, such as a human.

[0049] FIG. 4b illustrates a first stream of image frames 421 captured by a first camera 401 of a multi-camera system 400 and a second stream of image frames 422 captured by a second camera 402 of the camera system 400. The second camera 402 is different from the first camera 401. The first stream of image frames 421 includes an image frame 421_2 capturing a first object 431. The first object 431 may include a first portion 431a. The second stream of image frames 422 includes an image frame 422_2 capturing a second object 432. The second object 432 may include a second portion 432a. In the scenario herein, the first object 431 may correspond to the second object 432 in that they are the same real-world object captured. Also, the first portion 431a may correspond to the second portion 432a.

[0050] 5 shows a flow chart illustrating a method for determining a probability value indicative of the probability that an object captured in a stream of image frames 422 belongs to an object type, such as a particular object type. For example, the object type may be a human or an animal. In other words, the method may be for determining the probability that an object belongs to an object type.

[0051] The method may be performed in the camera system 400, and more specifically, for masking or counting objects captured in a stream of image frames. The masking or counting of objects may be performed when a probability value indicating that an object captured in the stream of image frames 422 belongs to an object type exceeds a threshold for masking objects or a threshold for counting objects. Also, the threshold for masking objects and the threshold for counting objects may be different.

[0052] In particular, the embodiment may be performed by an image processing device of the multi-camera system 400. The image processing device may be either the first camera 401 and the second camera 402 or the video server 460. The first camera 401 and the second camera 402 may each be a video camera such as a surveillance camera.

[0053] The following actions may be performed in any suitable order, such as in an order other than the order presented below.

[0054] Action 501 The method comprises detecting a first object 431 or a first portion 431a of the first object 431 in a first area 441 of a scene captured in a first stream 421 of image frames captured by a first camera 401 of the multi-camera system 400.

[0055] Action 502 The method further comprises determining a first probability value indicative of a first probability that the detected first object 431 or first portion 431a belongs to an object type, such as a human, based on a characteristic of the first object 431 or the first portion 431a of the first object 431. In other words, the method comprises determining a first probability that the first object belongs to an object type.

[0056] For example, the determination may be performed by weighting probability values ​​for several parts of the first object 431. The respective probability values ​​indicating the respective probabilities that the detected part 431a of the first object 431 belongs to an object part type may be determined based on characteristics of the detected part 431a of the first object 431.

[0057] Action 503 If the first probability value is above a first threshold, the detected first object 431 or first portion 431a may be determined to belong to an object type, such as a human. Furthermore, if the first probability value is above a first threshold, some further action may be performed. For example, a privacy mask, such as a masking filter or a masking overlay, may be applied to at least a portion of the first object 431, such as the first object 431 or the first portion 431a of the first object 431, in the first stream of image frames 421 to mask out, e.g., hide, at least a portion of the first object 431, such as the first object 431 or the first portion 431a of the first object 431, in the first stream of image frames 421. In particular, the privacy mask may be obtained by masking regions in each image frame, and may include color masking, also called solid color masking or monochrome masking, mosaic masking, also called pixelation, pixelated privacy masking or transparent pixelation, Gaussian blur, Sobel filter, background model masking by applying the background as a mask (objects appear transparent), chameleon mask (mask that changes color depending on the background). As a specific example, a color mask can be obtained by ignoring the image data and using other pixel values ​​for the pixels to be masked. For example, the other pixel values ​​may correspond to a particular color such as red or gray.

[0058] In another embodiment, the object may be counted as a particular predefined object if the first probability value exceeds a first threshold. Thus, the first threshold may be a masking threshold, a counting threshold, or another threshold associated with some other function to be performed on the object as a result of the first probability value exceeding the first threshold.

[0059] Action 504 The method further comprises detecting a second object 432 or a second portion 432a of the second object 432 in a second area 442 of the scene captured in a second stream of image frames 422 by a second camera 402 of the camera system 400. The second camera 402 is different from the first camera 401. The second area 442 at least partially overlaps with the first area 441.

[0060] Action 505 The method further comprises determining a second probability value indicative of a second probability that the detected second object 432 or second portion 432a belongs to the object type based on a characteristic of the second object 432 or the second portion 432a of the second object 432. In other words, the method comprises determining a second probability that the second object belongs to the object type.

[0061] If the second probability value is above a second threshold, the detected second object 432 or second portion 432a may be determined to belong to an object type such as a human. In other words, the detected second object 432 or second portion 432a may be detected as a human or as belonging to a human.

[0062] If the second probability value is below the second threshold, it may be determined that the second object 432 does not belong to the object type. However, in embodiments herein, the method continues to evaluate whether the second object 432 belongs to the object type even when the second probability value is below the second threshold by considering the first probability value from the first camera 401.

[0063] Action 506 Yet another condition for continued evaluation of the detection of the second object 432 based on the first probability value may be a determination of the co-location of the first object 431 and the second object 432 to ensure that the first object 431 is the same object as the second object 432 and therefore it is appropriate to take the first probability value into account when evaluating the detection of the second object 432.

[0064] Thus, the method may further comprise determining that the second object 432 or the second portion 432a of the second object 432 and the first object 431 or the first portion 431a of the first object 431 are co-located within an overlap area of ​​the second area 442 and the first area 441.

[0065] In other words, the method may further include determining that the first object 431 and the second object 432 are the same object.

[0066] For example, a first person in the first stream of image frames 421 may be determined to be co-located with a second person in the second stream of image frames 422. Based on the co-location determination, it may be determined that the first person is the same person as the second person. For example, if the first object 431 is determined to be a person with a 90% probability, and the second object 432 is further determined to be a person with a 70% probability, and there are no other objects detected in the image frames, it may be determined that the people are likely the same person.

[0067] In other words, the first object 431 and the second object 432 may be determined to be linked by, for example, determining that the motion track entries relating to the first object 431 and the second object 432 are linked (i.e., determined to belong to one real-world object) given a matching level based on a set of constraints. The track entries may be created based on the received first and second image series, the determined first and second character-descriptive feature sets, and the movement data.

[0068] In another example, the second object 432 may be determined to be co-located with the first portion 431 a of the first object 431 .

[0069] In another example, a second portion 432a of a second object 432 may be determined to be co-located with a first portion of a first object 431. For example, a head of the first object may be determined to be in the same location as a head of a second object. This method may be similarly applied to other parts such as arms, torsos, and legs. Operation 506 may be performed after or before operation 507.

[0070] Action 507 If the second probability value is lower than the second threshold and the first probability value is higher than the first threshold, the method includes determining an updated second probability value by increasing the second probability value.

[0071] The second threshold may also be a masking threshold, a counting threshold, or another threshold associated with some other function to be performed on the second object 432 or the second portion 432a as a result of the second probability value exceeding the second threshold.

[0072] The second threshold and the first threshold may be different. However, in some embodiments herein, they are the same. Figure 4c shows a simple example of how the second probability value can be increased from a value below the second threshold to a value above the second threshold when the first threshold and the second threshold are the same and the first probability value is above the first threshold.

[0073] In some embodiments herein, the first threshold is 90% and the second threshold is 70%. The second threshold may be lower than the first threshold. This may be advantageous, for example, when the second object 432 is partially occluded by some other object, such as a wall.

[0074] In one scenario, the first camera 401 determines a first probability value for the first object being a human being to be 95%. The second camera 402 determines a second probability value for the second object being a human being to be 68%. The second probability value may then be increased, for example, by 5 or 10%. The increase may be a fixed value as long as the first probability value is above the first threshold. However, in some embodiments herein, the increase in the second probability value is based on the first probability value. For example, the increase in the second probability value may be proportional to the first probability value. In some embodiments herein, when the first probability value is 60%, the increase in the second probability value is 5%, when the first probability value is 70%, the increase in the second probability value is 10%, and when the first probability value is 80%, the increase in the second probability value is 15%.

[0075] Increasing the second probability value based on the first probability value may comprise determining a difference between the first probability value and a first threshold and increasing the second probability value based on the difference. For example, if the difference between the first probability value and the first threshold is 5%, the second probability value may be increased by 10%. In another example, if the difference between the first probability value and the first threshold is 10%, the second probability value may be increased by 20%.

[0076] In some other embodiments, increasing the second probability value based on the first probability value includes increasing the second probability value with the difference. For example, if the difference between the first probability value and the first threshold is 5%, the second probability value may be increased by 5%. In another example, if the first probability value is 95% and a common threshold such as a common masking threshold is 80%, the difference is 15%, and thus, for example, a second probability value of 67% may be increased by 15%, resulting in 82% above the common threshold of 80%. As a result, the second object may also be masked or counted.

[0077] In some embodiments herein, the updated second probability value is determined in response to determining, in accordance with action 506 above, that the second object 432 or a portion 432a of the second object 432 and the first object 431 or a portion 431a of the first object 431 are co-located within the overlap area.

[0078] In some embodiments herein, each first and second threshold is specific to a type of object or part of an object. For example, each first and second threshold may be specific to different object types or to different object part types, such as head, arm, torso, leg, etc., or both. A masking threshold for the face may be lower than another masking threshold for the arm, since it is more likely to be important to mask the face than the arm. For example, the masking threshold for the head may be 70%, while the masking threshold for the arm may be 90% and the masking threshold for the torso may be 80%.

[0079] A further condition for determining the updated second probability value by increasing the second probability value may be that the second probability value is above a third threshold in addition to being below the second threshold. Thus, an additional condition may be that the second probability value is above a lower third threshold. The reason why it is advantageous for the second probability value to be higher than the lower third threshold is that if the second probability value is lower than the third threshold, e.g., less than 20%, the object detector is quite confident that the object is not an object to be masked, e.g., not a human. However, if the second probability value is between 20% and 80%, the object may be an occluded person, and therefore, if another camera is more confident in detecting the object as an object to be masked, the probability value may be increased.

[0080] In some embodiments herein, the second object part type of the second portion 432a is the same object part type as the first object part type of the first portion 431a. For example, the method can compare a first face to a second face, a first leg to a second leg, etc.

[0081] It may be advantageous to increase the second probability value only when the difference between the first probability value and the first threshold exceeds a fourth threshold, for example, when the difference exceeds 10%.In this way, the number of false positives can be controlled.For example, the larger the threshold difference, the smaller the number of false positives can be.

[0082] For example, in a scenario where the first threshold is 70%, the second threshold is 90% and the fourth threshold is 75%, a first probability value of 72% may be considered large enough to detect the first object as a human in the first stream of image frames 421 but too low to increase the second probability value.

[0083] Action 508 In some embodiments herein, in response to the updated second probability value exceeding the second threshold and in response to determining that the second object belongs to an object type, further actions may be taken. For example, if the object type is an object type to be masked out, the first and second thresholds may be for masking an object or a portion of an object of the object type to be masked out. The method then further includes applying a privacy mask to at least a portion of the second object 432, such as the second portion 432a in the second stream of image frames 422, to mask out the portion when the updated second probability value exceeds the second threshold. For example, if a human face is captured by both the first camera 401 and the second camera 402 and a face score is updated for the second camera 402, this may lead to privacy masking of the face in the second image stream 422. Thus, the method may further include anonymizing the unidentified person by removing the identifying features, for example, by any of the methods described above in action 503.

[0084] In some other embodiments, the object type is an object type to be counted. In this case, the first and second thresholds may be for counting objects or parts of objects of the object type to be counted. The method then further includes increasing 422 a counter value associated with the second stream of image frames when the updated second probability value is above the second threshold. For example, the counter value may be for an object or an object part.

[0085] FIG. 4d illustrates a scenario in which a scene includes a first human 451 and a second human 452, and a third object 453 between the two humans 451, 452. The shape of the third object 453 resembles the shape of a body part of the humans 451, 452. The third object may be, for example, a balloon resembling a head. In one example, the second camera 402 captures three objects 451, 452, 453, and the method of FIG. 5 may be used for each of the three objects 451, 452, 453. Thus, each of the three objects 451, 452, 453 in FIG. 4d may be a second object 432 in the second stream of image frames 422. Correspondingly, each of the three objects in FIG. 4d may be a first object 431 in the first stream of image frames 421.

[0086] The method may be performed by a second camera, or by the video server 460, or even by the first camera 401.

[0087] The method may include determining a second probability value. The second probability value may be 69%, 70%, and 80% for the first human 451 (left), the third object 453 (balloon), and the second human 452 (right), respectively. In one scenario, the second threshold for detecting and masking humans is 80%. This means that without the method of FIG. 5, only the second human 452 would be masked. If a general reduction of the second threshold is performed by 10%, the third object 453 would be masked but the first human 451 would not be masked. However, if the first probability value derived from the first stream of image frames 421 for each first object 431, such as the first human 451 and the second human 452, is higher than the first threshold and the second probability value is lower than the second threshold, the second probability value is increased. Thus, if a first human 451 is detected in the first stream of image frames 421 from the first camera 401 with the first probability value being 90%, indicating that the detected first human 451 belongs to a human type, the second probability value is increased, for example, by 15%. The updated second probability value is 84%, which is above the second threshold, and the first human 451 is masked. The first probability value of the third object 453, indicating the probability that the detected third object 453 belongs to a human object type, is low, for example 18%. This means that the second probability value is not increased, since the first probability value of the third object 453 is below the first threshold. Thus, the third object 453 is not masked in either the stream of images from the first camera 401 or the second camera 402.

[0088] 6, there is shown a schematic block diagram of an embodiment of an image processing device 600. As mentioned above, the image processing device 600 is configured to determine a probability value indicating that an object captured in a stream of image frames belongs to an object type. Furthermore, the image processing device 600 can be part of the multi-camera system 400.

[0089] As mentioned above, the image processing device 600 may include or be any of a camera, such as a surveillance camera, a camcorder, a network video recorder, and a wireless communication device 130. In particular, the image processing device 600 may be a first camera 401 or a second camera 402, such as a surveillance camera, or a video server 460, which may be part of a multi-camera system 400. The method of determining a probability value indicating that an object captured in a stream of image frames belongs to an object type may also be performed in a distributed manner in several image processing devices, such as in the first camera 401 and the second camera 402. For example, the actions 501 to 503 may be performed by the first camera 401, while the actions 504 to 508 may be performed by the second camera 402.

[0090] The image processing device 600 may further include a processing module 601, such as means for performing the methods described herein, which may be embodied in the form of one or more hardware modules and / or one or more software modules.

[0091] The image processing device 600 may further comprise a memory 602. The memory may include, for example contain or store, instructions, for example in the form of a computing program 603, which may comprise computer readable code units which, when executed on the image processing device 600, cause the image processing device 600 to perform a method for determining a probability value that is indicative of an object captured in a stream of image frames belonging to an object type. The image processing device 600 may comprise a computer, in which case the computer readable code units, when executed on the computer, cause the computer to perform the method for determining a probability value that is indicative of an object captured in a stream of image frames belonging to an object type.

[0092] According to some embodiments herein, image processing device 600 and / or processing module 601 include processing circuitry 604 as an example hardware module that may include one or more processors. Thus, processing module 601 may be embodied in the form of, or "implemented by," processing circuitry 604. Instructions are executable by processing circuitry 604 such that image processing device 600 operates to perform the method of FIG. 5 as described above. As another example, instructions, when executed by image processing device 600 and / or processing circuitry 604, may cause image processing device 600 to perform the method according to FIG. 5.

[0093] In view of the above, in one example, an image processing device 600 is provided for determining a probability value that indicates that an object captured in a stream of image frames belongs to an object type.

[0094] Again, the memory 602 includes instructions executable by the processing circuitry 604 such that the image processing apparatus 600 operates to perform the method according to FIG.

[0095] 6 further illustrates a carrier 605, or program carrier, that contains the immediately above described computing program 603. The carrier 605 may be one of an electronic signal, an optical signal, a radio signal, and a computer readable medium.

[0096] In some embodiments, the image processing device 600 and / or the processing module 601 may include one or more of the following exemplary hardware modules: a detection module 610, a determination module 620, a masking module 630, and a counting module 640. In other examples, one or more of the above-mentioned exemplary hardware modules may be implemented as one or more software modules.

[0097] Furthermore, the processing module 601 may comprise an input / output unit 606. According to one embodiment, the input / output unit 606 may comprise an image sensor configured to capture the raw image frames mentioned above, such as the raw image frames included in the video stream 310 from the image sensor 301.

[0098] According to various embodiments described above, the image processing device 600 and / or the processing module 601 and / or the detection module 610 are configured to receive captured image frames 311 , 312 of the image stream 310 from the image sensor 301 of the image processing device 600 .

[0099] The image processing device 600 and / or the processing module 601 and / or the detection module 610 are configured to detect a first object 431 or a first part 431a of the first object 431 in a first region 441 of a scene captured in a first stream 421 of image frames captured by a first camera 401 of the multi-camera system 400.

[0100] The image processing device 600 and / or the processing module 601 and / or the determination module 620 are further configured to determine a first probability value indicating a first probability that the detected first object 431 or first part 431a belongs to an object type based on characteristics of the first object 431 or the first part 431a of the first object 431.

[0101] The image processing device 600 and / or the processing module 601 and / or the detection module 610 are further configured to detect a second object 432 or a second portion 432a of the second object 432 in a second region 442 of the scene captured in the second stream of image frames 422 by a second camera 402 of the camera system 400. The second camera 402 is different from the first camera 401. The second region 442 at least partially overlaps with the first region 441.

[0102] The image processing device 600 and / or the processing module 601 and / or the determination module 620 are further configured to determine a second probability value indicating a second probability that the detected second object 432 or the second portion 432a of the second object 432 belongs to the object type based on characteristics of the second object 432 or the second portion 432a of the second object 432.

[0103] The image processing device 600 and / or the processing module 601 and / or the detection module 610 are further configured to determine an updated second probability value by increasing the second probability value when the second probability value is below a second threshold and the first probability value is above the first threshold.

[0104] The image processing device 600 and / or the processing module 601 and / or the decision module 620 may be further configured to increase the second probability value based on the first probability value.

[0105] The image processing device 600 and / or the processing module 601 and / or the decision module 620 may be further configured to increase the second probability value based on the first probability value by determining a difference between the first probability value and a first threshold and increasing the second probability value based on this difference.

[0106] The image processing device 600 and / or the processing module 601 and / or the masking module 630 may be further configured to apply a privacy mask to at least a portion of the second object 432 in the second stream of image frames 422 if the updated second probability value exceeds a second threshold.

[0107] The image processing device 600 and / or the processing module 601 and / or the counting module 640 may be further configured to increase a counter value associated with the second stream of image frames 422 if the updated second probability value exceeds a second threshold.

[0108] The image processing device 600 and / or the processing module 601 and / or the decision module 620 may be further configured to increase the second probability value if the second probability value is above a third threshold in addition to being below the second threshold.

[0109] The image processing device 600 and / or the processing module 601 and / or the determination module 620 may be further configured to determine that the second object 432 or the second portion 432a of the second object and the first object 431 or the first portion 431a of the first object 431 are co-located within the overlap region between the second region 442 and the first region 441, and to determine an updated second probability value in response to determining that the second object 432 or the second portion 432a of the second object 432 and the first object 431 or the first portion 431a of the first object 431 are co-located within the overlap region.

[0110] As used herein, the term "module" may refer to one or more functional modules, each of which may be implemented as one or more hardware modules and / or one or more software modules and / or combined software / hardware modules. In some examples, a module may represent a functional unit that is realized as software and / or hardware.

[0111] As used herein, the term "electronic program carrier," "program carrier," or "carrier" can refer to one of an electronic signal, an optical signal, a radio signal, and a computer-readable medium. In some examples, the electronic program carrier may exclude transitory propagating signals, such as electronic signals, optical signals, and / or radio signals. Thus, in these examples, the electronic program carrier may be a non-transitory carrier, such as a non-transitory computer-readable medium.

[0112] As used herein, the term "processing module" may include one or more hardware modules, one or more software modules, or a combination thereof. Any such module, be it a hardware, software, or combination hardware and software module, may be a connecting means, a providing means, a configuring means, a responding means, a disabling means, etc., as disclosed herein. As an example, the term "means" may be a module corresponding to the modules listed above in conjunction with the figures.

[0113] As used herein, the term "software module" may refer to a software application, a dynamic link library (DLL), a software component, a software object, an object under the Component Object Model® (COM), a software component, a software function, a software engine, an executable binary software file, and the like.

[0114] The terms "processing module" or "processing circuitry" as used herein may encompass processing units including, for example, one or more processors, application specific integrated circuits (ASICs), field programmable gate arrays (FPGAs), etc. A processing circuitry or the like may comprise one or more processor kernels.

[0115] As used herein, the phrase "configured / for" may mean that a processing circuit is configured, e.g., adapted, or operable, by a software and / or hardware configuration, to perform one or more of the actions described herein.

[0116] As used herein, the term "action" may refer to an action, step, operation, response, reaction, activity, etc. It should be noted that an action herein may be divided into two or more sub-actions, where applicable. Additionally, it should be noted that two or more actions described herein may be merged into a single action, where applicable.

[0117] As used herein, the term "memory" can refer to a hard disk, a magnetic storage medium, a portable computer diskette or disk, a flash memory, a random access memory (RAM), etc. Memory may also refer to a processor's internal register memory, etc.

[0118] As used herein, the term "computer-readable medium" may be a Universal Serial Bus memory, a DVD disc, a Blu-ray disc, a software module received as a stream of data, a flash memory, a hard drive, a memory stick, a Multimedia Card (MMC), a memory card such as a Secure Digital (SD) card, etc. One or more of the foregoing examples of computer-readable medium may be provided as one or more computing program products.

[0119] As used herein, the term "computer readable code unit" may refer to the text of a computer program, a portion or an entire computer program representing a binary file in compiled format, or anything in between. As used herein, the terms "number" and / or "value" may be any type of number, such as a binary number, a real number, an imaginary number, or a rational number. Furthermore, a "number" and / or "value" may be one or more characters, such as a character or a string of characters. A "number" and / or "value" may be represented by a string of bits, i.e., 0s and / or 1s.

[0120] As used herein, the phrase "in some embodiments" is used to indicate that features of the described embodiments can be combined with any other embodiment disclosed herein.

[0121] While embodiments of various aspects have been described, many different changes, modifications, etc. thereof will become apparent to those skilled in the art. Accordingly, the described embodiments are not intended to limit the scope of the present disclosure.

Claims

1. 1. A method for determining a probability value indicative of a probability that an object captured in a stream of image frames belongs to an object type, the method comprising the steps of: Detecting (501) a first object (431) or a first part (431a) of the first object (431) in a first region (441) of a scene captured in a first stream (421) of image frames captured by a first camera (401) of a multi-camera system (400); determining (502) a first probability value indicative of a first probability that the detected first object (431) or the first part (431a) belongs to an object type based on characteristics of the first object (431) or the first part (431a) of the first object (431); determining (503) that the detected first object (431) or first part (431a) belongs to the object type if the first probability value is above a first threshold; detecting (504) a second object (432) or a second portion (432a) of the second object (432) in a second region (442) of the scene captured in a second stream of image frames (422) by a second camera (402) of the camera system (400) different from the first camera (401), the second region (442) at least partially overlapping the first region (441); determining (505) a second probability value indicative of a second probability that the detected second object (432) or second portion (432a) belongs to the object type based on characteristics of the second object (432) or the second portion (432a) of the second object (432), and determining that the detected second object (432) or second portion (432a) belongs to the object type if the second probability value is above a second threshold; determining (507) an updated second probability value by increasing the second probability value if the second probability value is below the second threshold and the first probability value is above the first threshold, wherein increasing the second probability value is based on the first probability value; determining a difference between the first probability value and a first threshold; and increasing the second probability value based on a difference between the first probability value and a first threshold.

2. The method described in claim 1, wherein increasing the second probability value based on the difference between the first probability value and the first threshold includes increasing the second probability value in accordance with the difference.

3. The method of claim 1 or 2, wherein the second threshold is lower than the first threshold.

4. A method according to any one of claims 1 to 3, further comprising determining that the second object belongs to the object type in response to the updated second probability value exceeding a second threshold.

5. 5. The method of claim 1, further comprising: applying (508) a privacy mask to at least a portion of a second object (432) in a second stream of image frames (422) if the updated second probability value exceeds the second threshold; and wherein the object type is an object type to be masked out, and the first and second thresholds are for masking an object or a portion of an object of the object type to be masked out, and if the updated second probability value exceeds the second threshold, applying (508) a privacy mask to at least a portion of a second object (432) in a second stream of image frames (422).

6. 5. The method of claim 1, further comprising: increasing a counter value associated with the second stream of image frames (422) if the updated second probability value exceeds the second threshold value, the object type being a counted object type and the first and second threshold values ​​being for counting objects or parts of objects of the counted object type.

7. The method of any one of claims 1 to 6, wherein the first and second thresholds are specific to a type of object or part of an object, respectively.

8. 8. The method according to claim 1, wherein a further condition for determining an updated second probability value by increasing the second probability value is that the second probability value, in addition to being below the second threshold, is above a third threshold.

9. 9. The method of claim 1, further comprising: determining (506) that the second object (432) or the second portion (432a) of the second object and the first object (431) or the first portion (431a) of the first object (431) are co-located within an overlap region of the second region (442) and the first region (441); and determining an updated second probability value in response to determining that the second object (432) or the second portion (432a) of the second object (432) and the first object (431) or the first portion (431a) of the first object (431) are co-located within the overlap region.

10. 10. The method of any one of claims 1 to 9, wherein the second object part type of the second part (432a) is the same object part type as the first object part type of the first part (431a).

11. An image processing device (402, 460) of a multi-camera system (400) configured to perform a method according to any one of the preceding claims.

12. The image processing device (402, 460) according to claim 11, wherein the image processing device (402, 460) is a camera (402) such as a surveillance camera or a video server (460).

13. A computer program (603) comprising computer readable code units which, when executed on an image processing device (110), cause the image processing device (110) to perform a method according to any one of claims 1 to 10.

14. A computer readable medium (605) comprising a computer program according to claim 13.