Method and image processing apparatus for determining a probability value indicating that an object captured in a stream of image frames belongs to an object type

Through the information interaction adjustment in the multi-camera system, the problem of inaccurate detection of some occluded objects in video surveillance is solved, and more accurate object masking and counting is achieved.

CN117768605BActive Publication Date: 2025-08-29AXIS
View PDF 3 Cites 0 Cited by

Patent Information

Application Number
CN202311207032.4
Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Priority Date
2022-09-23
Filing Date
2023-09-19
Publication Date
2025-08-29
Estimated Expiration
2043-09-19

AI Technical Summary

Technical Problem

In the prior art, some of the obscured objects are difficult to be effectively detected and masked in video surveillance, resulting in the inability to accurately perform anonymization.

Method used

Through the multi-camera system, the detection information of the first camera and the second camera compensate each other, and adjust the detection probability value of the second camera to ensure that the partially obscured objects are correctly classified and masked in the image frame stream.

Benefits of technology

The detection accuracy of partially obscured objects is improved, false positive detection is reduced, and effective anonymization and counting of objects in the image frame stream is achieved.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN117768605B_ABST
    Figure CN117768605B_ABST
Patent Text Reader

Abstract

A method comprises: detecting (501) a first object or a first part of the first object in a first area of ​​a scene captured in a first image frame stream captured by a first camera of a multi-camera system; determining (502) a first probability value indicating a first probability that the detected first object or first part belongs to an object type based on characteristics of the first object or the part of the first object; detecting (504) a second object or a second part of the second object in a second area of ​​the scene captured in a second image frame stream by a second camera of the camera system that is different from the first camera, wherein the second area at least partially overlaps with the first area; determining (505) a second probability value indicating a second probability that the detected second object or second part belongs to the object type based on characteristics of the second object or the second part of the second object; and determining (507) an updated second probability value by increasing the second probability value when the second probability value is below a second threshold value and the first probability value is above the first threshold value.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] Embodiments herein relate to a method and an image processing device for determining a probability value indicating that an object captured in a stream of image frames belongs to an object type. A corresponding computer program and a computer program carrier are also disclosed. Background Art

[0002] The use of imaging, particularly video imaging, for public surveillance is common in many areas around the world. Examples of areas where surveillance may be necessary include banks, stores, and other areas requiring security, such as schools and government facilities. However, in many locations, installing cameras without a permit is illegal. Other areas where surveillance may be necessary include processing, manufacturing, and logistics applications, where video surveillance is primarily used to monitor processes.

[0003] However, there may be requirements that prevent the identification of people from video surveillance. This requirement may conflict with the requirement to be able to determine what is happening in the video. For example, performing people counting or queue monitoring on anonymized image data may be of interest. In practice, there is a trade-off between meeting these two requirements: non-identifiable video and extracting large amounts of data for different purposes (such as people counting).

[0004] Several image processing techniques have been described to avoid identifying people while still being able to identify activities. For example, edge detection / representation, edge enhancement, silhouetting objects, and different types of "blurring" (such as color changes or dilation) are examples of such operations. Privacy masking is another image processing technique used in video surveillance to protect the privacy of individuals by hiding parts of the image from view using masked areas.

[0005] Image processing refers to any manipulation applied to an image. This can include applying various effects, masks, filters, etc. to the image. In this way, the image can, for example, be sharpened, converted to grayscale, or altered in some way. The image is typically captured by a video camera, a still camera, or the like.

[0006] As mentioned above, one way to avoid identifying people is to mask moving people and objects in the image in real time. Masking in live and recorded video can be done by comparing the live camera view to a set background scene and applying dynamic masking to the changing areas (essentially moving people and objects). Color masking, which can also be called solid color masking or monochrome masking in which an object is masked by an overlaid solid mask of a certain color, provides privacy protection while allowing you to see movement. Mosaic masking, which is also called pixelation, pixelated privacy masking, or transparent pixelated masking, displays moving objects at a lower resolution and allows you to better distinguish the form by looking at the color of the object.

[0007] Masking live and recorded video is suitable for remote video surveillance or recording in areas where surveillance is problematic due to privacy rules and regulations. When video surveillance is primarily used to monitor processes, it is well-suited for processing, manufacturing, and logistics applications. Other potential applications include retail, education, and government facilities.

[0008] Before masking an object, it may be necessary to detect the object as an object to be masked, or in other words, classify it as an object to be masked. US20180268240A1 discloses a method comprising acquiring a video of a scene, detecting objects in the scene, determining an object detection probability value indicating a likelihood that the detected object belongs to an object category to be redacted, and editing the video by obfuscating the detected object belonging to the object category to be redacted.

[0009] One challenging issue encountered when using dynamic masking in a surveillance system is that detection of a partially occluded person will give a much lower detection score, e.g., a lower object detection probability value, than the detection score obtained for detection of a non-occluded person. The occlusion may be due to another object in the scene. The lower detection score may result in a partially occluded person not being masked in the captured video stream. For example, if half of an object (e.g., a person) captured in a video stream is occluded, the detection score obtained from the object detector may be 67% that the object is a human, whereas if the captured object is not occluded, the detection score may be 95%. If the threshold detection score for masking people in a video stream is set to, for example, 80%, then the partially occluded person will not be masked, while the non-occluded object will be. This is a problem because partially occluded people should also be masked to avoid identification. Summary of the Invention

[0010] Therefore, embodiments herein may aim to eliminate some of the aforementioned problems, or at least reduce their impact. Specifically, embodiments herein may aim to detect captured objects that are obscured by other objects in an image stream and to classify them according to known object types. For example, embodiments herein may aim to detect humans in a video image stream, even though the humans are not fully visible in the image stream. Once the object is classified as a human, this classification may result in the obscuration of the human or the counting of the human.

[0011] Therefore, another purpose of embodiments herein may be to de-identify or anonymize a person in a stream of image frames, for example by masking the person, while still being able to determine what is occurring in the stream of image frames.

[0012] Another purpose may be to improve the determination of a probability value indicating whether an object captured in a stream of image frames belongs to an object type (e.g., an object type to be masked, filtered, or counted). In other words, another purpose may be to improve the determination of an object detection probability value indicating the likelihood that a detected object belongs to a particular object class.

[0013] According to one aspect, the object is achieved by a method performed in a multi-camera system for determining a probability value indicating whether an object captured in a stream of image frames belongs to an object type.

[0014] The method includes detecting a first object or a first portion of the first object in a first area of ​​a scene captured in a first stream of image frames captured by a first camera of the multi-camera system.

[0015] The method further comprises determining a first probability value indicating a probability that the detected first object or first part belongs to the object type based on characteristics of the first object or the first part of the first object.

[0016] The method also includes detecting a second object or a second portion of the second object in a second area of ​​the scene captured in a second image frame stream captured by a second camera of the camera system that is different from the first camera, wherein the second area at least partially overlaps the first area.

[0017] The method further comprises determining a second probability value indicating a probability that the detected second object or the second part belongs to the object type based on characteristics of the second object or the second part of the second object.

[0018] The method further includes determining an updated second probability value by increasing the second probability value when the second probability value is below a second threshold and the first probability value is above a first threshold.

[0019] According to another aspect, the object is achieved by an image processing apparatus configured to perform the above method.

[0020] According to a further aspect, the object is achieved by a computer program and a computer program carrier corresponding to the above aspects.

[0021] Embodiments herein use a probability value from another camera indicating that the detected first object or first part belongs to the object type.

[0022] Since the second probability value is increased when the second probability value is below the second threshold and the first probability value is above the first threshold, the second probability value can be compensated when the second object is occluded. By doing so, the second probability value of a particular object from the first camera with a high first probability value can be compensated. This means that the probability of detecting that a second object belongs to a particular object type is increased, even if it is partially occluded in the second camera, while still having a small number of second objects falsely detected as the object type to be detected, i.e., a small number of false positive detections.

[0023] Thus, an advantage of embodiments herein is an increased probability of detecting an object of a particular object type that is partially occluded in a particular camera. Thus, detection of certain types or classes of objects can be improved. Detection of an object can be determined when a probability value indicates that an object captured in the image frame stream belongs to an object type. For example, a determination is made that the second object belongs to the object type to be detected when the second probability value indicates that the second object or the second portion belongs to the object type. For example, a high second probability value can indicate that the second object or the second portion belongs to the object type, while a low second probability value can indicate that the second object or the second portion does not belong to the object type. High and low probability values ​​can be identified by one or more threshold values.

[0024] Another advantage is improved masking of certain types of objects. Yet another advantage is improved counting of certain types of objects. BRIEF DESCRIPTION OF THE DRAWINGS

[0025] Various aspects of the embodiments disclosed herein, including their specific features and advantages, will be readily understood from the following detailed description and accompanying drawings, in which:

[0026] Figure 1 An exemplary embodiment of an image capture device is shown,

[0027] Figure 2a shows an exemplary embodiment of a video network system,

[0028] Figure 2b An exemplary embodiment of a video network system and user equipment is shown.

[0029] Figure 3 is a schematic block diagram illustrating an exemplary embodiment of an imaging system,

[0030] Figure 4a is a schematic block diagram illustrating an embodiment of a method in an image processing device,

[0031] Figure 4b is a schematic block diagram illustrating an embodiment of a method in an image processing device,

[0032] Figure 4c is a schematic block diagram illustrating an embodiment of a method in an image processing device,

[0033] Figure 4d is a schematic block diagram illustrating an embodiment of a method in an image processing device,

[0034] Figure 5 is a flowchart illustrating an embodiment of a method in an image processing device,

[0035] Figure 6 is a block diagram illustrating an embodiment of an image processing apparatus. DETAILED DESCRIPTION

[0036] The embodiments herein may be implemented in one or more image processing devices. In some embodiments herein, the one or more image processing devices may include or be one or more image capture devices, such as digital cameras. Figure 1 Various exemplary image capture devices 110 are depicted. Image capture device 110 may be, for example, or include any of a camcorder, a network video recorder, a camera, a video camera 120 (e.g., a security camera or surveillance camera), a digital camera, a wireless communication device 130 (e.g., a smartphone) including an image sensor, or an automobile 140 including an image sensor.

[0037] Figure 2a An exemplary video network system 250 is depicted in which embodiments herein may be implemented. The video network system 250 may include an image capture device, such as a video camera 120, which may capture a digital image 201 (eg, a digital video image) and perform image processing thereon. Figure 2a The video server 260 in the embodiment can obtain images from the video camera 120, for example, through a network. Figure 2a Indicated by a double-headed arrow.

[0038] Video server 260 is a computer-based device dedicated to transmitting video. Video servers are used in many applications and often have additional features and capabilities to meet the needs of specific applications. For example, video servers used in security, monitoring, and inspection applications are typically designed to capture video from one or more cameras and transmit that video over a computer network. In video production and broadcast applications, a video server may be able to record and play recorded video and transmit many video streams simultaneously. Today, many video server functions can be built into video cameras 120.

[0039] However, in Figure 2a In the embodiment, the video server 260 is connected to the image capture device, exemplified herein by the video camera 120, via the video network system 250. The video server 260 may also be connected to a video storage 270 for storing video images and / or to a monitor 280 for displaying video images. In some embodiments, the video camera 120 is directly connected to the video storage 270 and / or the monitor 280, such as Figure 2a In some other embodiments, the video camera 120 is connected to the video storage 270 and / or the monitor 280 via the video server 260, as shown by the arrows between the video server 260 and the other devices.

[0040] Figure 2b A user device 295 is depicted connected to the video camera 120 via the video network system 250. The user device 295 may be, for example, a computer or a mobile phone. The user device 295 may, for example, control the video camera 120 and / or display the video from the video camera 120. The user device 295 may also include the functionality of both the monitor 280 and the video storage 270.

[0041] To better understand the embodiments herein, an imaging system will first be described.

[0042] Figure 3 is a schematic diagram of an imaging system 300, in this case a digital camera, such as video camera 120. The imaging system images a scene onto an image sensor 301. Image sensor 301 can be equipped with a Bayer filter so that different pixels will receive radiation in a specific wavelength region in a known pattern. Typically, each pixel of a captured image is represented by one or more values ​​that represent the intensity of the captured light within a certain wavelength band. These values ​​are often referred to as color components or color channels. The term "image" can refer to an image frame or video frame that includes information from the image sensor that has captured the image.

[0043] After reading the signals of the individual sensor pixels of the image sensor 301, various image processing actions may be performed by the image signal processor 302. The image signal processor 302 may include an image processing portion 302a, sometimes referred to as an image processing pipeline, and a video post-processing portion 302b.

[0044] Typically, for video processing, images are included in an image stream. Figure 3 A first video stream 310 is shown from the image sensor 301. The first image stream 310 may include a plurality of captured image frames, such as a first captured image frame 311 and a second captured image frame 312.

[0045] Image processing may include demosaicing, color correction, noise filtering (for removing spatial and / or temporal noise), distortion correction (for removing the effects of, for example, barrel distortion), global and / or local tone mapping (for example, to enable imaging of scenes containing a wide range of intensities), transformations (for example, rectification and rotation), flat-field correction (for example, for removing the effects of vignetting), applying overlays (for example, privacy masks, explanatory text), etc. The image signal processor 302 may also be associated with an analysis engine that performs object detection, recognition, alerting, etc.

[0046] The image processing portion 302a may for example perform image stabilization, apply noise filtering, distortion correction, global and / or local tone mapping, transformations and flat field correction. The video post-processing portion 302b may for example crop portions of an image, apply overlays and include an analysis engine.

[0047] After the image signal processor 302, the image may be forwarded to the encoder 303, where the information in the image frame is encoded according to a coding protocol such as H.264. The encoded image frame is then forwarded to, for example, a receiving client (here, monitor 280 is used as an example), a video server 260, a storage 270, etc.

[0048] The video encoding process produces multiple values ​​that can be encoded to form a compressed bitstream. These values ​​may include:

[0049] Quantized transform coefficients,

[0050] Information that enables the decoder to recreate the prediction,

[0051] Information about the structure of the compressed data and the compression tools used during the encoding process, and

[0052] Information about the complete video sequence.

[0053] These values ​​and parameters (syntax elements) are converted into binary codes using, for example, variable length coding and / or arithmetic coding. Each of these coding methods produces an efficient, compact binary representation of the information, also known as a coded bitstream. The coded bitstream can then be stored and / or transmitted.

[0054] In a camera system with multiple cameras, embodiments herein propose using detection information from surrounding cameras to enhance mask quality within a camera. For example, mask quality can be enhanced by increasing the probability of detecting an object of a particular object type, even if the object is partially occluded in one of the cameras, while still ensuring a low number of objects are incorrectly detected as the object type being detected, i.e., a low number of false positive detections.

[0055] Detection information can include detection scores or detection probabilities. Different cameras have different pre-conditions for detecting people due to camera resolution, camera installation locations, and processing capabilities, which in turn affect the characteristics of the video network used with the cameras. An example of such characteristics is the size of the image input to the video network. A person that is slightly obscured in one camera's image frame may be fully visible in another camera's image frame. A person with too low a pixel density to be detected in one camera may be detectable in another camera with a higher pixel density.

[0056] The camera system may need to find spatial overlap between cameras. For example, the camera system may be presented with or determine the overlap of the areas captured by the cameras that want to share detection information. This can be done using known scene segmentation methods, using information such as geolocation, pan and tilt angles, zoom, field of view, and mounting height.

[0057] An image processing device (e.g., a camera) collects detection scores of objects in a region captured by two or more cameras from surrounding cameras. This region can be called a spatial overlap region.

[0058] • The image processing device adjusts the detection score accordingly based on the detection score from at least one other camera.

[0059] The image processing device may apply a privacy mask to the detected object based on the adjusted detection score. This will be described below, for example, with respect to Figure 5 Action 503 is explained in more detail.

[0060] Now refer to Figure 4a 、 4b , 4c, 4d and 5 and further reference Figure 1 、 2a , 2b and 3 to describe the exemplary embodiments of this invention.

[0061] Figure 4a A scene including an object, such as a person, that is obscured by another object, such as a wall, and a multi-camera system 400 are shown. Multi-camera system 400 includes at least a first camera 401 and a second camera 402. First camera 401 and second camera 402 may be video cameras. Multi-camera system 400 may also include a video server 460. First camera 401 and second camera 402 may capture the object from different angles.

[0062] In a scenario in which an embodiment can be implemented, the first camera 401 can capture a large portion of a person, while the second camera 402 can capture only a smaller portion of the person, such as the person's head and arms. Each of the first and second cameras 401, 402 can be assigned a detection score, such as a probability value indicating that the detected object belongs to a type of object, such as a human.

[0063] Figure 4b 1 shows a first image frame stream 421 captured by first camera 401 of multi-camera system 400 and a second image frame stream 422 captured by second camera 402 of camera system 400. Second camera 402 is different from first camera 401. First image frame stream 421 includes image frames 421_2 capturing a first object 431. First object 431 may include a first portion 431a. Second image frame stream 422 includes image frames 422_2 capturing a second object 432. Second object 432 may include a second portion 432a. In this scenario, first object 431 may correspond to second object 432 because they are the same real-world object that has been captured. Furthermore, first portion 431a may correspond to second portion 432a.

[0064] Figure 5 A flow chart is shown describing a method for determining a probability value indicating the probability that an object captured in the image frame stream 422 belongs to an object type, such as a particular object type. For example, the object type may be a person or an animal. In other words, the method can be used to determine the probability that the object belongs to the object type.

[0065] The method can be performed in camera system 400 and can be used more specifically to mask or count objects captured in an image frame stream. When a probability value indicating that an object captured in image frame stream 422 belongs to an object type is higher than a threshold for masking the object or a threshold for counting the objects, masking or counting the objects can be performed. The threshold for masking the object and the threshold for counting the objects can be different.

[0066] Specifically, these embodiments may be performed by an image processing device of the multi-camera system 400. The image processing device may be any one of the first camera 401 and the second camera 402 or the video server 460. The first camera 401 and the second camera 402 may each be a video camera, such as a surveillance camera.

[0067] The following actions may be taken in any suitable order (eg, in another order than that presented below).

[0068] Action 501

[0069] The method includes detecting a first object 431 or a first portion 431 a of the first object 431 in a first area 441 of a scene captured in a first stream of image frames 421 captured by a first camera 401 of the multi-camera system 400 .

[0070] Action 502

[0071] The method also includes determining a first probability value based on the characteristics of the first object 431 or the first portion 431a of the first object 431, the first probability value indicating a first probability that the detected first object 431 or the first portion 431a belongs to a type of object, such as a human. In other words, the method includes determining a first probability that the first object belongs to the type of object.

[0072] For example, the determination may be performed by weighting probability values ​​of several parts of the first object 431. A respective probability value indicating a respective probability that the detected part 431a of the first object 431 belongs to the object part type may be determined based on characteristics of the detected part 431a of the first object 431.

[0073] Action 503

[0074] When the first probability value is higher than a first threshold, it can be determined that the detected first object 431 or first portion 431a belongs to an object type, such as a human. Furthermore, when the first probability value is higher than the first threshold, further actions can be performed. For example, a privacy mask, such as a masking filter or a masking overlay, can be applied to the first object 431 or at least a portion of the first object 431 (e.g., the first portion 431a of the first object 431) in the first image frame stream 421 to mask out (e.g., conceal) the first object 431 or at least a portion of the first object 431 (e.g., the first portion 431a of the first object 431) in the first image frame stream 421. Specifically, the privacy mask can be obtained by masking an area in each image frame and can include: color masking, which can also be referred to as solid color masking or monochrome masking; mosaic masking, which can also be referred to as pixelation, pixelated privacy masking, or transparent pixelation; Gaussian blurring; a Sobel filter; a background model masking by applying a background as a mask (making the object appear transparent); and a chameleon masking (a mask that changes color depending on the background). In a detailed example, a color mask can be obtained by ignoring the image data and using another pixel value instead of the pixel to be masked. For example, the other pixel value can correspond to a specific color, such as red or gray.

[0075] In another embodiment, when the first probability value is above a first threshold, the object may be counted as a specific predefined object. Thus, the first threshold may be a masking threshold, a counting threshold, or another threshold associated with some other function performed on the object due to the first probability value being above the first threshold.

[0076] Action 504

[0077] The method also includes detecting a second object 432 or a second portion 432a of the second object 432 in a second area 442 of the scene captured in the second image frame stream 422 by a second camera 402 of the camera system 400. The second camera 402 is different from the first camera 401. The second area 442 at least partially overlaps the first area 441.

[0078] Action 505

[0079] The method also includes determining a second probability value based on the characteristics of the second object 432 or the second portion 432a of the second object 432, the second probability value indicating a second probability that the detected second object 432 or the second portion 432a belongs to the object type. In other words, the method includes determining a second probability that the second object belongs to the object type.

[0080] When the second probability value is higher than the second threshold, it can be determined that the detected second object 432 or second part 432a belongs to an object type, such as a human. In other words, the detected second object 432 or second part 432a can be detected as a human or belongs to a human.

[0081] When the second probability value is lower than the second threshold, it can be determined that the second object 432 does not belong to the object type. However, in the embodiment of this invention, when the second probability value is lower than the second threshold, the method further evaluates whether the second object 432 belongs to the object type by considering the first probability value from the first camera 401.

[0082] Action 506

[0083] Another condition for continuing to evaluate the detection of the second object 432 based on the first probability value may be determining the co-location of the first object 431 and the second object 432 to ensure that the first object 431 is the same object as the second object 432, and therefore it is appropriate to consider the first probability value when evaluating the detection of the second object 432.

[0084] Therefore, the method may further include determining that the second object 432 or the second portion 432a of the second object 432 and the first object 431 or the first portion 431a of the first object 431 are co-located in an overlapping area of ​​the second region 442 and the first region 441 .

[0085] In other words, the method may further include determining that the first object 431 and the second object 432 are the same object.

[0086] For example, it can be determined that second object 432 and first object 431 are co-located. For example, it can be determined that a first person in first image frame stream 421 and a second person in second image frame stream 422 are co-located. Based on the determination of co-location, it can be determined that the first and second persons are the same person. For example, if the probability that first object 431 is a person is determined to be 90%, and the probability that second object 432 is a person is further determined to be 70%, and no other objects are detected in the image frames, it can be determined that the persons are likely the same person.

[0087] In other words, the first and second objects 431, 432 can be determined to be linked, for example, by determining that motion trajectory entries associated with the first and second objects 431, 432 are linked (i.e., determined to belong to one real-world object) given a matching level based on a set of constraints. The trajectory entries can be created based on the received first and second image series and the determined first and second characterizing feature sets and motion data.

[0088] In another example, it may be determined that the second object 432 is co-located with the first portion 431 a of the first object 431 .

[0089] In another example, it may be determined that the second portion 432 a of the second object 432 is co-located with the first object 431 .

[0090] In another example, it can be determined that the second portion 432a of the second object 432 is co-located with the first portion of the first object 431. For example, it can be determined that the head of the first object is co-located with the head of the second object. This method can also be applied to other parts, such as arms, torsos, and legs.

[0091] Action 506 may be performed after or before action 507 .

[0092] Action 507

[0093] When the second probability value is below the second threshold and the first probability value is above the first threshold, the method includes determining an updated second probability value by increasing the second probability value.

[0094] The second threshold may also be a masking threshold, a counting threshold, or another threshold associated with some other function performed on the second object 432 or the second portion 432a due to the second probability value being above the second threshold.

[0095] The second threshold and the first threshold may be different. However, in some embodiments herein they are the same. Figure 4c A simple example is shown of how the second probability value may increase from below the second threshold to above the second threshold when the first and second thresholds are the same and the first probability value is above the first threshold.

[0096] In some embodiments herein, the first threshold is 90% and the second threshold is 70%. The second threshold can be lower than the first threshold. For example, this can be an advantage when the second object 432 is partially obscured by some other object (e.g., a wall).

[0097] In one scenario, first camera 401 determines a first probability value of 95% that a first object is a person. Second camera 402 determines a second probability value of 68% that a second object is a person. The second probability value may then be increased, for example, by 5% or 10%. This increase may be a fixed value as long as the first probability value is above a first threshold. However, in some embodiments herein, the increase in the second probability value is based on the first probability value. For example, the increase in the second probability value may be proportional to the first probability value. In some embodiments herein, when the first probability value is 60%, the second probability value may be increased by 5%, when the first probability value is 70%, the second probability value may be increased by 10%, and when the first probability value is 80%, the second probability value may be increased by 15%.

[0098] Increasing the second probability value based on the first probability value may include determining a difference between the first probability value and a first threshold value, and increasing the second probability value based on the difference. For example, if the difference between the first probability value and the first threshold value is 5%, the second probability value may be increased by 10%. In another example, if the difference between the first probability value and the first threshold value is 10%, the second probability value may be increased by 20%.

[0099] In some other embodiments, increasing the second probability value based on the first probability value includes increasing the second probability value by the difference. For example, if the difference between the first probability value and the first threshold is 5%, the second probability value may be increased by 5%. In another example, if the first probability value is 95% and a common threshold, such as a common masking threshold, is 80%, the difference is 15%, and thus, the second probability value, such as 67%, may be increased by 15%, resulting in a value of 82% above the common threshold of 80%. Thus, the second object may also be masked or counted.

[0100] In some embodiments herein, the updated second probability value is determined in response to determining that the second object 432 or portion 432a of the second object 432 is co-located with the first object 431 or portion 431a of the first object 431 in the overlapping region according to act 506 above.

[0101] In some embodiments herein, the respective first and second thresholds are specific to the type of object or object part. For example, the respective first and second thresholds can be specific to different object types or different object part types, such as a head, an arm, a torso, a leg, or both. Because masking a face may be more important than masking an arm, the masking threshold for the face may be lower than another masking threshold for the arm. For example, the masking threshold for the head may be 70%, the masking threshold for the arm may be 90%, and the threshold for the torso may be 80%.

[0102] Another condition for determining the updated second probability value by increasing the second probability value can be that the second probability value is above a third threshold in addition to being below the second threshold. Thus, having the second probability value be above the lower third threshold can be an additional condition. The reason why having the second probability value be above the lower third threshold can be advantageous is that if the second probability value is below the third threshold, for example, below 20%, the object detector is highly confident that the object is not an object that is supposed to be obscured, e.g., not a person. However, if the second probability value is between 20% and 80%, the object is likely an occluded person, so if the other camera is more confident in detecting the object as an object that is supposed to be obscured, its probability value may be increased.

[0103] In some embodiments herein, the second object part type of the second portion 432a is the same object part type as the first object part type of the first portion 431a. For example, the method can compare a first face to a second face, a first leg to a second leg, and so on.

[0104] It may be advantageous to increase the second probability value only when the difference between the first probability value and the first threshold value is above a fourth threshold value, for example, when the difference is above 10%. In this way, the number of false positives can be controlled. For example, the greater the threshold difference, the lower the number of false positives may be.

[0105] For example, in a scenario where the first threshold is 70%, the second threshold is 90%, and the fourth threshold is 75%, the first probability value of 72% is large enough to detect the first object as a human in the first image frame stream 421, but may be considered too low to increase the second probability value.

[0106] Action 508

[0107] In some embodiments herein, further actions may be performed in response to the updated second probability value being above a second threshold and in response to determining that the second object belongs to a particular object type. For example, when the object type is an object type to be masked, the first threshold and the second threshold may be used to mask an object or portion of an object of the object type to be masked. The method then further includes applying a privacy mask to at least a portion (e.g., second portion 432a) of the second object 432 in the second image frame stream 422 when the updated second probability value is above the second threshold to mask that portion. For example, if a person's face is captured in both the first and second cameras 401, 402 and the face score of the second camera 402 is updated, this may result in a privacy masking of the face in the second image stream 422. For example, the method may also include anonymizing unidentified persons by removing identifying features, such as by any of the methods mentioned above in action 503.

[0108] In some other embodiments, the object type is an object type to be counted. The first threshold and the second threshold can then be used to count objects or object portions of the object type to be counted. The method then further includes: when the updated second probability value is higher than the second threshold, increasing a counter value associated with the second image frame stream 422. For example, the counter value can be for the object or the object portion.

[0109] Figure 4dA scene is shown in which the scene includes a first person 451 and a second person 452 and a third object 453 between the two people 451, 452. The shape of the third object 453 is similar to the shape of the body parts of the people 451, 452. The third object can be, for example, a balloon similar to a head. In the example, the second camera 402 captures the three objects 451, 452, 453, and Figure 5 The method can be used for each of the three objects 451, 452, 453. Therefore, Figure 4d Each of the three objects 451, 452, 453 in may be the second object 432 in the second image frame stream 422. Correspondingly, Figure 4d Each of the three objects in may be the first object 431 in the first image frame stream 421 .

[0110] The method may be performed by the second camera or by the video server 460 , or even by the first camera 401 .

[0111] The method may then include determining a second probability value. For the first person 451 (left), the third object 453 (balloon), and the second person 452 (right), the second probability values ​​may be 69%, 70%, and 80%, respectively. In one scenario, the second threshold for detecting and masking humans is 80%. This means that if there is no Figure 5 If the second threshold is reduced by 10%, only the second person 452 will be obscured. If a 10% overall decrease in the second threshold is performed, the third object 453 will be obscured, but the first person 451 will not be obscured. However, if the first probability value derived from the first image frame stream 421 is above the first threshold and the second probability value is below the second threshold for the corresponding first object 431 (e.g., the first person 451 and the second person 452), the second probability value is increased. Therefore, if the first person 451 is detected in the first image frame stream 421 from the first camera 401 with a first probability value of 90%, indicating that the detected first person 451 is of the human type, the second probability value is increased by, for example, 15%. The updated second probability value will be 84%, which is above the second threshold, and the first person 451 will be obscured. The first probability value of the third object 453, indicating the probability that the detected third object 453 is of the human type, will be lower, for example, 18%. This means that since the first probability value of the third object 453 is below the first threshold, the second probability value will not be increased. Therefore, the third object 453 will not be obscured in any image stream from the first and second cameras 401 , 402 .

[0112] refer to Figure 6, shows a schematic block diagram of an embodiment of an image processing device 600. As described above, the image processing device 600 is configured to determine a probability value indicating that an object captured in a stream of image frames belongs to an object type. In addition, the image processing device 600 can be part of the multi-camera system 400.

[0113] As described above, the image processing device 600 may include or be any one of a camera (e.g., a surveillance camera), a monitoring camera, a video camera, a network video recorder, and the wireless communication device 130. The image processing device 600 may be the first camera 401 or the second camera 402, such as a surveillance camera, or may be a video server 460 that is part of the multi-camera system 400. The method for determining a probability value indicating that an object captured in an image frame stream belongs to a certain object type may also be performed in a distributed manner in multiple image processing devices (e.g., in the first camera 401 and the second camera 402). For example, actions 501-503 may be performed by the first camera 401, while actions 504-508 may be performed by the second camera 402.

[0114] The image processing device 600 may further include a processing module 601, such as a device for executing the method described herein, which may be implemented in the form of one or more hardware modules and / or one or more software modules.

[0115] The image processing device 600 may further include a memory 602. The memory may include instructions, for example, contain or store instructions, for example, in the form of a computer program 603, which may include computer-readable code means that, when executed on the image processing device 600, causes the image processing device 600 to perform a method for determining a probability value indicating that an object captured in a stream of image frames belongs to an object type. The image processing device 600 may include a computer, and the computer-readable code means may then be executed on the computer and cause the computer to perform the method for determining a probability value indicating that an object captured in a stream of image frames belongs to an object type.

[0116] According to some embodiments herein, the image processing device 600 and / or the processing module 601 includes a processing circuit 604 as an exemplary hardware module, which may include one or more processors. Thus, the processing module 601 may be embodied in the form of the processing circuit 604, or "implemented" by the processing circuit 604. Instructions may be executed by the processing circuit 604, whereby the image processing device 600 may be operable to perform the operations described above. Figure 5 As another example, the instructions, when executed by the image processing device 600 and / or the processing circuit 604, may cause the image processing device 600 to perform a method according to Figure 5 method.

[0117] In view of the above, in one example, an image processing apparatus 600 is provided for determining a probability value indicating that an object captured in a stream of image frames belongs to an object type.

[0118] Likewise, the memory 602 contains instructions executable by the processing circuit 604, whereby the image processing apparatus 600 is operable to perform operations according to Figure 5 method.

[0119] Figure 6 Also shown is a carrier 605 or program carrier, which includes the computer program 603 as described directly above. The carrier 605 may be one of an electric signal, an optical signal, a radio signal, and a computer-readable medium.

[0120] In some embodiments, the image processing device 600 and / or the processing module 601 may include one or more of the detection module 610, the determination module 620, the masking module 630, and the counting module 640 as exemplary hardware modules. In other examples, one or more of the aforementioned exemplary hardware modules may be implemented as one or more software modules.

[0121] Furthermore, the processing module 601 may include an input / output unit 606. According to an embodiment, the input / output unit 606 may include an image sensor configured to capture the above-mentioned raw image frames, such as the raw image frames included in the video stream 310 from the image sensor 301.

[0122] According to various embodiments described above, the image processing device 600 and / or the processing module 601 and / or the detection module 610 are configured to receive captured image frames 311 , 312 of the image stream 310 from the image sensor 301 of the image processing device 600 .

[0123] The image processing device 600 and / or the processing module 601 and / or the detection module 610 are configured to detect a first object 431 or a first part 431a of the first object 431 in a first area 441 of a scene captured in a first image frame stream 421 captured by the first camera 401 of the multi-camera system 400.

[0124] The image processing device 600 and / or the processing module 601 and / or the determination module 620 are also configured to determine a first probability value indicating a first probability that the detected first object 431 or the first part 431a belongs to the object type based on the characteristics of the first object 431 or the first part 431a of the first object 431.

[0125] The image processing device 600 and / or the processing module 601 and / or the detection module 610 are further configured to detect a second object 432 or a second portion 432a of the second object 432 in a second area 442 of a scene captured in a second image frame stream 422 captured by the second camera 402 of the camera system 400. The second camera 402 is different from the first camera 401. The second area 442 at least partially overlaps with the first area 441.

[0126] The image processing device 600 and / or the processing module 601 and / or the determination module 620 are also configured to determine a second probability value indicating a second probability that the detected second object 432 or the second part 432a belongs to the object type based on the characteristics of the second object 432 or the second part 432a of the second object 432.

[0127] The image processing apparatus 600 and / or the processing module 601 and / or the detection module 610 are further configured to determine an updated second probability value by increasing the second probability value when the second probability value is lower than the second threshold and the first probability value is higher than the first threshold.

[0128] The image processing device 600 and / or the processing module 601 and / or the determination module 620 may also be configured to increase the second probability value based on the first probability value.

[0129] The image processing device 600 and / or the processing module 601 and / or the determination module 620 may also be configured to increase the second probability value based on the first probability value by determining a difference between the first probability value and a first threshold and increasing the second probability value based on the difference.

[0130] The image processing device 600 and / or the processing module 601 and / or the masking module 630 may also be configured to apply a privacy mask to at least a portion of the second object 432 in the second image frame stream 422 when the updated second probability value is higher than a second threshold.

[0131] The image processing device 600 and / or the processing module 601 and / or the counting module 640 may further be configured to increase a counter value associated with the second image frame stream 422 when the updated second probability value is above a second threshold.

[0132] The image processing device 600 and / or the processing module 601 and / or the determination module 620 may also be configured to increase the second probability value when the second probability value is higher than a third threshold in addition to being lower than the second threshold.

[0133] The image processing device 600 and / or the processing module 601 and / or the determination module 620 can also be configured to determine that the second object 432 or the second part 432a of the second object and the first object 431 or the first part 431a of the first object 431 are co-located in the overlapping area of ​​the second area 442 and the first area 441, and determine an updated second probability value in response to determining that the second object 432 or the second part 432a of the second object 432 and the first object 431 or the first part 431a of the first object 431 are co-located in the overlapping area.

[0134] As used herein, the term "module" may refer to one or more functional modules, each of which may be implemented as one or more hardware modules and / or one or more software modules and / or a combined software / hardware module. In some examples, a module may represent a functional unit implemented as software and / or hardware.

[0135] As used herein, the term "computer program carrier," "program carrier," or "carrier" may refer to one of an electronic signal, an optical signal, a radio signal, and a computer-readable medium. In some examples, the computer program carrier may exclude transient, propagated signals such as electronic, optical, and / or radio signals. Thus, in these examples, the computer program carrier may be a non-transitory carrier, such as a non-transitory computer-readable medium.

[0136] As used herein, the term "processing module" may include one or more hardware modules, one or more software modules, or a combination thereof. Any such module, whether hardware, software, or a combined hardware-software module, may be a connection means, a provision means, a configuration means, a response means, a disabling means, etc., as disclosed herein. By way of example, the term "means" may refer to a module corresponding to the modules listed above in conjunction with the accompanying drawings.

[0137] As used herein, the term "software module" may refer to a software application, a dynamic link library (DLL), a software component, a software object, an object according to the Component Object Model (COM), a software component, a software function, a software engine, an executable binary software file, etc.

[0138] The term "processing module" or "processing circuit" herein may encompass a processing unit including, for example, one or more processors, application specific integrated circuits (ASICs), field programmable gate arrays (FPGAs), etc. The processing circuit, etc. may include one or more processor cores.

[0139] As used herein, the expression “configured to / for” may mean that the processing circuit is configured to (eg, adapted to or operable to) perform one or more of the actions described herein through software configuration and / or hardware configuration.

[0140] As used herein, the term "action" may refer to an action, step, operation, response, reaction, activity, etc. It should be noted that, if applicable, an action herein may be divided into two or more sub-actions. Furthermore, it should also be noted that two or more of the actions described herein may be combined into a single action.

[0141] As used herein, the term "memory" may refer to a hard disk, magnetic storage media, a portable computer floppy disk or disk, flash memory, random access memory (RAM), etc. Additionally, the term "memory" may refer to a processor's internal register memory, etc.

[0142] As used herein, the term "computer-readable medium" may be a universal serial bus (USB) memory, a DVD disk, a Blu-ray disc, a software module received as a data stream, a flash memory, a hard disk drive, a memory card (such as a memory stick, a multimedia card (MMC), a secure digital (SD) card, etc.) One or more of the foregoing examples of computer-readable media may be provided as one or more computer program products.

[0143] As used herein, the term "computer readable code unit" may be the text of a computer program, a binary file representing part or all of a computer program in compiled format, or anything in between.

[0144] As used herein, the terms "number" and / or "value" may refer to any type of number, such as a binary number, a real number, an imaginary number, or a rational number. Furthermore, a "number" and / or "value" may refer to one or more characters, such as a letter or a string of letters. A "number" and / or "value" may also be represented by a string of bits, i.e., a plurality of zeros and / or a plurality of ones.

[0145] As used herein, the expression "in some embodiments" has been used to indicate that features of a described embodiment can be combined with any other embodiments disclosed herein.

[0146] Although various embodiments have been described, many different changes, modifications, etc. thereof will become apparent to those skilled in the art. Therefore, the described embodiments are not intended to limit the scope of the present disclosure.

Claims

1. A method, performed in a multi-camera system (400), for determining a probability value indicating a probability that an object captured in a stream of image frames (422) belongs to an object type, the method comprising: - detecting (501) a first object (431) or a first part (431a) of the first object (431) in a first area (441) of a scene captured in a first stream of image frames (421) captured by a first camera (401) of the multi-camera system (400); - determining (502) a first probability value indicating a first probability that the detected first object (431) or first part (431a) belongs to said object type based on a characteristic of said first object (431) or said first part (431a) of said first object (431); - when the first probability value is higher than a first threshold, determining (503) that the detected first object (431) or first part (431a) belongs to the object type; - detecting (504) a second object (432) or a second part (432a) of the second object (432) in a second area (442) of the scene captured in a second stream of image frames (422) captured by a second camera (402) of the camera system (400) different from the first camera (401), wherein the second area (442) at least partially overlaps the first area (441); - determining (505) a second probability value indicating a second probability that the detected second object (432) or the second part (432a) belongs to the object type based on a characteristic of the second object (432) or the second part (432a) of the second object (432), wherein when the second probability value is above a second threshold value, it is determined that the detected second object (432) or the second part (432a) belongs to the object type; and - when the second probability value is below the second threshold value and the first probability value is above the first threshold value, determining (507) an updated second probability value by increasing the second probability value, wherein increasing the second probability value is based on the first probability value and comprises: - determining a difference between said first probability value and said first threshold value; and - increasing the second probability value based on the difference between the first probability value and the first threshold value.

2. The method according to claim 1, wherein Increasing the second probability value based on the difference between the first probability value and the first threshold value includes: The second probability value is increased by the difference.

3. The method according to any one of claims 1 to 2, wherein The second threshold is lower than the first threshold. 4 . The method according to claim 1 , further comprising determining that the second object belongs to the object type in response to the updated second probability value being higher than the second threshold.

5. The method according to any one of claims 1 to 2, wherein The object type is an object type to be masked, the first threshold and the second threshold are used to mask an object or a portion of an object of the object type to be masked, and the method further includes: when the updated second probability value is higher than the second threshold, applying (508) a privacy mask to at least a portion of the second object (432) in the second image frame stream (422).

6. The method according to any one of claims 1 to 2, wherein: The object type is an object type to be counted, the first threshold and the second threshold are used to count objects or parts of objects of the object type to be counted, and the method further includes: when the updated second probability value is higher than the second threshold, increasing the counter value associated with the second image frame stream (422).

7. The method according to any one of claims 1 to 2, wherein: The respective first and second threshold values ​​are specific to the object type or object part type.

8. The method according to any one of claims 1 to 2, wherein: Another condition for determining the updated second probability value by increasing the second probability value is that the second probability value is higher than a third threshold value in addition to being lower than the second threshold value.

9. The method according to any one of claims 1 to 2, further comprising: Determining (506) that the second object (432) or the second part (432a) of the second object and the first object (431) or the first part (431a) of the first object (431) are co-located in an overlapping area of ​​the second area (442) and the first area (441); and determining the updated second probability value in response to determining that the second object (432) or the second part (432a) of the second object (432) and the first object (431) or the first part (431a) of the first object (431) are co-located in the overlapping area.

10. The method according to any one of claims 1 to 2, wherein: The second object part type of the second part (432a) is the same object part type as the first object part type of the first part (431a).

11. An image processing device (402, 460) of a multi-camera system (400), configured to perform the method according to any one of claims 1-10.

12. The image processing device (402, 460) according to claim 11, wherein The image processing device (402, 460) is a camera (402) or a video server (460).

13. A computer program product comprising computer readable code means which, when executed on an image processing device (110), causes the image processing device (110) to perform the method according to any one of claims 1-10.

14. A computer-readable medium (605) having computer-readable code means stored thereon, which, when executed on an image processing device (110), causes the image processing device (110) to perform the method according to any one of claims 1-10.

Citation Information

Patent Citations

  • Video redaction method and system

    US20180268240A1

  • Optimized object detection

    CN108027884A

  • Systems and methods for video surveillance

    CN113243015A