Method and image processing device for detecting reflections of an identified object in an image frame
Patent Information
- Application Number
- CN202410373677.3
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Priority Date
- 2023-03-30
- Filing Date
- 2024-03-29
- Publication Date
- 2026-10-09
- Estimated Expiration
- 2044-03-29
AI Technical Summary
可能出现的一个问题是,场景中的反射表面可能将隐私掩模后面的东西反射到捕获场景的相机中
[0022] The embodiments described herein find candidate reflective pixels in an image frame by tracing rays from an object to the camera in a 3D model of the scene captured in the image frame, and confirm the detection of reflective pixels by comparing the color values of the object's pixels with those of the reflective surface's pixels. The actual data values of the reflective pixels are also used to achieve unbiased results.
Smart Images

Figure CN118736089B_ABST
Abstract
Description
Technical Field
[0001] The embodiments described herein relate to a method and image processing apparatus for detecting reflections of identified objects in an image frame. Corresponding computer programs and computer program carriers are also disclosed. Background Technology
[0002] The use of imaging, particularly video imaging, to monitor the public is common in many parts of the world. Areas that may need to be monitored include, for example, banks, shops, and other areas requiring security, such as schools and government facilities. Other areas that may need to be monitored are processing, manufacturing, and logistics applications, where video surveillance is primarily used to monitor processes.
[0003] However, there may be requirements that make it impossible to identify people from video surveillance. This inability to identify people may conflict with the requirement to be able to determine what is happening in the video. For example, performing people counting or queue monitoring on anonymized image data might be of interest. In practice, there is a trade-off between satisfying these two requirements: unidentifiable video and extracting large amounts of data for different purposes, such as people counting.
[0004] Several image processing techniques have been described to avoid identifying people while still allowing for the identification of activity. Examples of such manipulation include edge detection / representation, edge enhancement, silhouette objects, and different types of "color blurring" (such as color shifting or dilation). Privacy masks are another image processing technique used in video surveillance to protect personal privacy by hiding portions of an image with masked areas.
[0005] Image processing refers to any processing applied to an image. Processing can include applying various effects, masks, filters, etc., to an image. In this way, an image can be sharpened, converted to grayscale, or altered in some other way, for example. Images are typically captured by video cameras, still image cameras, etc.
[0006] As mentioned above, one way to avoid identifying people is by masking moving people and objects in a live image. Masking in live and recorded video is done by comparing the live camera view to the set background scene and applying a dynamic mask to areas of change (basically moving people and objects). Color masks, also known as solid color masks or monochrome masks, mask objects with a solid color overlay, providing privacy while allowing you to see motion. Mosaic masks, also called pixelated, pixelated privacy masks, or transparent pixelated masks, show moving objects at a lower resolution and allow you to better distinguish forms by seeing the object's color.
[0007] In areas where surveillance is problematic due to privacy rules and regulations, masked live and recorded video is suitable for remote video surveillance or recording. When video surveillance is primarily used to monitor processes, it is ideal for processing, manufacturing, and logistics applications. Other potential applications include retail, education, and government facilities.
[0008] Despite the continuous development of masking technology, there is still room for improvement. One potential problem is that reflective surfaces in the scene might reflect things behind the privacy mask into the camera capturing the scene.
[0009] This problem is particularly challenging for dynamic masks that are expected to move with objects. Document CN 108 090947A discloses a ray tracing optimization method for 3D scenes. Summary of the Invention
[0010] Therefore, the purpose of the embodiments described herein may be to avoid some of the problems mentioned above, or at least to reduce their impact. Specifically, the purpose of the embodiments described herein may be to identify pixels in an image that represent surfaces in a scene that reflect objects into the camera, making it possible to apply image processing to pixels representing those reflective surfaces. For example, reflections may also be masked to provide improved anonymization.
[0011] The embodiments described herein address the aforementioned problem by creating a 3D representation of the scene (including identified mask objects) captured by image frames from a camera and tracing light rays from the identified objects to the camera within the 3D representation of the scene via reflective surfaces in the scene. However, only those reflective surfaces that are sufficiently similar to the actual objects themselves will be detected as reflective objects. Specifically, only those reflective surfaces that produce reflections with color values matching the color values of the reflected object will be detected. Color value comparison is performed by mixing the color values of pixels from the image frame representing the object, pixels from the reflective surface, and pixels from the background image frame representing the reflective surface, without the influence of the object.
[0012] According to one aspect, this objective is achieved by a method performed by an image processing device for detecting the reflection of an identified object in an image frame captured by a camera. The method includes generating a three-dimensional model of the background scene of the image frame based on obtained three-dimensional information about the background scene.
[0013] The method also includes: defining the identified object in the image frame based on the image information in the image frame.
[0014] The method also includes defining the three-dimensional bounding box of the defined object in the three-dimensional model of the background scene.
[0015] The method also includes defining surface elements of a 3D bounding box, wherein the corresponding surface elements are defined by the center coordinates and color values in the 3D model of the background scene.
[0016] The method further includes: determining the three-dimensional coordinates of a surface in a three-dimensional model of a background scene that reflects light from surface elements of a three-dimensional bounding box of an object into a camera, wherein the determination is performed by tracing the light rays starting from the center coordinates of the surface elements of the object's three-dimensional bounding box and based on the normals of the surface in the three-dimensional model of the background scene at said three-dimensional coordinates.
[0017] The method further includes: identifying a first pixel in the image frame that corresponds to the three-dimensional coordinates of the determined surface.
[0018] The method further includes detecting the reflection of the object when the mixture of the first color value of the identified first pixel, the color value of the surface element of the object, and the ground truth color value of the identified first pixel satisfies the mixing criterion.
[0019] The true data represents a reflective surface without the influence of an object. The true data color value can be obtained from one or more background image frames or from one or more neighboring pixels of the identified first pixel within an image frame.
[0020] Alternatively, this objective can be achieved by an image processing device configured to perform the methods described above.
[0021] According to a further aspect, this objective is achieved by a computer program and a computer program carrier corresponding to the foregoing aspects. Although embodiments have been outlined above, the claimed subject matter is defined by the appended claims 1 to 14.
[0022] The embodiments described herein find candidate reflective pixels in an image frame by tracing rays from an object to the camera in a 3D model of the scene captured in the image frame, and confirm the detection of reflective pixels by comparing the color values of the object's pixels with those of the reflective surface's pixels. The actual data values of the reflective pixels are also used to achieve unbiased results.
[0023] Therefore, the image processing device will only detect reflective pixels that are sufficiently similar to the object. Attached Figure Description
[0024] Aspects of the embodiments disclosed herein, including their particular features and advantages, will be readily understood from the following detailed description and accompanying drawings, wherein:
[0025] Figure 1 An exemplary embodiment of the image capture device is shown.
[0026] Figure 2An exemplary embodiment of a video network system is shown.
[0027] Figure 3 This is a schematic block diagram illustrating an exemplary embodiment of the imaging system.
[0028] Figure 4 This is a schematic block diagram illustrating the image frame stream and its contents.
[0029] Figure 5 This is a schematic block diagram illustrating an embodiment of a method in an image processing apparatus.
[0030] Figure 6 This is a schematic block diagram illustrating a reference method for masking objects in an image frame. Figure 7a This is a schematic block diagram illustrating a first scenario in which the embodiments disclosed herein can be implemented.
[0031] Figure 7b This is a schematic block diagram illustrating a second scenario in which the embodiments disclosed herein can be implemented.
[0032] Figure 8 This is a schematic block diagram illustrating an embodiment of a method in an image processing apparatus.
[0033] Figure 9a This is a schematic block diagram illustrating an embodiment of a method in an image processing apparatus.
[0034] Figure 9b This is a schematic block diagram illustrating an embodiment of a method in an image processing apparatus.
[0035] Figure 10 This is a flowchart illustrating an embodiment of a method in an image processing device.
[0036] Figure 11 This is a flowchart illustrating an embodiment of a method in an image processing device.
[0037] Figure 12 This is a schematic block diagram illustrating a further embodiment of a method in an image processing apparatus.
[0038] Figure 13 This is a schematic block diagram illustrating a further embodiment of a method in an image processing apparatus.
[0039] Figure 14 This is a schematic block diagram illustrating an embodiment of an image processing device. Detailed Implementation
[0040] The embodiments disclosed herein relate to improving the detection of pixels representing reflections of objects detected in image frames (such as image frames in a video stream).
[0041] Specifically, the embodiments disclosed herein relate to improved anonymization of image frames.
[0042] Therefore, the embodiments described herein can be implemented in an image processing device. In some embodiments herein, the image processing device may include, or be, an image capture device such as a digital camera. Figure 1 Various exemplary image capture devices 110 are depicted. Image capture device 110 may be, for example, or include any of the following: a camera, a network video recorder, a video camera 120 such as a surveillance camera or monitoring camera, a digital camera, a wireless communication device 130 such as a smartphone that includes an image sensor, or a car 140 that includes an image sensor.
[0043] Figure 2 An exemplary video network system 250 is depicted, in which the embodiments described herein may be implemented. The video network system 250 may include an image capture device, such as a video camera 120, which may capture digital images 201, such as digital video frames, and perform image processing thereon. Figure 2 The video server 260 can obtain video frames from the video camera 120, for example, via a network. Figure 2 The image processing device may be indicated by a double-pointing arrow. In some embodiments herein, the image processing device may include or be a video server 260.
[0044] The video server 260 is a computer-based device specifically designed for delivering video.
[0045] However, in Figure 2 In this embodiment, video server 260 is connected to an image capture device, exemplified herein by video camera 120, via video network system 250. Video server 260 may also be connected to video storage 270 for storing video frames, and / or to monitor 280 for displaying video frames. In some embodiments, video camera 120 is directly connected to video storage 270 and / or monitor 280, such as... Figure 2 The direct arrows between these devices indicate this. In some other embodiments, the video camera 120 is connected to the video storage 270 and / or the monitor 280 via the video server 260, as indicated by the arrows between the video server 260 and other devices.
[0046] Figure 3This is a schematic diagram of an imaging system 300, in this case a digital video camera, such as video camera 120. The imaging system 300 images a scene on an image sensor 301. The image sensor 301 may be equipped with a Bayer filter, such that different pixels will receive radiation in a specific wavelength region in a known pattern. Typically, each pixel of the captured image is represented by one or more values that represent the intensity of light captured within a specific wavelength band. These values are often referred to as color components or color channels. The term "image" can refer to an image frame or video frame that includes information derived from the image sensor that has captured the image.
[0047] After the signals of each sensor pixel of the image sensor 301 have been read, different image processing actions can be performed by the image signal processor 302. The image signal processor 302 may include an image processing section 302a (sometimes referred to as the image processing pipeline) and a video post-processing section 302b.
[0048] Typically, for video processing, images are included in an image stream (also known as a video frame stream). Figure 3 A video stream 310 from an image sensor 301 is shown. The video stream 310 may include multiple captured video frames, such as a first captured video frame 311 and a second captured video frame 312. The image stream may also be referred to as a video sequence.
[0049] Image processing may include applying overlays (e.g., privacy masks, explanatory text). The image signal processor 302 may also be associated with an analysis engine that performs object detection, recognition, alerts, etc.
[0050] Image processing unit 302a may, for example, perform image stabilization, apply noise filtering, distortion correction, global and / or local tone mapping, transform, and flat field correction. Video post-processing unit 302b may, for example, crop portions of the image, apply overlays, and includes an analysis engine. Therefore, the embodiments disclosed herein can be implemented by video post-processing unit 302b.
[0051] Following the image signal processor 302, the image can be forwarded to the encoder 303, where information in the video frames is encoded according to an encoding protocol such as H.264. The encoded video frames are then forwarded to, for example, receiving clients, such as monitor 280, video server 260, storage 270, etc.
[0052] As mentioned above, the purpose of the embodiments described herein may be to improve the detection of pixels representing reflections of objects detected in an image frame.
[0053] Figure 4 It shows the corresponding Figure 3The video stream 310 includes a video sequence 400. The video sequence 400 comprises multiple image frames. For example, the video sequence 400 includes image frame 402, which may also be referred to as the first image frame 402. The video sequence 400 also includes a second image frame 403. The video sequence 400 may also include a background image frame 401b. In some embodiments herein, the background image frame 401b may be derived from the multiple image frames of the video sequence 400. The background image frame 401b is not necessarily part of the video sequence 400. Instead, it may be stored on an image processing device and retrieved when needed. The background image frame 401b can be derived from the multiple image frames of the video sequence 400 by averaging the multiple image frames.
[0054] Figure 4 The lower part shows the contents of background image frame 401b, image frame 402, and second image frame 403. The content of background image frame 401b may include background objects such as a house. Preferably, the background objects of background image frame 401b are stationary, such that they also exist at the same pixel positions in other image frames. The contents of image frames 402 and second image frames 403 may include background objects and foreground objects. Foreground objects may move in the scene and therefore may be represented by different pixels in different image frames, such as... Figure 4 As shown by the people moving in China.
[0055] Figure 5 A simplified method for obtaining a background image frame 401b from an image frame (such as image frame 401) is illustrated. The image frame includes a scene with different objects, which can be classified as background objects and foreground objects according to some criteria. Background image frame 401b can be generated from image frame 401 and may include the background objects of image frame 401. Correspondingly, foreground image frame 401f can be generated and may include the foreground objects of image frame 401.
[0056] To understand the advantages of the embodiments disclosed herein, the reference method will first be described. Figure 6 This is a schematic block diagram illustrating a reference method for masking objects in an image frame. More specifically, Figure 6 A portion of an image frame captured by video camera 120 is shown.
[0057] Video camera 120 captures a scene with background and foreground objects. Specifically, object 410 is captured and detected as a foreground object. If object 410 is detected as a person, it can be masked to anonymize that person. As mentioned above, one way to avoid identifying people is by masking moving people and objects in a live image. Masking in live and recorded video can be done by comparing the live camera view with the set background scene and applying dynamic masks to areas of change (basically moving people and objects). Color masks, also known as solid color masks or monochrome masks, where objects are masked by a solid color mask overlaid with a certain color, provide privacy while allowing you to see movement. Mosaic masks, also known as pixelated, pixelated privacy masks, or transparent pixelated masks, show moving objects at a lower resolution and allow you to better distinguish forms by seeing the object's color.
[0058] However, one potential problem is that reflective surfaces in the scene, such as... Figure 6 The window shown may reflect what is behind the privacy mask onto the video camera 120 that captures the scene.
[0059] Therefore, the purpose of the embodiments herein may be to identify pixels in an image frame that represent surfaces in the scene that reflect object 410 onto video camera 120, making it possible to apply image processing to pixels representing those reflective surfaces. For example, reflections may also be masked to provide improved anonymization.
[0060] Now refer to Figure 7a , Figure 7b , Figure 8 , Figure 9a and Figure 9b and further reference Figure 2 , Figure 3 and Figure 4 An exemplary embodiment for detecting the reflection of an identified object 410 in an image frame 402 captured by camera 120 is described.
[0061] In a scenario where the embodiments described herein can be implemented, video camera 120 captures video sequence 400. Video sequence 400 captures a scene including background and foreground objects. Figure 7a In the case shown, a portion of image frames from video sequence 400 is illustrated, where one of the background objects is a house and one of the foreground objects is a person. The detected person will be masked. However, the background house includes reflective surfaces, such as windows, which reflect the person into the video camera 120. Therefore, the image of the person exists within the reflective surfaces. The background may, for example, include a first reflective surface 411 and a second reflective surface 412. Typically, the reflective surfaces are part of the background objects.
[0062] Image frame 402 includes pixels representing an image of object 410, which is detected as a person, i.e., the object to be masked. However, image frame 402 also includes pixels representing images of the first reflective surface 411 and the second reflective surface 412.
[0063] Figure 7b This indicates that the image was captured using the second image frame 403. Figure 7a The scene is where the second image frame 403 is slightly after the first image frame 402. Figure 7b In this process, object 410 has moved slightly relative to the background. Furthermore, the image of object 410 in the reflective surface 411 has been moved.
[0064] Figure 8 The contents of the first image frame 402 and the second image frame 403 are shown. Image processing devices 120 and 260 can define the identified object 410 in image frame 402 using a two-dimensional bounding box 413. Figure 8 For simplicity, object 410 is shown as a rectangle. A two-dimensional bounding box 413 can be defined by its coordinates in image frame 402. Furthermore, the surface of the two-dimensional bounding box 413 can be divided into surface elements 414. Each surface element 414 of the two-dimensional bounding box 413 can be defined by its two-dimensional coordinates in image frame 402 and a color value such as a hue value. The color value can be obtained by calculating the average color value of the pixels within the surface element 414.
[0065] The embodiments described herein are based on finding a reflective surface by tracing rays in a 3D model of the scene, which could potentially reflect an object into the camera 120. To do this, the embodiments herein define a 3D bounding box 415 of an object 410 defined in a 3D model of the background scene, for example, based on the position of the object 410 in the obtained 3D model, as defined in image frame 402, and by extrapolating the object 410 defined in a plane extending along the normal to the image plane (i.e., along the depth plane of the image).
[0066] Figure 9a A simplified 3D model of the scene is shown, including a 3D bounding box 415, a camera 120, and a reflective surface 411 in the background. For simplicity, the reflective surface 411 in the background is shown... Figure 9a It is drawn as a two-dimensional piece.
[0067] The two-dimensional bounding box 413 and its surface element 414 can be used to extrapolate the defined object 410 in the depth plane. For example, a three-dimensional bounding box 415 can be generated by extrapolating the two-dimensional bounding box 413. The surface of the three-dimensional bounding box 415 can be generated based on the surface element 414 of the two-dimensional bounding box 413.
[0068] Figure 9b A portion of the background model, including the first reflective surface 411, is schematically shown. Figure 9b Further schematically illustrating how light from the rear surface 416 of the three-dimensional bounding box 415 is reflected from the first reflective surface 411 into the camera 120. Figure 9b A portion of the background image frame 401b overlaid on the background model is further illustrated schematically, corresponding to at least a portion of the first reflective surface 411. The rear surface 416 of the 3D bounding box 415 may include surface elements 417 of the 3D bounding box 415, which are copies of surface elements 414 of the 2D bounding box 413. The side surfaces of the 3D bounding box 415 may be generated as copies of the boundary surface elements 414 of the 2D bounding box 413. The colors of the copied surface elements may be preserved.
[0069] Figure 10 A flowchart is shown describing a method performed in an image processing device for detecting the reflection of an identified object 410 in an image frame 402 captured by camera 120.
[0070] This method can be performed by an image processing device, such as a video camera 120 or a video server 260.
[0071] The following actions can be taken in any suitable order, for example, in an order different from that presented below.
[0072] Action 1001
[0073] Background image frame 401b can be generated according to known methods. Background objects can be identified and defined in background image frame 401b.
[0074] Preferably, the background image frame 401b does not include the foreground object 410.
[0075] Action 1002
[0076] Image frame 402 can be obtained. Object 410 in image frame 402 can be identified from the image information in image frame 402. For example, object identification algorithms can be used to find objects in a scene.
[0077] Action 1003
[0078] Based on the obtained 3D information about the background scene, a 3D model of the background scene of image frame 401 is generated. The background scene can be, for example, a room.
[0079] The 3D model of the background scene includes spatial information about the background objects, such as the position, size, and orientation of their surfaces. For example, the 3D model of the background scene may include information about the directions of the normals to the surfaces of the background objects. Normals can be used for ray tracing.
[0080] The 3D information about the background scene may include the 3D coordinates of surfaces in the background scene, the corresponding surface normals, the 3D coordinates of camera 120, and the orientation of camera 120. For example, the 3D model of the background scene may include information about the position of camera 120 relative to the background objects.
[0081] It is possible to generate 3D information about the background scene from light detection and ranging (LIDAR).
[0082] In some other embodiments, three-dimensional information about the background scene can be obtained by running a neural network with background image frame 401b as input data.
[0083] A 3D model of the background scene is linked to a background image frame 401b. For example, there is a mapping between the 3D coordinates of the surface of the background model and the corresponding pixels of the background image frame 401b. Mapping image data over, for example, LiDAR data is known in the art.
[0084] When updating background image frame 401b, the 3D model of the background scene can be updated.
[0085] In some embodiments described herein, the 3D model of the background scene is modified by removing surfaces of background objects whose normals point towards camera 120, since reflected light from these surfaces is less likely to reach the camera.
[0086] Action 1004a
[0087] Image processing devices 120 and 260 define the identified object 410 in image frame 402 based on the image information in image frame 402.
[0088] For example, image processing devices 120 and 260 can define the identified object 410 in image frame 402 using a two-dimensional bounding box 413. The position of the identified object 410 can be defined in image frame 402.
[0089] Action 1004b
[0090] Image processing devices 120 and 260 define a 3D bounding box 415 for the defined object 410 in a 3D model of the background scene. The 3D bounding box 415 can be a cuboid. The bounding box can contain any shape, which simplifies calculations. Any shape can fit inside the bounding box 415. In some embodiments herein, the 3D bounding box 415 includes one or more voxels. The 3D bounding box 415 can be based on the object 410 defined in image frame 402, the position of the object 410 obtained in the 3D model, and the object 410 defined by extrapolation in a plane extending along the normal to the image plane. The plane extending along the normal to the image plane can also be referred to as the depth plane of the image.
[0091] Image processing devices 120 and 260 can extrapolate object 410 by extrapolating the boundary pixels of object 410 from captured image frame 402.
[0092] In some example embodiments, the pixels of object 410 are extrapolated a certain distance in the plane normal to the image plane. For example, a car has a fairly uniform length, which can be used to perform the extrapolation.
[0093] The rear surface 416 of the 3D bounding box 415 may include pixels that are copies of the pixels of the object 410 of the image frame 402, i.e., pixels within the 2D bounding box 413. The side surfaces of the 3D bounding box 415 may include pixels that are copies of the pixels of the boundary pixels of the object 410 of the image frame 402. The colors of the copied pixels are preserved.
[0094] The position of object 410 in the 3D model can be obtained through 3D detection of object 410, such as artificial intelligence (AI) based on depth estimation based on LIDAR, RADAR, or image information in image frame 402.
[0095] In some other embodiments, the 3D bounding box 415 is directly based on the 3D mapping of the object 410, such as based on LiDAR or depth estimation artificial intelligence (AI).
[0096] Action 1005
[0097] Image processing devices 120 and 260 define surface elements 417 of the three-dimensional bounding box 415. Smaller surface elements 417 mean better accuracy but poorer performance.
[0098] The corresponding surface element 417 is defined by the center coordinates 418 and color value in the 3D model of the background scene.
[0099] Color values can be hue values in YUV format or a combination of Cb and Cr values. Other color values are also possible.
[0100] Hue is one of the main characteristics of the color appearance parameter, referred to as color. In the CIECAM02 model, it is technically defined as, within certain color vision theories, "the degree to which a stimulus can be described as similar to or different from a stimulus described as red, orange, yellow, green, blue, or purple." Hue can usually be quantitatively represented by a single number, which typically corresponds to an angular position on a color space coordinate diagram (such as a chromaticity diagram or color wheel) around a center or neutral point or axis.
[0101] The corresponding surface element 417 of the 3D bounding box 415 corresponds to a plurality of pixels from the captured image frame 402. The color value of the corresponding surface element 417 is calculated as the average of the color values of the corresponding plurality of pixels.
[0102] The color value of surface element 417 can be defined by averaging the color values of pixels within another surface element from which surface element 417 is derived. For example, surface element 417 can be on the back side and can be derived from another surface element on the front side, which is derived from a set of pixels in image frame 402.
[0103] Action 1006
[0104] Image processing devices 120 and 260 determine the three-dimensional coordinates of surface 411 in a three-dimensional model of the background scene, which reflects light from surface elements 417 of the three-dimensional bounding box 415 of object 410 into camera 120. That is, the three-dimensional coordinates of surface 411 determined in the three-dimensional model of the background scene are positioned such that light from surface elements 417 and reflected by surface 411 at those three-dimensional coordinates will be captured by camera 120. However, the actual lighting conditions of the scene can determine whether it is possible to detect the reflection of object 410 in pixels of image frame 402 corresponding to the three-dimensional coordinates of surface 411. Therefore, action 1006 concerns finding candidate reflection coordinates of surface 411.
[0105] It can store the three-dimensional coordinates of surface 411 determined in the three-dimensional model of the background scene.
[0106] The determination is performed by tracing light rays starting from the center coordinates 418 of the surface element 417 of the three-dimensional bounding box 415 of object 410 and based on the surface normals of the three-dimensional model of the background scene at said three-dimensional coordinates.
[0107] Ray tracing can be repeated for multiple surface elements 417. For example, ray tracing can be performed from all surface elements of the 3D bounding box 415. Then, all the corresponding 3D coordinates of the surfaces 411 in the 3D model of the background scene that reflect light from the surface elements 417 can be found.
[0108] In some other embodiments, ray tracing is performed only on the surface elements 417 that define the 3D bounding box 415. In this way, the contour onto which the 3D bounding box 415 is projected onto the surface 411 can be found and later used to determine which pixels of the image frame 402 should be masked. This latter option requires less computation.
[0109] Ray tracing can be performed using known methods. Ray tracing is a method for calculating the path of a wave or particle through regions of surfaces with different propagation speeds, absorption properties, and reflective properties. In these cases, the wavefront may bend, change direction, or be reflected from the surface. Ray tracing solves this problem by repeatedly advancing a discrete amount of an idealized, narrow beam, called a ray, through a medium.
[0110] When applied to problems involving electromagnetic radiation (such as light), ray tracing typically relies on approximate solutions to Maxwell's equations, which are valid as long as the light wave propagates through and around an object whose size is much larger than the wavelength of the light.
[0111] Ray tracing works by assuming that particles or waves can be modeled as a large number of very narrow beams (light rays) at a distance, perhaps very small, such that the light rays are locally straight. The ray tracer can advance the light rays over this distance and then use the local derivative of the medium to calculate the new direction of the light rays. From that position, a new ray is sent out, and the process is repeated until a complete path is generated. If the simulation includes solid objects, the ray can be tested for intersections with them at each step, and if a collision is detected, the direction of the ray can be adjusted.
[0112] Action 1007
[0113] Image processing devices 120 and 260 identify a first pixel 431 in image frame 402, which corresponds to the three-dimensional coordinates of the determined surface 411. The identified first pixel may be a candidate reflective pixel.
[0114] Image processing devices 120 and 260 can identify a plurality of first pixels 431 in image frame 402, the plurality of first pixels 431 corresponding to a plurality of defined three-dimensional coordinates of surface 411.
[0115] Action 1008
[0116] When the mixture of the first color value of the identified first pixel 431, the color value of the surface element 417 of the object 410, and the real data color value of the identified first pixel 431 satisfies the mixing criterion, the image processing devices 120 and 260 detect the reflection of the object 410 in the image frame 402. That is, if the mixture of the first color value of the identified first pixel 431, the color value of the surface element 417 of the object 410, and the real data color value of the identified first pixel 431 satisfies the mixing criterion, then the identified first pixel 431 is detected as a pixel reflecting the object. The mixing criterion can be, for example, the color value of the identified first pixel 431 in the background image frame 401b is roughly the sum of the color value of the surface element 417 of the object 410 and the real data color value. This can be checked by subtracting the real data color value from the color value of the identified first pixel 431 and comparing the resulting value with the color value of the surface element 417 of the object 410.
[0117] The true data color value is the color value of the first pixel 431 when there is no reflection from the foreground object.
[0118] The true data color value can be obtained from one or more background image frames 401b or from one or more neighboring pixels of the first pixel 431 identified in image frame 402.
[0119] The actual data color values from one or more background image frames 401b can be values taken over time to produce an average value. The actual data values can be stored in memory.
[0120] The real data can be a heatmap built up over time. If some of the background image frames 401b in one or more background image frames 401b include object 410, then one or more background image frames 401b can be averaged to not contain any significant reflections of object 410.
[0121] Action 1008 can be repeated for multiple identified first pixels 431.
[0122] Figure 11 Another flowchart is shown, illustrating another method performed in an image processing device for detecting the reflection of an identified object 410 in an image frame 402 captured by camera 120. Figure 11 The method is also used when the mask has been determined to represent the pixel region of the reflection of object 410.
[0123] Figure 11 The method can be performed by an image processing device, such as a video camera 120 or a video server 260.
[0124] The following actions can be taken in any suitable order, for example, in an order different from that presented below.
[0125] Action 1109
[0126] In response to the detection of reflection from object 410, the image processing device applies a mask to pixel region 441 of image frame 402. Pixel region 441 includes the identified first pixel 431. Pixel region 441 is shown in Figure 9B. As mentioned above, regarding action 1008, the image processing devices 120 and 260 detect the reflection of object 410 in image frame 402 when the mixture of the first color value of the identified first pixel 431, the color value of the surface element 417 of object 410, and the real data color value of the identified first pixel 431 satisfies a mixing criterion. For example, if the color value of the identified first pixel 431 of the background image frame 401b is roughly the sum of the color value of the surface element 417 of object 410 and the real data color value, then the identified first pixel 431 and surrounding pixels of pixel region 441 can be masked.
[0127] Figure 12 A portion of the background model, including the first reflective surface 411, is schematically shown. Figure 12 A portion of the background image frame 401b overlaid on the background model is further schematically shown, corresponding to at least a portion of the first reflective surface 411. In some embodiments herein, the masked pixel region 441 includes all pixels of the projection 421 of the surface element 417 corresponding to the 3D bounding box 415 onto the surface 411 or the second surface 412 of the 3D model of the background scene. In some embodiments herein, the masked pixel region 441 is limited to the pixels of the projection 421 of the surface element 417 corresponding to the 3D bounding box 415 onto the surface 411 or the second surface 412. The projection is reflected into the camera 120. A magnified view of the projection 421 is also shown... Figure 12 As shown, a magnified view of projection 421 has been shown as pixels filling image frame 402, which correspond to the portion of surface 411 covered by projection 421. Therefore, the masked pixel region 441 may include pixels of image frame 402, which correspond to the portion of surface 411 covered by projection 421.
[0128] Now refer to Figure 11 and Figure 13 A method for verifying the detection of reflections of an identified object 410 in an image frame 402 is described, which can be performed by analyzing whether the detected reflections move between a first image frame 402 and at least one other image frame (such as a second image frame 403).
[0129] When movement of object 410 has been detected based on second image frame 403 and image frame 402, the following method can be triggered.
[0130] If movement of object 410 and movement of reflection are detected, image processing devices 120 and 260 can increase the probability value of the reflection of object 410 that has been detected.
[0131] Action 1110
[0132] In some embodiments herein, image processing devices 120, 260 obtain a second image frame 403 of a video sequence 400 including image frame 402. The second image frame 403 includes an identified object 410. The second image frame 403 may be an image frame immediately following the first image frame 402 in the video sequence 400.
[0133] Action 1111
[0134] Image processing devices 120 and 260 define the identified object 410 in the second image frame 403 based on the image information in the second image frame 403.
[0135] Action 1112
[0136] Image processing devices 120 and 260 determine the corresponding second center coordinates 420 of the surface elements 417 of the 3D bounding box 415 of the object 410 in the 3D model of the background scene based on the obtained second position of the object 410 in the 3D model. The second position of the object 410 in the 3D model can be obtained in the same manner as described above with respect to action 1004b. The second position of the object 410 in the 3D model can correspond to the position of the object 410 in the second image frame 403.
[0137] Action 1113
[0138] Image processing devices 120 and 260 determine the second three-dimensional coordinates of surface 411 or second surface 412 in the three-dimensional model of the background scene. Surface 411 or second surface 412 reflects light from the surface element 417 of the three-dimensional bounding box 415 of object 410 into camera 120, and the second three-dimensional coordinates are different from the three-dimensional coordinates of the determined surface 411.
[0139] The determination is performed by tracing the light rays from the second center coordinates 420 of the surface element 417 and based on the second normal of the surface 411 or the second surface 412 in the 3D model of the background scene at the second 3D coordinates.
[0140] Action 1114
[0141] Image processing devices 120 and 260 identify a second pixel 432 in a second image frame 403, the second pixel corresponding to a second three-dimensional coordinate of a determined surface or second surface 412.
[0142] Action 1115
[0143] Image processing devices 120 and 260 obtain the second color value of the second pixel 432. The second color value of the second pixel 432 can be stored.
[0144] Action 1116a
[0145] When the mixture of the first color value and the second color value meets the second mixing criterion, the image processing devices 120 and 260 can confirm the detection of the reflection of the object 410.
[0146] For example, for the next frame 403, examine the saved 3D position from the previous frame 402 and compare the saved color value with its current color value. If there is a color change that roughly corresponds to that of object 410, it strongly indicates that the reflection is moving. If this conclusion is reached, it indicates that the area is reflective, which may increase the probability of adding a mask.
[0147] Action 1116a
[0148] In some other embodiments herein, image processing devices 120 and 260 reject the detection of reflection from object 410 when the mixture of the first color value and the second color value does not meet a second mixing criterion. The color value of the first pixel 431 in the second image frame 403 is equal to the first color value.
[0149] Therefore, if the second color value is different from the first color value, but the color value is the same for the pixel region analyzed in the previous frame 402, this indicates a false positive.
[0150] This means the object is moving, but there is no reflection. If the object is moving but there is no reflection, this indicates a false positive, and the removal of the reflection mask in that area should be considered.
[0151] Action 1117
[0152] In response to confirmation of the detection of a reflection from object 410, image processing devices 120, 260 may apply a mask to a second pixel region 442 of a second image frame 403. The second pixel region 442 includes the identified second pixel 432. In some embodiments herein, the masked second pixel region 442 is limited to pixels corresponding to the second projection of the surface element 417 of the three-dimensional bounding box 415 onto surface 411 or the second surface 412.
[0153] Action 1118
[0154] In response to the detection of reflection of the rejected object 410, the image processing devices 120 and 260 may determine that the mask will not be applied to the second pixel region 442 including the identified second pixel 432.
[0155] refer to Figure 14 A schematic block diagram of an embodiment of the image processing device 600 is shown. The image processing device 600 corresponds to... Figure 1 or Figure 2 The image processing device 600 can be any of the following image processing devices: 120, such as a surveillance camera, monitoring camera, video camera, network video recorder, and wireless communication device 130. Specifically, the image processing device 600 can be camera 120, such as a surveillance video camera, or video server 260.
[0156] As mentioned above, the image processing device 600 is configured to perform according to Figure 10 or Figure 11 The method.
[0157] The image processing apparatus 600 may also include a processing module 601, such as means for performing the methods described herein. The means may be implemented as one or more hardware modules and / or one or more software modules.
[0158] The image processing apparatus 600 may also include a memory 602. The memory may include instructions, for example, in the form of a computer program 603, such as including or storing the instructions, which may include computer-readable code units that, when executed on the image processing apparatus 600, cause the image processing apparatus 600 to perform the methods described above, such as regarding... Figure 10 or Figure 11 .
[0159] The image processing device 600 may include a computer, and then computer-readable code units can be executed on the computer, causing the computer to perform... Figure 10 or Figure 11 The method.
[0160] According to some embodiments herein, image processing device 600 and / or processing module 601 include processing circuitry 604 as an exemplary hardware module, which may include one or more processors. Therefore, processing module 601 may be implemented as, or may be “implemented” by, processing circuitry 604. Instructions may be executed by processing circuitry 604, thereby enabling image processing device 600 to perform as described above. Figure 10 or Figure 11The method. As another example, when executed by the image processing device 600 and / or the processing circuit 604, the instructions can cause the image processing device 600 to perform according to Figure 10 or Figure 11 The method.
[0161] In view of the above, in one example, an image processing device 600 is provided for detecting the reflection of an identified object 410 in an image frame 402 captured by a camera 120.
[0162] Furthermore, the memory 602 contains instructions executable by the processing circuitry 604, thereby enabling the image processing device 600 to operate for executing instructions according to... Figure 10 or Figure 11 The method.
[0163] Figure 6 A carrier 605, or program carrier, is further shown, which includes a computer program 603 as directly described above. The carrier 605 may be one of an electronic signal, an optical signal, a radio signal, and a computer-readable medium.
[0164] Furthermore, the processing module 601 may include an input / output unit 606. According to an embodiment, the input / output unit 606 may include an image sensor configured to capture the raw video frames described above, such as raw video frames included in the video stream 310 from the image sensor 301.
[0165] According to the various embodiments described above, the image processing device 600 and / or processing module 601 are configured to generate a three-dimensional model of the background scene of the image frame 402 based on the obtained three-dimensional information about the background scene.
[0166] The image processing device 600 and / or processing module 601 are also configured to define an identified object 410 in the image frame 402 based on the image information in the image frame 402.
[0167] The image processing device 600 and / or processing module 601 are also configured to define an identified object 410 in the image frame 402 based on the image information in the image frame 402.
[0168] The image processing device 600 and / or processing module 601 are further configured to define a three-dimensional bounding box 415 of the defined object 410 in a three-dimensional model of the background scene. The three-dimensional bounding box 415 may be based on the defined object 410 in the image frame 402, the obtained position of the object 410 in the three-dimensional model, and the object 410 defined by extrapolation in a plane extending along the normal to the image plane.
[0169] The image processing device 600 and / or processing module 601 are also configured to define surface elements 416 of a three-dimensional bounding box 415, the corresponding surface elements 416 being defined by center coordinates 418 and color values in a three-dimensional model of the background scene.
[0170] The image processing device 600 and / or processing module 601 are also configured to determine the three-dimensional coordinates of a surface 411 in a three-dimensional model of a background scene, which reflects light from surface elements 416 of the three-dimensional bounding box 415 of the object 410 into the camera 120, wherein the determination is performed by tracing the light rays starting from the center coordinates 418 of the surface elements 416 of the three-dimensional bounding box 415 of the object 410 and based on the normals of the surface in the three-dimensional model of the background scene at said three-dimensional coordinates.
[0171] The image processing device 600 and / or processing module 601 are further configured to identify a first pixel 431 in the image frame 402, the first pixel 431 corresponding to the three-dimensional coordinates of the determined surface 411.
[0172] The image processing device 600 and / or processing module 601 are further configured to detect the reflection of the object 410 in the image frame 402 when the mixture of the first color value of the identified first pixel 431, the color value of the surface element 416 of the object 410, and the real data color value of the identified first pixel 431 satisfies a mixing criterion.
[0173] The image processing device 600 and / or processing module 601 are also configured to apply a mask to a pixel region 441 of the image frame 402 in response to the detection of a reflection of the object 410, the pixel region 441 including the identified first pixel 431.
[0174] In some embodiments herein, the image processing device 600 and / or processing module 601 are configured to obtain a second image frame 403 comprising a video sequence 400 including image frame 402, the second image frame 403 including the identified object 410, and
[0175] Based on the image information in the second image frame 403, the identified object 410 is defined in the second image frame 403, and based on the second position of the object 410 in the 3D model, the corresponding second center coordinates 420 of the surface elements 416 of the 3D bounding box 415 of the object 410 in the 3D model of the background scene are determined.
[0176] In a 3D model of the background scene, the second 3D coordinates of surface 411 or second surface 412 are determined. Surface 411 or second surface 412 reflects light from surface element 416 of the 3D bounding box 415 of object 410 into camera 120. The second 3D coordinates are different from the determined 3D coordinates of surface 411. The determination is performed by tracing the light rays from the second center coordinates 420 of surface element 416 and based on the second normal of surface 411 or second surface 412 in the 3D model of the background scene at the second 3D coordinates. A second pixel 432 in the second image frame 403 is identified, corresponding to the determined second 3D coordinates of the surface or second surface 412.
[0177] Obtain the second color value of the second pixel 432, and when the mixture of the first color value and the second color value satisfies the second mixing criterion, confirm the reflection of the detected object 410, or
[0178] When the mixing of the first color value and the second color value does not meet the second mixing criterion, the detection of the reflection of object 410 is rejected, and the color value of the first pixel 431 in the second image frame 403 is equal to the first color value.
[0179] In some embodiments herein, the image processing device 600 and / or processing module 601 are configured to apply a mask to a second pixel region 442 of a second image frame 403 in response to the detection of a reflection of the confirmed object 410, the second pixel region 442 including the identified second pixel 432.
[0180] In some embodiments herein, the image processing device 600 and / or processing module 601 are configured to determine, in response to the detection of reflections from the rejected object 410, not to apply a mask to the second pixel region 442 including the identified second pixel 432.
[0181] In some embodiments herein, the image processing device 600 and / or the processing module 601 are configured to extrapolate the object 410 by being configured to extrapolate the boundary pixels of the object 410 from the captured image frame 402.
[0182] As used herein, the term "module" can refer to one or more functional modules, each of which can be implemented as one or more hardware modules and / or one or more software modules and / or combinations of software / hardware modules. In some examples, a module can represent a functional unit implemented as software and / or hardware.
[0183] As used herein, the terms "computer program carrier," "program carrier," or "carrier" can refer to one of electronic signals, optical signals, radio signals, and computer-readable media. In some examples, a computer program carrier may exclude transient propagating signals, such as electronic, optical, and / or radio signals. Therefore, in these examples, a computer program carrier may be a non-transitory carrier, such as a non-transitory computer-readable medium.
[0184] As used herein, the term "processing module" can include one or more hardware modules, one or more software modules, or a combination thereof. Any such module, whether hardware, software, or a combination of hardware and software modules, can be a connection means, providing means, configuring means, responding means, disabling means, etc., as disclosed herein. By way of example, the expression "means" can refer to a module corresponding to the modules listed above in conjunction with the accompanying drawings.
[0185] As used in this article, the term "software module" can refer to software applications, dynamic link libraries (DLLs), software components, software objects, objects based on the Component Object Model (COM), software functions, software engines, executable binary software files, etc.
[0186] The terms "processing module" or "processing circuitry" may include processing units herein, including, for example, one or more processors, application-specific integrated circuits (ASICs), field-programmable gate arrays (FPGAs), etc. Processing circuitry, etc., may include one or more processor cores.
[0187] As used herein, the expression “configured to / for” can mean that the processing circuitry is configured, such as to be adapted or operable to, by means of software and / or hardware configuration, perform one or more of the actions described herein.
[0188] As used herein, the term "action" can refer to a movement, step, operation, response, reaction, activity, etc. It should be noted that an action in this document may be divided into two or more sub-actions (where applicable). Furthermore, it should also be noted, where applicable, that two or more actions described herein may be combined into a single action.
[0189] As used herein, the term "memory" can refer to hard disks, magnetic storage media, portable computer floppy disks or optical disks, flash memory, random access memory (RAM), etc. Additionally, the term "memory" can refer to the internal register memory of processors, etc.
[0190] As used herein, the term "computer-readable medium" can refer to a Universal Serial Bus (USB) memory, a DVD, a Blu-ray disc, a software module that receives data as a stream, flash memory, a hard disk drive, a memory card such as a Memory Stick, a Multimedia Card (MMC), a Secure Digital Card (SD), etc. One or more of the foregoing examples of computer-readable media can be provided as one or more computer program products.
[0191] As used herein, the term “computer-readable code unit” can be the text of a computer program, a portion or the entire binary file representing the computer program in a compiled format, or anything in between.
[0192] As used herein, the terms “number” and / or “value” can be any kind of number, such as binary, real, imaginary, or rational numbers. Furthermore, a “number” and / or “value” can be one or more characters, such as letters or a string of letters. A “number” and / or “value” can also be represented by a string of bits, i.e., zero and / or one.
[0193] As used herein, the expression "in some embodiments" has been used to indicate that features of the described embodiments may be combined with any other embodiments disclosed herein.
[0194] Although various embodiments have been described, many different changes, modifications, etc., will become apparent to those skilled in the art. Therefore, the described embodiments are not intended to limit the scope of this disclosure.
Claims
1. A method performed by an image processing device (120, 260) for detecting the reflection of an identified object (410) in an image frame (402) captured by a camera (120), the method comprising: Based on the obtained three-dimensional information about the background scene, a three-dimensional model of the background scene of the image frame (401) is generated (1003); Based on the image information in the image frame (402), the identified object (410) is defined (1004a) in the image frame (402). The three-dimensional bounding box (415) of the object (410) defined in the three-dimensional model of the background scene is defined (1004b). Define (1005) the surface elements (416) of the three-dimensional bounding box (415), and the corresponding surface elements (416) are defined by the center coordinates (418) and color values in the three-dimensional model of the background scene; Determine (1006) the three-dimensional coordinates of surface (411) in the three-dimensional model of the background scene, the surface (411) reflecting light from the surface elements (416) of the three-dimensional bounding box (415) of the object (410) into the camera (120), wherein the determination is performed by tracing the light from the center coordinates (418) of the surface elements (416) of the three-dimensional bounding box (415) of the object (410) based on the normal of the surface in the three-dimensional model of the background scene at the three-dimensional coordinates; Identify (1007) the first pixel (431) in the image frame (402) corresponding to the three-dimensional coordinates of the determined surface (411). When the mixture of the first color value of the identified first pixel (431), the color value of the surface element (416) of the object (410), and the real data color value of the identified first pixel (431) satisfies the mixing criterion, the reflection of the object (410) in the image frame (402) is detected (1008); and In response to the detection of reflection of the object (410), a mask is applied (1109) to a pixel region (441) of the image frame (402), the pixel region (441) including the identified first pixel (431).
2. The method according to claim 1, wherein, The actual data color value is obtained from one or more background image frames (401b) or from one or more neighboring pixels of the first pixel (431) identified in the image frame (402).
3. The method according to claim 1, further comprising: Obtain (1110) a second image frame (403) of a video sequence (400) including the image frame (402), the second image frame (403) including the identified object (410); Based on the image information in the second image frame (403), the identified object (410) is defined (1111) in the second image frame (403). Based on the second position obtained by the object (410) in the three-dimensional model, the corresponding second center coordinates (420) of the surface element (416) of the three-dimensional bounding box (415) of the object (410) in the three-dimensional model of the background scene are determined. Determine (1113) the second three-dimensional coordinates of the surface (411) or the second surface (412) in the three-dimensional model of the background scene, the surface (411) or the second surface (412) reflecting light from the surface element (416) of the three-dimensional bounding box (415) of the object (410) into the camera (120), and the second three-dimensional coordinates are different from the determined three-dimensional coordinates of the surface (411), wherein the determination is performed by tracing light starting from the second center coordinate (420) of the surface element (416) and based on the second normal of the surface (411) or the second surface (412) in the three-dimensional model of the background scene at the second three-dimensional coordinates, and Identify (1114) the second pixel (432) in the second image frame (403), the second pixel (432) corresponding to the determined second three-dimensional coordinates of the surface or the second surface (412); Obtain the second color value of the second pixel (432) (1115); and When the mixture of the first color value and the second color value satisfies the second mixing criterion, the detection of the reflection of the object (410) is confirmed (1116a); or When the mixing of the first color value and the second color value does not meet the second mixing criterion, the detection of the reflection of the object (410) is rejected (1116b), and the color value of the first pixel (431) in the second image frame (403) is equal to the first color value.
4. The method according to claim 3, further comprising: In response to the detection of the reflection of the object (410), a mask is applied (1117) to the second pixel region (442) of the second image frame (403), the second pixel region (442) including the identified second pixel (432).
5. The method according to claim 3, further comprising: In response to the detection of a rejection of the reflection of the object (410), it is determined (1118) that the mask is not applied to the second pixel region (442) including the identified second pixel (432).
6. The method according to any one of claims 1 to 5, wherein, The masked pixel region (441, 442) includes all pixels corresponding to the projection (421) of the surface element (416) of the three-dimensional bounding box (415) onto the surface (411) or the second surface (412) in the three-dimensional model of the background scene, wherein the projection is reflected into the camera (120).
7. The method according to claim 1, wherein, The three-dimensional information about the background scene includes the three-dimensional coordinates of the surface in the background scene, the corresponding normal vector of the surface, the three-dimensional coordinates of the camera (120), and the orientation of the camera (120).
8. The method according to claim 1, wherein, The color value is either a hue value or a combination of Cb and Cr values in the YUV format.
9. The method according to claim 1, wherein, Extrapolation of the object (410) is performed by extrapolating the boundary pixels of the object (410) from the captured image frame (402).
10. The method according to claim 1, wherein, The corresponding surface element (416) of the three-dimensional bounding box (415) corresponds to a plurality of pixels from the captured image frame (402), and wherein the color value of the corresponding surface element (416) is calculated as the average of the color values of the corresponding plurality of pixels.
11. An image processing apparatus (120, 260) configured to perform the method according to any one of claims 1 to 10.
12. The image processing apparatus (120, 260) according to claim 11, wherein, The image processing device (120, 260) is a video camera (120) or a video server (260).
13. The image processing apparatus (120, 260) according to claim 12, wherein, The video camera (120) includes a surveillance camera.
14. A computer-readable medium (605) comprising a computer program (603) including computer-readable code units that, when executed on an image processing apparatus (120, 260), cause the image processing apparatus (120, 260) to perform the method according to any one of claims 1 to 10.
Citation Information
Patent Citations
Light tracking optimization method facing 3D scenes
CN108090947A
Method and system for identifying reflective surfaces in scene
CN108140255A
Dynamic image compensation for pre-touch localization on a reflective surface
US20170147142A1