Image processing device

The image processing apparatus enhances human body detection and motion detection sensitivity in specific regions, ensuring accurate replacement of human bodies with different images while maintaining real-time performance.

JP7869735B2Active Publication Date: 2026-06-03NTT DOCOMO INC

Patent Information

Authority / Receiving Office
JP · JP
Patent Type
Patents
Current Assignee / Owner
NTT DOCOMO INC
Filing Date
2022-11-08
Publication Date
2026-06-03

AI Technical Summary

Technical Problem

Existing image processing methods struggle to accurately replace regions containing human bodies with different images while maintaining real-time performance, often leading to unnecessary replacement of areas without human bodies and loss of real-time image quality.

Method used

An image processing apparatus that includes an object detection unit for human bodies, a motion detection unit, an image processing unit for replacing detected objects with different images, and a control unit that adjusts the frame interval for motion detection in regions with decreased detection accuracy to enhance detection sensitivity.

Benefits of technology

The apparatus effectively replaces regions with human bodies with different images while preserving real-time image quality by improving detection accuracy in challenging areas and maintaining normal frame intervals elsewhere.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure 0007869735000001
    Figure 0007869735000001
  • Figure 0007869735000002
    Figure 0007869735000002
  • Figure 0007869735000003
    Figure 0007869735000003
Patent Text Reader

Abstract

To provide an image processing device capable of highly accurately replacing a region including a specific object in a video with another image while securing a real-time property of the video.SOLUTION: An image processing device 10 comprises: a human body detection unit 12 which performs human body detection processing for detecting a human body included in a video; a dynamic body detection unit 13 which performs dynamic body detection processing for detecting a dynamic body included in the video by comparing a first frame of the video with a second frame older than the first frame; an image processing unit 14 which changes a human body region to another image not including a human body and changes a dynamic body region to another image not including a dynamic body; an extraction unit 15 which extracts a partial region PA in which detection accuracy of the human body detection processing is likely to decrease on the basis of the results of the human body detection processing; and a control unit 16 which makes a frame interval between the first and second frames applied to the dynamic body detection processing for the partial region PA in the video longer than a frame interval between the first and second frames applied to the dynamic body detection processing for a region other than the partial region PA in the video.SELECTED DRAWING: Figure 1
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] One aspect of the present invention relates to an image processing apparatus.

Background Art

[0002] In Patent Document 1, both the human body area detected by human body detection processing (object detection processing) and the moving body area detected by moving body detection processing are replaced with a separate image that does not include the human body (for example, an image subjected to blurring processing, watermark processing, mosaic processing, filling processing, etc.) to obtain an image that does not include the protection target (human body). A method is disclosed.

Prior Art Documents

Patent Documents

[0003]

Patent Document 1

Summary of the Invention

Problems to be Solved by the Invention

[0004] According to the method described in Patent Document 1, for example, if a human body is captured in an image taken by a camera installed at a certain monitoring location, the area containing the human body can be replaced with a different image, thereby protecting the privacy of individuals while monitoring objects other than people. In such a method, it is necessary to detect human bodies in images with high accuracy and minimize the possibility of missing human body detections. For example, one possible approach is to increase the sensitivity of the motion detection process (i.e., to make it easier to detect even relatively slow-moving objects as moving objects) so that even if the human body detection process fails, the motion detection process can still detect the human body as a moving object. However, if the sensitivity of the motion detection process is made too high, there is a high possibility that areas where there are no human bodies will be detected as moving areas and replaced with a different image. Replacing areas where there are no human bodies with a different image in this way is a wasteful process. Furthermore, if areas where there are no human bodies and where it is preferable to display a real-time landscape image are replaced with a different image (for example, a past landscape image), the real-time nature of the image is lost.

[0005] Therefore, one aspect of the present invention aims to provide an image processing device that can replace a region containing a specific object in an image with another image with high precision, while ensuring the real-time nature of the image. [Means for solving the problem]

[0006] An image processing apparatus according to one aspect of the present invention includes: an object detection unit that performs object detection processing to detect a predetermined specific object included in an image; a motion detection unit that performs motion detection processing to detect a moving object that has moved a certain amount or more between the second frame and the first frame by comparing the first frame of the image with a second frame that is earlier than the first frame; an image processing unit that changes an object region containing the specific object detected by the object detection unit into a separate image that does not contain the specific object, and changes a motion region containing a moving object detected by the motion detection unit into a separate image that does not contain the moving object; an extraction unit that extracts a partial region in the image that is prone to a decrease in detection accuracy of the object detection processing based on the results of the object detection processing; and a control unit that makes the frame interval between the first frame and the second frame applied to motion detection processing for a partial region in the image longer than the frame interval between the first frame and the second frame applied to motion detection processing for regions other than the partial region in the image.

[0007] According to one aspect of the present invention, an image processing apparatus can extract a partial region where the detection accuracy of the object detection process tends to decrease, and make the frame interval (frame interval between the first frame and the second frame) of the motion detection process applied to that partial region longer than the normal frame interval applied to other regions. With the above configuration, even if a specific object present in the partial region is not detected by the object detection process, the possibility of that specific object present in the partial region being detected as a moving object by the motion detection process can be increased. As a result, the failure to detect the specific object (moving object) can be suppressed, and an image in which the specific object has been removed can be appropriately generated. Furthermore, since the frame interval of the motion detection process applied to regions other than the partial region is maintained at the normal length (a shorter frame interval than the frame interval applied to the partial region), it is possible to prevent deterioration of the real-time performance of regions other than the partial region (i.e., objects other than the specific object that actually stopped moving a considerable time ago continue to be detected as moving objects, causing the region in which the moving object was detected to be replaced with a different image that is not a real-time image). Therefore, according to the above image processing apparatus, it is possible to replace the region containing the specific object in the video with a different image with high accuracy while ensuring the real-time performance of the video. [Effects of the Invention]

[0008] According to one aspect of the present invention, it is possible to provide an image processing device that can replace a region containing a specific object in an image with another image with high precision while ensuring the real-time nature of the image. [Brief explanation of the drawing]

[0009] [Figure 1] This figure shows an example of the functional configuration of an image processing apparatus according to one embodiment. [Figure 2] This diagram schematically shows the processing pattern of the image processing device (extraction unit). [Figure 3] This figure is intended to explain the third pattern in Figure 2. [Figure 4] This figure is intended to explain the third pattern in Figure 2. [Figure 5] This figure is intended to explain the third pattern in Figure 2. [Figure 6] This is a flowchart illustrating an example of how an image processing device operates. [Figure 7] This figure shows an example of the hardware configuration of an image processing device. [Modes for carrying out the invention]

[0010] Hereinafter, an embodiment of the present invention will be described in detail with reference to the attached drawings. In the description of the drawings, the same or equivalent elements will be denoted by the same reference numerals, and redundant descriptions will be omitted.

[0011] Figure 1 shows an example of the functional configuration of an image processing device 10 according to one embodiment. The image processing device 10 is used, for example, in a system for monitoring a certain place while protecting the privacy of people appearing in the video. For example, the image processing device 10 detects a predetermined specific object or moving body from images (video) sequentially acquired by an imaging unit such as a camera that is installed (fixed) in a predetermined location (e.g., a monitoring area) and has a fixed shooting direction, and performs a process to replace the area in which the specific object or moving body is detected with another image that does not contain the specific object or moving body. In this embodiment, as an example, the specific object is a human body. However, the specific object may be a living organism other than a human, such as a dog or a cat, or a non-living thing such as a robot.

[0012] As shown in Figure 1, the image processing device 10 includes an image acquisition unit 11, a human body detection unit 12 (object detection unit), a motion detection unit 13, an image processing unit 14, an extraction unit 15, and a control unit 16.

[0013] The image acquisition unit 11 sequentially acquires images at predetermined intervals from an imaging unit such as a camera installed (fixed) in a predetermined location (e.g., a monitoring area). The images acquired sequentially by the image acquisition unit 11 (i.e., frames that constitute a video (moving image)) are provided to the human body detection unit 12 and the motion detection unit 13. The imaging unit may be provided within the image processing device 10 itself, or it may be an external device configured to communicate data with the image processing device 10.

[0014] The human body detection unit 12 performs human body detection processing (object detection processing) on ​​the video (each frame of the video) acquired by the image acquisition unit 11. Human body detection processing is the process of detecting a human body (specific object) contained in the video. When the human body detection unit 12 detects a human body through human body detection processing, it sets a human body region (object region) that includes the detected human body. Human body detection processing may be performed, for example, for each frame of the video, or it may be performed at a predetermined execution interval (frame interval).

[0015] The human body detection unit 12 performs human body detection processing by, for example, using known methods such as pattern matching. For example, the human body detection unit 12 detects human bodies in the video by comparing each frame contained in the video with a pre-prepared pattern image (an image representing a human body). However, various other known methods can be used for human body detection processing. For example, the human body detection unit 12 may input feature quantities of frames in the video into a machine learning model trained by deep learning or the like, and determine the human body region based on the output result from the machine learning model (for example, the region in the frame that is estimated to be a human body). In this embodiment, as an example, the human body region is a rectangular region containing the detected human body (see human body region A1 in Figure 2, etc.). However, the shape of the human body region is not limited to the above and may be a shape corresponding to the shape of the detected human body. For example, the outer edge of the human body region may be extracted along the boundary line between the detected human body and the background.

[0016] The motion detection unit 13 performs motion detection processing on the video acquired by the image acquisition unit 11. Motion detection processing is a process that detects motion objects that have moved a certain amount or more between the second frame and the first frame by comparing the first frame of the video (for example, the current frame, which is the most recent at the time of processing) with the second frame, which is earlier than the first frame. When motion is detected by the motion detection processing, the motion detection unit 13 sets a motion region that includes the detected motion object. Motion detection processing may be performed, for example, for every frame of the video, or at a predetermined execution interval (frame interval).

[0017] The moving object detection unit 13 detects moving objects in the video by using known methods such as, for example, the background subtraction method, the statistical background subtraction method, etc. For example, the moving object detection unit 13 compares the first frame with one or more frames up to the second frame acquired before the first frame to detect the moving objects shown in the first frame. Note that the detected moving objects may be the same object as the human body detected by the human body detection unit 12, or may be an object other than the human body. Also, the moving object area may be a rectangular shape including the detected moving object, similar to the above-described human body area, or may be a shape corresponding to the shape of the detected moving object. For example, the outer edge of the moving object area may be extracted along the boundary line between the detected moving object and the background.

[0018] Here, by increasing the frame interval between the first frame and the second frame applied to the moving object detection process, the sensitivity of the moving object detection process can be improved. Specifically, by increasing the frame interval (that is, by referring to more past frames for moving object detection), for example, an object that moves slowly over a relatively long time can be detected as a moving object. On the other hand, by shortening the frame interval, only an object that has moved more than a certain amount in the most recent few frames (that is, an object with a relatively large amount of movement in a short period) can be detected as a moving object. The frame interval applied to the moving object detection process can be arbitrarily set. In the following description, the default frame interval between the first frame and the second frame applied to the moving object detection process is referred to as the "standard frame interval".

[0019] The image processing unit 14 changes the human body area including the human body detected by the human body detection unit 12 to a different image that does not include the human body, and changes the moving object area including the moving object detected by the moving object detection unit 13 to a different image that does not include the moving object. The different image may be, for example, a landscape image (an image in which no human body is captured) taken in the most recent past, or an image to which blurring processing, mosaic processing, filling processing, etc. have been applied so that the human body cannot be recognized.

[0020] Through the processing of the image processing unit 14, in the human body area and the moving object area, instead of real-time video, the above-mentioned separate images are arranged (displayed). In this way, privacy is protected by replacing the areas where a human body may be reflected (human body area and moving object area) with separate images. On the other hand, in the areas other than the human body area and the moving object area in the frame, the image (video) acquired by the image acquisition unit 11 is displayed as it is.

[0021] Based on the result of the human body detection process by the human body detection unit 12, the extraction unit 15 extracts a partial area PA (see FIG. 2) which is a part of the area in the video and where the detection accuracy of the human body detection process is likely to decrease.

[0022] FIG. 2 is a diagram schematically showing the processing pattern of the extraction unit 15. Referring to FIG. 2, specific examples (Pattern 1 to Pattern 3) of the processing by the extraction unit 15 will be described. Note that the processing of Pattern 1 to Pattern 3 may be used in combination. For example, the extraction unit 15 may execute two or more of the processing of Pattern 1 to Pattern 3. In the example of FIG. 2, five human bodies P are reflected in the image IM (one frame in the video) acquired by the image acquisition unit 11. In such a situation where a plurality of people are densely packed and have little movement (for example, a situation where a plurality of people are lined up to form a queue and a plurality of people overlap each other in the video), the detection accuracy of the human body detection process is likely to decrease. Also, in such a situation (that is, when the movement of each human body P is relatively slow), in the moving object detection process with normal sensitivity (standard frame interval), it is difficult to detect the human body P as a moving object. Therefore, the extraction unit 15 extracts an area that is likely to apply to such a situation as the partial area PA. As a result, by the processing of the control unit 16 described later, the possibility that the human body P included in the partial area PA is detected as a moving object by the moving object detection process can be increased.

[0023] (Pattern 1) As shown in the upper part of Figure 2, when the human body detection process detects multiple overlapping human body regions A1, A2, and A3, the extraction unit 15 may extract the region containing the multiple human body regions A1, A2, and A3 as a partial region PA. In the area where multiple human body regions A1, A2, and A3 overlap in this way, there is a possibility that a situation is occurring where multiple people are densely packed together and moving little, as described above. According to the first pattern, a partial region PA can be suitably extracted based on the above considerations.

[0024] (Pattern 2) The human body detection processing by the human body detection unit 12 may be configured to output the confidence level of the detected human body region (for example, a numerical value (percentage, etc.) indicating the probability that it is a human body). In this case, as shown in the middle section of Figure 2, if the human body detection processing detects a low-confidence region (human body regions A1, A3 in the example of Figure 2), which is a human body region with a confidence level below a predetermined threshold, the extraction unit 15 may extract the region containing the low-confidence region as a partial region PA. In a region where multiple people are densely packed together as described above, people overlap, and parts of people located behind are hidden by parts of people located in front, making it highly likely that the confidence level of human body detection will be low (i.e., human body regions with low confidence level will be detected). According to the second pattern, partial region PA can be suitably extracted based on the above considerations.

[0025] (Pattern 3) As shown in the lower part of Figure 2, the extraction unit 15 may extract a region as a sub-region PA if the detection results of the human body detection process vary by a certain amount or more among multiple frames included in the most recent predetermined period. In the example in Figure 2, three human body regions A1, A2, and A3 are detected by the human body detection process for one frame (the left frame of the two frames arranged in the lower part of Figure 2), whereas only one human body region A3 is detected in the human body detection process for a frame temporally adjacent to that frame (the right frame of the two frames arranged in the lower part of Figure 2). Regions in which the detection results of the human body detection process change significantly between adjacent frames are likely to be regions where multiple people are densely packed together (for example, regions where multiple people are densely packed together, and the degree of overlap between people repeatedly changes, making it easy for the detection results of the human body detection process to change). For example, consider a situation where multiple people (in this case, five people) are lined up in a queue, as shown in Figure 3. In this situation, while the human body regions A1-A5 corresponding to all five people are correctly detected in the first frame FA, in the next frame FB, a subtle change in the overlap between adjacent people may cause human body regions A2 and A4 corresponding to some people (in this example, the second and fourth people from the left) to no longer be detected. The third pattern is a process that extracts a partial region PA based on this behavior of the human body detection process.

[0026] More specifically, in the third processing pattern described above, the extraction unit 15 may perform the following processing. That is, the extraction unit 15 may detect differences in human body regions between adjacent frames in a video for a predetermined period (i.e., a video including the current frame and past frames within a predetermined period from the current frame), and if the frequency of occurrence of such differences (for example, the number of frames in which the same difference appears) is greater than or equal to a predetermined threshold, it may extract the region in the video that includes the difference region corresponding to the difference as a partial region PA. With the above configuration, partial region PAs can be suitably extracted based on the inter-frame variation of the detection results of the human body detection processing described above.

[0027] Referring to Figures 4 and 5, an example of the processing of the extraction unit 15 in the third pattern described above will be explained. In the example in Figures 4 and 5, the predetermined period is a period that includes four temporally consecutive frames F1 to F4, and the threshold value described above is "2". Also, frame F1 is the oldest frame, and frame F4 is the most recent frame (current frame).

[0028] In the example in Figure 4, two human body regions A1 and A2 are detected in the first and third frames F1 and F3, and one human body region A1 is detected in the second and fourth frames F2 and F4. In this case, the difference in human body regions between frames F1 and F2 is the difference region D1 corresponding to human body region A2. Similarly, the difference in human body regions between frames F2 and F3 and between frames F3 and F4 is also the difference region D1 corresponding to human body region A2. Within a predetermined period, the difference region D1 appears three times. That is, the number of frames in which the difference region D1 appears within a predetermined period is "3". Therefore, the extraction unit 15 extracts the region containing the difference region D1 that has an appearance frequency of two times or more within a predetermined period (in the example in Figure 4, the same rectangular region as the difference region D1) as a partial region PA.

[0029] In the example in Figure 5, three human body regions A1, A2, and A3 are detected in the first frame F1, one human body region A2 is detected in the second and fourth frames F2 and F4, and two human body regions A1 and A3 are detected in the third frame F3. In this case, the difference in human body regions between frames F1 and F2 is the difference regions D1 and D3 corresponding to human body regions A1 and A3. Also, the difference in human body regions between frames F2 and F3 and between frames F3 and F4 is the difference regions D1, D2, and D3 corresponding to human body regions A1, A2, and A3. Within a predetermined period, difference regions D1 and D3 appear 3 times, and difference region D2 appears 2 times. Therefore, the extraction unit 15 extracts the region containing difference regions D1, D2, and D3 that have an occurrence frequency of 2 times or more within a predetermined period (in the example in Figure 4, the rectangular region encompassing difference regions D1, D2, and D3) as a partial region PA.

[0030] As shown in Figure 2, in any of the first to third patterns, if the extraction unit 15 extracts multiple regions within a predetermined distance range as regions corresponding to partial regions, it may extract a continuous region including these multiple regions as a partial region PA. Here, the threshold for whether or not a region is within the predetermined distance range (the upper limit of the distance between the furthest regions) can be arbitrarily selected. The distance between regions can be, for example, the distance between the centers of the regions. For example, in the first pattern, if the extraction unit 15 determines that the human body regions A1, A2, and A3 extracted as regions corresponding to partial regions are within the predetermined distance range, it may extract a rectangular partial region PA including these multiple human body regions A1, A2, and A3. In the second pattern as well, if the extraction unit 15 determines that the human body regions A1 and A3 extracted as regions corresponding to partial regions are within the predetermined distance range, it may extract a rectangular partial region PA including these multiple human body regions A1 and A3. For the third pattern, as shown in the example in Figure 5 above, the extraction unit 15 may also extract a rectangular sub-region PA containing multiple regions (difference regions D1, D2, D3) when it is determined that these regions are within a predetermined distance range. By extracting a single continuous region containing multiple regions within a predetermined distance range as a sub-region PA, the number of sub-region PAs can be reduced, making it easier to manage the sub-region PAs. Furthermore, if two regions determined to be sub-regions are separated from each other, and the distance between these two regions is within a predetermined distance range, the region located between these two regions is also highly likely to be a region where the detection accuracy of the human body detection process is likely to decrease. According to the above process, such a region (the region located between the two regions mentioned above) can be appropriately included in the sub-region PA.

[0031] The control unit 16 makes the frame interval between the first and second frames applied to motion detection processing for a partial region PA in the video longer than the standard frame interval applied to motion detection processing for areas other than the partial region PA in the video. In other words, the control unit 16 controls the operation of the motion detection unit 13 to cause motion detection processing to be performed on the partial region PA at a frame interval longer than the standard frame interval. As a result, the sensitivity of motion detection processing can be improved for partial regions PA where the detection accuracy of human body detection processing tends to decrease (i.e., areas where a human body P is present but is difficult to detect by human body detection processing), and the possibility of human body P being detected as a moving object in the partial region PA can be increased. As a result, the failure to detect human body P in the partial region PA can be suppressed. On the other hand, by applying the standard frame interval to areas other than the partial region PA, deterioration of real-time performance can be prevented (i.e., objects other than human body P that actually stopped moving a considerable time ago continue to be detected as moving objects, causing the area where the moving object was detected to be replaced with a different image that is not a real-time image).

[0032] The partial region PA may be released for the following reasons, for example. For example, the control unit 16 may release the partial region PA if, after the time the partial region PA was extracted by the extraction unit 15, no human body or moving body is detected by either the human body detection process or the motion detection process for the partial region PA. Alternatively, the control unit 16 may release the partial region PA after a certain period of time (a predetermined threshold period) has elapsed since the time the partial region PA was extracted.

[0033] Next, an example of the operation of the image processing device 10 will be described with reference to Figure 6. Note that the series of processes in steps S1 to S5 are repeatedly executed as long as the image acquisition unit 11 continues to acquire video.

[0034] In step S1, the image acquisition unit 11 acquires the video to be processed. For example, the image acquisition unit 11 sequentially acquires images (frames that make up the video) from a camera or other shooting unit installed at a predetermined location.

[0035] In step S2, the human body detection unit 12 and the motion detection unit 13 perform human body detection processing and motion detection processing on the image acquired by the image acquisition unit 11. If a human body is present in the image, the human body detection processing may detect human body regions (for example, human body regions A1, A2, A3 in Figure 2). Similarly, if a moving object is present in the image, the motion detection processing will detect the moving object region. If a partial region PA is extracted, the motion detection unit 13 performs motion detection processing on the partial region PA using a longer frame interval than the standard frame interval applied to motion detection processing for regions other than the partial region PA.

[0036] In step S3, the image processing unit 14, when a human body is detected by the human body detection unit 12, changes the area containing the human body in the video to a different image that does not contain the human body, and when a moving object is detected by the motion detection unit 13, changes the area containing the moving object in the video to a different image that does not contain the moving object.

[0037] In step S4, the extraction unit 15 extracts a partial region PA in which the detection accuracy of the human body detection process is likely to decrease, for example, by executing at least one of the first to third patterns described above.

[0038] In step S5, the control unit 16 makes the frame interval between the first and second frames applied to the motion detection processing for the partial region PA in the video longer than the standard frame interval applied to the motion detection processing for areas other than the partial region PA in the video. As a result, in step S2 while the partial region PA is active (from the time the partial region PA is extracted until it is deactivated), motion detection processing with normal sensitivity (standard frame interval) is performed for areas other than the partial region PA, and motion detection processing with higher sensitivity than normal is performed for the partial region PA. In other words, for the partial region PA, even human bodies (objects) with movements that would not be detected with the standard frame interval are detected as moving objects.

[0039] According to the image processing device 10 described above, a partial region PA in which the detection accuracy of the human body detection process is likely to decrease can be extracted, and the frame interval of the motion detection process applied to the partial region PA (the frame interval between the first frame and the second frame) can be made longer than the standard frame interval applied to other regions. With the above configuration, even if a human body P present in the partial region PA is not detected by the human body detection process, the possibility of the human body P present in the partial region PA being detected as a moving object by the motion detection process can be increased. As a result, the failure to detect human bodies P can be suppressed, and an image in which the human body P (moving object) has been erased can be appropriately generated. In addition, since the frame interval of the motion detection process applied to regions other than the partial region PA (standard frame interval) is maintained at the normal length (a shorter frame interval than the frame interval applied to the partial region PA), it is possible to prevent deterioration of the real-time performance of regions other than the partial region PA (i.e., objects other than human bodies that actually stopped moving a considerable time ago continue to be detected as moving objects, causing the region in which the moving object was detected to be replaced with a different image that is not a real-time image). Therefore, the image processing device 10 can replace areas containing human bodies in a video with other images with high accuracy while ensuring real-time video.

[0040] The image processing apparatus 10 of this disclosure has the following configuration.

[0041] [1] An object detection unit that performs object detection processing to detect a predetermined specific object contained in the video, A motion detection unit performs motion detection processing to detect a moving object that has moved a certain amount or more between the second frame and the first frame by comparing the first frame of the aforementioned video with a second frame that is earlier than the first frame. An image processing unit that changes the object region containing the specific object detected by the object detection unit to a separate image that does not contain the specific object, and changes the motion region containing the motion detected by the motion detection unit to a separate image that does not contain the motion, An extraction unit that extracts a portion of the video image where the detection accuracy of the object detection process is likely to decrease, based on the results of the object detection process, A control unit that makes the frame interval between the first frame and the second frame applied to the motion detection processing for the partial region in the video longer than the frame interval between the first frame and the second frame applied to the motion detection processing for a region other than the partial region in the video, An image processing device equipped with the following features.

[0042] [2] When the object detection process detects a plurality of overlapping object regions, the extraction unit extracts a region containing the plurality of object regions as the partial region. [1] Image processing device.

[0043] [3] The object detection process is configured to output the confidence level of the detected object region, The extraction unit, when the object detection process detects a low-confidence region which is an object region whose confidence level is below a predetermined threshold, extracts the region including the low-confidence region as the sub-region. Image processing device of [1] or [2].

[0044] [4] The extraction unit detects the difference between adjacent object regions in the video over a predetermined period and extracts a region including the difference region corresponding to the difference whose occurrence frequency is above a predetermined threshold as the partial region. One of the following image processing devices: [1] to [3].

[0045] [5] When a plurality of regions within a predetermined distance range are extracted as regions corresponding to the subregion, the extraction unit extracts a continuous region including the plurality of regions as the subregion. [2] to [4], one of the image processing devices.

[0046] Furthermore, the block diagrams used in the description of the above embodiments show functional units. These functional blocks (components) are realized by any combination of at least one of hardware and software. Moreover, the method of realizing each functional block is not particularly limited. That is, each functional block may be realized using one device that is physically or logically coupled, or it may be realized using two or more physically or logically separated devices that are directly or indirectly connected (for example, using wired or wireless connections). A functional block may be realized by combining the above one device or the above multiple devices with software.

[0047] Functions include, but are not limited to, judgment, decision, judgment, calculation, calculation, processing, derivation, investigation, exploration, confirmation, reception, transmission, output, access, resolution, selection, selection, establishment, comparison, assumption, expectation, deem, broadcasting, notifying, communicating, forwarding, configuring, reconfiguring, allocating (mapping), and assigning.

[0048] For example, the image processing apparatus 10 in one embodiment of the present disclosure may function as a computer that performs the image processing method of the present disclosure. Figure 7 is a diagram showing an example of the hardware configuration of the image processing apparatus 10 according to one embodiment of the present disclosure. The image processing apparatus 10 may be physically configured as a computer device including a processor 1001, memory 1002, storage 1003, communication device 1004, input device 1005, output device 1006, bus 1007, etc.

[0049] In the following explanation, the term "device" can be replaced with "circuit," "device," "unit," etc. The hardware configuration of the image processing device 10 may include one or more of the devices shown in Figure 7, or it may be configured to omit some of the devices.

[0050] Each function in the image processing device 10 is realized by loading predetermined software (programs) onto hardware such as the processor 1001 and memory 1002, which allows the processor 1001 to perform calculations, control communication by the communication device 1004, and control at least one of data reading and writing in the memory 1002 and storage 1003.

[0051] The processor 1001 controls the entire computer, for example, by running an operating system. The processor 1001 may consist of a central processing unit (CPU) that includes interfaces with peripheral devices, control units, arithmetic units, registers, and so on.

[0052] Furthermore, the processor 1001 reads programs (program code), software modules, data, etc., from at least one of the storage 1003 and the communication device 1004 into the memory 1002 and executes various processes accordingly. The program used is one that causes the computer to execute at least a part of the operations described in the above embodiment. For example, each functional unit of the image processing device 10 (e.g., the extraction unit 15) may be stored in the memory 1002 and implemented by a control program that runs on the processor 1001, and other functional blocks may be implemented similarly. The above-described various processes have been explained as being executed by one processor 1001, but they may be executed simultaneously or sequentially by two or more processors 1001. The processor 1001 may be implemented by one or more chips. The program may be transmitted from a network via a telecommunications line.

[0053] Memory 1002 is a computer-readable recording medium and may consist of at least one of the following: ROM (Read Only Memory), EPROM (Erasable Programmable ROM), EEPROM (Electrically Erasable Programmable ROM), RAM (Random Access Memory), etc. Memory 1002 may also be called a register, cache, main memory, etc. Memory 1002 can store executable programs (program code), software modules, etc., for carrying out an image processing method according to one embodiment of the present disclosure.

[0054] Storage 1003 is a computer-readable recording medium and may consist of at least one of the following: an optical disc such as a CD-ROM (Compact Disc ROM), a hard disk drive, a flexible disk, a magneto-optical disk (e.g., a compact disc, a digital multipurpose disc, a Blu-ray® disc), a smart card, flash memory (e.g., a card, a stick, a key drive), a floppy® disk, a magnetic strip, etc. Storage 1003 may also be called an auxiliary storage device. The above-mentioned storage medium may be, for example, a database, server, or other suitable medium including at least one of memory 1002 and storage 1003.

[0055] The communication device 1004 is hardware (transceiver / receiver device) for communicating between computers via at least one of a wired network and a wireless network, and is also called a network device, network controller, network card, communication module, etc.

[0056] The input device 1005 is an input device that accepts input from an external source (e.g., a keyboard, mouse, microphone, switch, button, sensor, etc.). The output device 1006 is an output device that outputs to an external source (e.g., a display, speaker, LED lamp, etc.). The input device 1005 and the output device 1006 may be configured as an integrated unit (e.g., a touch panel).

[0057] Furthermore, each device, such as the processor 1001 and memory 1002, is connected by a bus 1007 for communicating information. The bus 1007 may be configured using a single bus, or different buses may be configured for each device.

[0058] Furthermore, the image processing device 10 may be configured to include hardware such as a microprocessor, a digital signal processor (DSP), an ASIC (Application Specific Integrated Circuit), a PLD (Programmable Logic Device), and an FPGA (Field Programmable Gate Array), and some or all of each functional block may be realized by such hardware. For example, the processor 1001 may be implemented using at least one of these hardware components.

[0059] Although this embodiment has been described in detail above, it will be clear to those skilled in the art that this embodiment is not limited to the embodiments described herein. This embodiment can be implemented as a modified and altered form without departing from the spirit and scope of the invention as defined by the claims. Therefore, the description herein is for illustrative purposes only and is not intended to be restrictive in any way to this embodiment.

[0060] The processing procedures, sequences, flowcharts, etc., of each aspect / embodiment described herein may be reordered, provided they are consistent with each other. For example, the methods described herein present various step elements in an exemplary order and are not limited to that specific order.

[0061] Input and output information may be stored in a specific location (e.g., memory) or managed using a management table. Input and output information may be overwritten, updated, or appended to. Output information may be deleted. Input information may be transmitted to other devices.

[0062] The determination may be made by a value represented by 1 bit (0 or 1), by a boolean value (true or false), or by a numerical comparison (for example, a comparison with a predetermined value).

[0063] Each aspect / embodiment described herein may be used individually, in combination, or switched between as needed during implementation. Furthermore, notification of specific information (e.g., notification that "X is") is not limited to explicit notification, but may also be implicit (e.g., by not providing such notification).

[0064] Software should be broadly interpreted to mean instructions, instruction sets, code, code segments, program code, programs, subprograms, software modules, applications, software applications, software packages, routines, subroutines, objects, executable files, execution threads, procedures, functions, and so on, whether they are called software, firmware, middleware, microcode, hardware description languages, or by any other name.

[0065] Furthermore, software, instructions, information, etc., may be transmitted and received via a transmission medium. For example, if software is transmitted from a website, server, or other remote source using at least one of wired technology (such as coaxial cable, fiber optic cable, twisted pair, or digital subscriber line (DSL)) and wireless technology (such as infrared or microwave), then at least one of these wired and wireless technologies is included in the definition of a transmission medium.

[0066] The information, signals, etc. described in this disclosure may be represented using any of the various different techniques. For example, the data, instructions, commands, information, signals, bits, symbols, chips, etc. that may be referred to throughout the above description may be represented by voltage, current, electromagnetic waves, magnetic fields or magnetic particles, optical fields or photons, or any combination thereof.

[0067] Furthermore, the information, parameters, etc., described in this disclosure may be expressed using absolute values, relative values ​​from a predetermined value, or corresponding other information.

[0068] The names used for the parameters described above are not restrictive in any way. Furthermore, the formulas and other expressions using these parameters may differ from those expressly disclosed in this disclosure. Since various information elements can be identified by any suitable name, the various names assigned to these various information elements are not restrictive in any way.

[0069] In this disclosure, the phrase "based on" does not mean "based solely on" unless otherwise specified. In other words, the phrase "based on" means both "based solely on" and "based at least on."

[0070] Any reference to elements using the designations “first,” “second,” etc., as used in this disclosure does not generally limit the quantity or order of those elements. These designations may be used in this disclosure as a convenient way to distinguish between two or more elements. Accordingly, references to the first and second elements do not imply that only two elements may be employed, or that the first element must precede the second element in any way.

[0071] Where the terms “include,” “including,” and variations thereof are used in this disclosure, these terms are intended to be inclusive, as is the term “comprising.” Furthermore, the term “or” as used in this disclosure is not intended to mean exclusive OR.

[0072] In this disclosure, if articles are added through translation, such as a, an, and the in English, this disclosure may include the fact that the noun following these articles is plural.

[0073] In this disclosure, the term "A and B are different" may mean "A and B are different from each other." The term may also mean "A and B are each different from C." Terms such as "separate" and "combine" may be interpreted similarly to "different." [Explanation of symbols]

[0074] 10...Image processing device, 11...Image acquisition unit, 12...Human body detection unit (object detection unit), 13...Motion detection unit, 14...Image processing unit, 15...Extraction unit, 16...Control unit, A1, A2, A3...Human body region (object region), D1, D2, D3...Difference region, F1, F2, F3, F4...Frame, IM...Image (video), P...Human body (specific object), PA...Partial region.

Claims

1. An object detection unit that performs object detection processing to detect predetermined specific objects contained in the video, A motion detection unit performs motion detection processing to detect a moving object that has moved by a certain amount or more between the second frame and the first frame by comparing the first frame of the aforementioned video with a second frame that is earlier than the first frame. An image processing unit that changes the object region containing the specific object detected by the object detection unit to a separate image that does not contain the specific object, and changes the motion region containing the motion detected by the motion detection unit to a separate image that does not contain the motion, An extraction unit that extracts a portion of the video image where the detection accuracy of the object detection process is likely to decrease, based on the results of the object detection process, A control unit that makes the frame interval between the first frame and the second frame applied to the motion detection processing for the partial region in the video longer than the frame interval between the first frame and the second frame applied to the motion detection processing for a region other than the partial region in the video, An image processing device equipped with the following features.

2. When the object detection process detects a plurality of overlapping object regions, the extraction unit extracts a region containing the plurality of object regions as the partial region. The image processing apparatus according to claim 1.

3. The object detection process is configured to output the confidence level of the detected object region. The extraction unit, when the object detection process detects a low-confidence region which is an object region whose confidence level is below a predetermined threshold, extracts the region including the low-confidence region as the sub-region. The image processing apparatus according to claim 1.

4. The extraction unit detects the difference between adjacent object regions in the video over a predetermined period, and extracts a region containing the difference region corresponding to the difference whose occurrence frequency is above a predetermined threshold as the partial region. The image processing apparatus according to claim 1.

5. When multiple regions within a predetermined distance range are extracted as regions corresponding to the sub-region, the extraction unit extracts a continuous region including the multiple regions as the sub-region. The image processing apparatus according to any one of claims 2 to 4.