Image processing device and control method thereof, and imaging device

The image processing device integrates user-defined trajectories with detected objects to refine the region of interest, addressing the challenge of erroneous tracking by accurately setting the region of interest and improving subject capture.

JP7784253B2Active Publication Date: 2025-12-11CANON KK
View PDF 5 Cites 0 Cited by

Patent Information

Application Number
JP2021142699
Authority / Receiving Office
JP · JP
Patent Type
Patents
Current Assignee / Owner
Filing Date
2021-09-01
Publication Date
2025-12-11
Estimated Expiration
2041-09-01

AI Technical Summary

Technical Problem

Existing image processing systems struggle to accurately set the region of interest, leading to erroneous tracking when partial regions of an object are mistakenly identified, making it difficult to distinguish from other regions and resulting in missed subjects.

Method used

An image processing device that integrates user input trajectories with detected object candidates to generate an integrated region of interest, using neural networks for object detection and user-defined trajectories to refine the tracking area.

Benefits of technology

Enables more accurate setting of the region of interest, ensuring appropriate tracking by integrating user-defined trajectories with detected objects, reducing the likelihood of erroneous tracking and improving subject capture.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure 0007784253000001
    Figure 0007784253000001
  • Figure 0007784253000002
    Figure 0007784253000002
  • Figure 0007784253000003
    Figure 0007784253000003
Patent Text Reader

Abstract

To provide an image processing system that further appropriately sets an attention area in an image.SOLUTION: An image processing system includes: image input means for inputting an image; detection means for detecting an object from the image; acceptance means for accepting an input of a trajectory for the image; selection means for selecting, on the basis of a trajectory area determined by the trajectory, two or more objects included in a plurality of detected objects; and integration means for generating an integrated area by integrating two or more areas in an image corresponding to two or more selected objects and setting it as an attention area in the image.SELECTED DRAWING: Figure 2
Need to check novelty before this filing date? Find Prior Art

Description

[Technical Field]

[0001] The present invention relates to a technique for setting an area of ​​interest within an image. [Background technology]

[0002] Current cameras have the ability to detect subject areas with specific characteristics from an image and automatically determine the exposure and focus distance to ensure appropriate capture. Some cameras also have a tracking function that continuously adjusts focus, brightness, and color by continuing to track a pre-selected subject area in subsequent frames. These functions are performed using information about the area of ​​interest in the input image where the subject is located, so it is necessary to set the area of ​​interest appropriately.

[0003] To extract information about a region of interest of a subject from an input image, a target object detection technology is required. For example, technology is used to detect target objects of specific categories, such as a person's face, facial organs (eyes, nose, mouth), or the entire body of a person. In recent years, with the development of deep learning, technology has been realized to detect any subject, such as animals or vehicles, by learning using information about objects of various categories (Non-Patent Documents 1 to 3).

[0004] On the other hand, when the attention area is automatically set using the above-mentioned detection technology, an object unintended by the user may be set as the tracking target area. From this perspective, a method for correcting the attention area through user operation has been proposed. For example, Patent Document 1 discloses a method for switching the tracking target to a specific subject based on a user's touch operation. [Prior art documents] [Patent documents]

[0005] [Patent Document 1] Patent No. 6397454 [Non-patent literature]

[0006] [Non-Patent Document 1] Ross Girshick et al., "Rich Feature Hierarchies for Accurate Object Detection and Semantic Segmentation.", 2014 IEEE Conference on Computer Vision and Pattern Recognition [Non-patent document 2] Wei Liu et al., "SSD: Single Shot MultiBox Detector", Computer Vision-ECCV 2016 [Non-patent document 3] Joseph Redmon et al., "You Only Look Once: Unified, Real-Time Object Detection", 2016 IEEE Conference on Computer Vision and Pattern Recognition Summary of the Invention [Problem to be solved by the invention]

[0007] However, when automatically setting a region of interest using the above-mentioned detection technology, a partial region of an object intended by a user may be set as the region of interest. Performing a tracking process using a partial region as the region of interest can lead to an erroneous result because it is difficult to distinguish it from other regions in the image, resulting in a problem of missing the subject. However, Patent Document 1 only describes a technology for switching subjects and is unable to address such a problem.

[0008] The present invention has been made in view of the above problems, and aims to provide a technique that enables a region of interest in an image to be set more appropriately. [Means for solving the problem]

[0009] In order to solve the above-mentioned problems, an image processing device according to the present invention has the following arrangement. image input means for inputting an image; Object from the image candidate a detection means for detecting the a receiving means for receiving an input of a trajectory for the image; A plurality of objects detected by the detection means based on a trajectory area determined by the trajectory. candidate Included in , which are parts of the same object 2 or more objects candidate a selection means for selecting Two or more objects selected by the selection means candidate two or more regions in the image corresponding to Based on Generate an integrated region in the image Corresponding to the object to be tracked an integration means for setting as an area of ​​interest; Equipped with. [Effects of the Invention]

[0010] An object of the present invention is to provide a technique that enables a region of interest in an image to be set more appropriately. [Brief explanation of the drawings]

[0011] [Figure 1] FIG. 1 is a diagram illustrating an overall configuration of a camera system. [Figure 2] FIG. 2 is a diagram illustrating a functional configuration of the camera system. [Figure 3] 10 is a flowchart illustrating processing during photography in the camera system. [Figure 4] FIG. 1 is a diagram illustrating the structure of a neural network. [Figure 5] 10 is a detailed flowchart of the first detection frame selection (S403). [Figure 6] 10 is a detailed flowchart of second detection frame selection (S411). [Figure 7] 10 is a detailed flowchart of the integrated frame selection (S413). [Figure 8]1A and 1B are diagrams illustrating an example of an input image and a detection result of an object candidate. [Figure 9] FIG. 10 is a diagram showing an example of a result of first detection frame selection. [Figure 10] 10A and 10B are diagrams illustrating setting of a trajectory frame based on a trajectory input by a user. [Figure 11] FIG. 10 is a diagram illustrating an example of generating a combination of frames. [Figure 12] FIG. 10 is a diagram illustrating an example of generating an integrated frame. [Figure 13] FIG. 10 is a diagram illustrating the degree of overlap between an integrated frame and a trajectory frame. [Figure 14] 10 is a flowchart illustrating the processing of a selection unit 270 in Modification 1. [Figure 15] 10A and 10B are diagrams illustrating examples of a center map and a size map. [Figure 16] 10A and 10B are diagrams illustrating another example of setting a trajectory frame based on a trajectory input by a user. DETAILED DESCRIPTION OF THE INVENTION

[0012] Hereinafter, embodiments will be described in detail with reference to the accompanying drawings. Note that the following embodiments do not limit the scope of the invention claimed. Although multiple features are described in the embodiments, not all of these multiple features are necessarily essential to the invention, and multiple features may be combined arbitrarily. Furthermore, in the accompanying drawings, the same reference numerals are used to designate the same or similar components, and redundant explanations will be omitted.

[0013] (First embodiment) A first embodiment of an image processing device according to the present invention will be described below using a camera system as an example. However, the present invention can be implemented in any electronic device that tracks an object area in a moving image. Such electronic devices include, but are not limited to, imaging devices such as digital cameras and digital video cameras, as well as personal computers with camera functions, mobile phones, drive recorders, robots, drones, etc.

[0014] <System configuration> Figure 1 shows the overall configuration of a camera system. The camera system includes an imaging device 101, RAM 102, ROM 103, image processing device 104, input / output device 105, and control device 106. Each component is configured to be able to communicate with each other and is connected via a bus or the like. Note that, although it is assumed here that each component shown in Figure 1 constitutes an integrated device (camera), they may also be configured as a distributed system connected via a network.

[0015] The imaging device 101 is composed of a photographing lens, an image sensor, an A / D converter, an aperture control device, and a focus control device. The photographing lens includes a fixed lens, a zoom lens, a focus lens, an aperture, and an aperture motor. The image sensor includes a CCD or CMOS that converts an optical image of a subject into an electrical signal. The A / D converter converts an analog signal into a digital signal. The imaging device 101 converts the subject image formed on the imaging plane of the image sensor by the photographing lens into an electrical signal, and the A / D converter applies signal processing to the electrical signal through A / D conversion processing, supplying the resulting image data to RAM 102. The aperture control device controls the operation of the aperture motor and changes the aperture opening diameter to control the aperture of the photographing lens. The focus control device controls the operation of the focus motor based on the phase difference between a pair of focus detection signals obtained from the image sensor, and drives the focus lens to control the focus state of the photographing lens.

[0016] The RAM 102 stores image data obtained by the imaging device 101 and image data to be displayed on the input / output device 105. The RAM 102 has a storage capacity sufficient to store a predetermined number of still images and a predetermined length of video. The RAM 102 also serves as a memory (video memory) for image display, and supplies image data for display to the input / output device 105.

[0017] The ROM 103 is a storage device such as a magnetic storage device or semiconductor memory, and stores programs loaded based on the operations of the image processing device 104 and the control device 106, data to be stored for a long period of time, and the like.

[0018] The image processing device 104 detects and selects object candidate regions from the image, superimposes the image with the object candidate regions, and outputs the results to the input / output device 105 and the control device 106. The object candidates referred to here refer to unspecified objects of various categories such as animals, vehicles, insects, and aquatic animals. In this embodiment, the image processing device 104 performs object detection by outputting the position, size, and likelihood representing the object-likeliness of the unspecified object candidate regions as detection results. The configuration and operation of the image processing device 104 will be described in detail later.

[0019] The input / output device 105 is a device through which the camera system 100 receives instructions from a user and through which the user obtains various pieces of information from the camera system 100. The input / output device 105 is composed of, for example, a group of input devices such as switches, buttons, keys, and a touch panel, and a display such as an LCD or an organic EL display. The control device 106 detects inputs from the group of input devices via a bus, and the control device 106 controls each component to perform operations according to the inputs. In addition, the touch-sensing surface of the touch panel of the input / output device 105 serves as the display surface of the display. The touch panel may be any of various touch panel types, such as a resistive type, a capacitive type, or an optical sensor type. In addition, the input / output device 105 displays a live view image by sequentially transferring and displaying image data. In the following description, the input / output device 105 is described as being configured as a touch display in which the touch panel and display are integrated.

[0020] The control device 106 is composed of a CPU (Central Processing Unit). The control device 106 executes programs stored in the ROM 103 and realizes the functions of the camera system 100. The control device 106 also controls the image capture device 101, performing aperture control, focus control, and exposure control. For example, the control device 106 performs AE (auto exposure) processing, which automatically determines exposure conditions (shutter speed or accumulation time, aperture value, and sensitivity) based on subject brightness information of image data obtained by the image capture device 101. The control device 106 can also automatically set a focus detection area using the subject area detection results obtained by the image processing device 104, thereby achieving a tracking AF processing function for any subject area. Furthermore, the control device 106 can perform AE processing based on brightness information of the focus detection area and image processing (e.g., gamma correction processing and AWB (auto white balance) adjustment processing) based on pixel values ​​of the focus detection area. The control device 106 also controls the display of the input / output device 105. For example, the control device 106 superimposes an indicator indicating the current position of the subject area (e.g., a rectangular frame surrounding the area) on the displayed image.

[0021] The input / output device 105 can detect the following five states (operations) on a touch panel, which is an input device. Touchdown: A finger or pen that has not been touching the touch panel touches the touch panel again (i.e., the start of a touch). Touch on: Touch panel with your finger or pen. Touch Move: Touching the touch panel with your finger or pen and moving it. Touch up: When you release your finger or pen from the touch panel (i.e., the end of touch). Touch-off: When nothing is touching the touch panel.

[0022] When touch-down is detected, touch-on is also detected at the same time. After touch-down, touch-on will usually continue to be detected unless touch-up is detected. Touch-move is also detected when touch-on is detected. Even if touch-on is detected, touch-move will not be detected unless the touch position moves. Once it is detected that all fingers or pens that were touching have touched up, touch-off occurs.

[0023] These operation states and the position coordinates of the finger or pen touching the touch panel are notified to the control device 106 via the internal bus. Based on the notified information, the control device 106 determines what kind of touch operation the user performed on the touch panel.

[0024] <Functional configuration> 2 is a diagram showing the functional configuration of the camera system. Here, functions corresponding to the image processing device 104, input / output device 105, and control device 106 are shown. The camera system includes an image input unit 210, a detection unit 220, a selection unit 230, a superimposition unit 240, an image display unit 250, an operation acquisition unit 260, a selection unit 270, an integration unit 280, a selection unit 285, and a tracking unit 290.

[0025] The image input unit 210 inputs time-series moving images captured by the imaging device 101 to the image processing device 104. For example, the image input unit 210 inputs frame images constituting a full HD (1920×1280 pixels) moving image in real time (60 frames per second).

[0026] The detection unit 220 processes the image input by the image input unit 210 to detect object candidates. For example, the object candidates are detected by estimating the detection area of ​​the object. The detection area is estimated based on the image coordinate values ​​of the center of the frame, the width and height of the frame, and a likelihood indicating the probability of the existence of the object.

[0027] The selection unit 230 selects one frame that has a high likelihood and is close to the center of the image from the object candidates detected by the detection unit 220, and sets this as the first detection frame selection result. The selection unit 270 selects a combination of object candidates from the object candidates detected by the detection unit 220 and the information acquired by the operation acquisition unit 260. The integration unit 280 integrates the object candidates based on the combination of object candidates selected by the selection unit 270. The selection unit 285 selects one integrated frame from the integration result obtained by the integration unit 280. Details will be described later with reference to FIG. 3.

[0028] Superimposing unit 240 superimposes the image input by image input unit 210 on an object frame selected by selection unit 230 or selection unit 285, or an object frame that is the processing result of tracking unit 290. Image display unit 250 displays the image superimposed by superimposing unit 240. Operation acquisition unit 260 acquires an operation input by the user for the image displayed on image display unit 250.

[0029] Tracking unit 290 performs tracking processing based on the image input by image input unit 210 and the object candidates obtained by selection unit 230 or selection unit 285. In addition, tracking unit 290 outputs an object frame, which is the processing result, to superimposition unit 240.

[0030] <Camera system operation> 3 is a flowchart illustrating the processing performed by the camera system during shooting. More specifically, it illustrates the operations performed by the camera system when selecting a frame of interest from a captured video and performing tracking and AF processing. Note that the camera system does not necessarily have to perform all of the steps described in this flowchart.

[0031] In S400, the image input unit 210 inputs an image from a time-series video captured by the imaging device 101 to the detection unit 220. The image acquired in S400 is, for example, bitmap data expressed in 8 bits each for RGB. In S401, the detection unit 220 processes the image input by the image input unit 210 and detects object candidates.

[0032] 8A and 8B are diagrams showing examples of an input image and the detection results of object candidates. Fig. 8A shows an image 800 input by the image input unit 210 and displayed on the input / output device 105. The image 800 includes a formula car 810, which is an unspecified object candidate. Fig. 8B shows detection frames 820 to 832 corresponding to the object candidates detected by the detection unit 220 for the image 800.

[0033] In this embodiment, object candidate detection is achieved using a neural network. Figure 4 is a diagram illustrating the structure of a neural network. The neural network has a network structure used in object detection described in any of Non-Patent Documents 1 to 3. Such a network outputs intermediate features by inputting an image into a network called a backbone. The features obtained through the backbone are input into networks separated by tasks for estimating the object position and object frame of an object (such as a vehicle or an animal). The network shown in Figure 4 obtains a "center map" indicating the center position of the object and two "size maps" indicating the width and height of the frame (object frame) surrounding the object. Each map is a two-dimensional array represented by a grid. The center map infers a likelihood representing the center position of the object from the array.

[0034] FIG. 15 shows examples of a center map and a size map. FIG. 15(a) shows a center map 1500. In the center map 1500, the likelihood magnitudes of a chair, a person's face, a car, a light, and a tire are indicated by black circles 1501 to 1505. FIG. 15(b) shows a size map 1506 representing the width (horizontal size) of an object. In the size map 1506, the widths of the chair, a person's face, a car, a light, and a tire are indicated by double-headed arrows 1507 to 1511. FIG. 15(c) shows a size map 1512 representing the height (vertical size) of an object. In the size map 1512, the heights of the chair, a person's face, a car, a light, and a tire are indicated by double-headed arrows 1513 to 1517.

[0035] The center map indicates that the closer to the center of the black circle, the higher the likelihood of the corresponding object (and each location). The size map consists of two maps, one for width and one for height, and infers the width and height of the object when that position is taken as the center of the object (and each location). The size map represents the magnitude of the value with the length of the double-headed arrow, indicating that values ​​indicating width and height are inferred at the center position of the object (and each location).

[0036] An object frame is defined by the center coordinates, width, and height of a rectangle that surrounds an object in the image. The center map estimates the likelihood that it is the center of the object. A threshold is set in advance for the likelihood, and elements with values ​​exceeding the threshold are obtained as candidates for the center of the object. If center position candidates are obtained for multiple adjacent elements, the element with the higher likelihood is determined to be the center of the object. Because the resolution of the center map is lower than that of the original image, the center position of the object on the image can be obtained by scaling the center position obtained from the center map to the image size. In addition, the width and height of the frame surrounding the object can be obtained from the element in the size map that corresponds to the detected object center position, and the object frame (detection frame) can be obtained.

[0037] In S402, the selection unit 230 determines whether a tracking template has been set in accordance with a control signal from the control device 106. A tracking template refers to an object frame used in tracking processing. The tracking processing method will be described later. If it is not determined in S402 that a tracking template has been set, the process proceeds to S403. If it is determined in S402 that a tracking template has been set, the process skips S403 to S405 and proceeds to S406.

[0038] In S403, selection unit 230 selects one detection result from the object candidates detected by detection unit 220. Here, the detection result refers to the detection frame obtained in S401. In order to improve the visibility of the object and detection frames in the image, in S403, the detection frames to be displayed on input / output device 105 are narrowed down to one of detection frames 820 to 832. Figure 5 is a detailed flowchart of the first detection frame selection (S403).

[0039] In S500, the selection unit 230 selects only detection frames whose distance from the image center is equal to or less than a preset threshold for the object positions of the object candidates obtained by the detection unit 220 in accordance with a control signal from the control device 106. The reason for selecting object candidates near the image center is to automatically select object candidates using only the framing of the camera system, without any user input operation. In S501, the selection unit 230 selects one detection frame with the highest likelihood in the center map from the one or more detection frames selected in S500, and initially sets it as a tracking template (initial region of interest).

[0040] FIG. 9 is a diagram showing an example of the results of the first detection frame selection. Dashed circle 900 represents the distance threshold from the image center. Detection frame 826 indicates a detection frame selected by the above-mentioned selection method whose distance from the image center is equal to or less than the threshold and whose likelihood is the highest. Frames 821 and 825 are detection frames not selected by the above-mentioned selection method (i.e., detection frames whose distance from the image center is equal to or less than the threshold or whose likelihood is not the highest). Note that, regardless of the above-mentioned selection method, one detection frame may be selected based on a selection instruction from the user. In general, it is desirable for a tracking template to be a frame that surrounds the entire subject and can capture the features of the subject. However, detection frame 826 with the highest likelihood corresponds to a part of the body of formula car 810 and is not suitable as a tracking template. Therefore, the tracking template is modified in steps S407 to S415, which will be described later.

[0041] In S404, selection unit 230 determines whether a selected detection frame exists in accordance with a control signal from control device 106. While FIG. 9 illustrates an example in which there is an object candidate, there may not be a detection frame if an image of only the background or a uniform image is input. If it is determined that a selected detection frame exists, the process proceeds to S405. If it is not determined that a selected detection frame exists, the process skips S405 and proceeds to S406. In S405, selection unit 230 sets the selected detection frame as a tracking template.

[0042] In S406, the operation acquisition unit 260 determines whether or not the image display unit 250 has detected a user input to the image displayed on the input / output device 105. Specifically, the operation acquisition unit 260 acquires input operation information by the user from the control device 106 and determines whether or not a touch-down has been detected. If it is determined that a touch-down has not occurred, the process skips S407 to S415 and proceeds to S416. If it is determined that a touch-down has occurred, the process proceeds to S407.

[0043] In S407, the operation acquisition unit 260 stores the image obtained by the image input unit 210 and the detection frame obtained by the detection unit 220 in the RAM 102. In S408, the image display unit 250 displays the image stored in S407 on the input / output device 105. If the moving image obtained from the image input unit is displayed on the image display unit 250 as is, the object to be tracked will move, making it difficult for the user to select the object to be tracked. For this reason, it is preferable to store the image at the time when touch-down is detected (the start of touch) and control the display so that this image continues to be displayed as a still image. This allows the user to easily select the object to be tracked and input a trajectory, which will be described later.

[0044] In S409, the operation acquisition unit 260 determines whether the end of the user input has been detected in accordance with a control signal from the control device 106. Specifically, it determines whether a touch-up has been detected, and if a touch-up has not been detected, the image stored in S407 continues to be displayed on the input / output device 105. If a touch-up is detected in S409, the process proceeds to S410. In S410, the operation acquisition unit 260 generates a frame surrounding the user input. In this embodiment, the user input is information including a series of coordinates (trajectory) input by the user with a touch-move. Hereinafter, the frame surrounding the user input (a rectangular area that includes the entire trajectory) will be referred to as the trajectory frame. Furthermore, the area within the trajectory frame will be referred to as the trajectory area.

[0045] Fig. 10 is a diagram illustrating the setting of a trajectory frame based on the trajectory of a user input. In Fig. 10(a), an arrow 1000 represents the trajectory input by the user with a touch-move. Fig. 10(b) shows a touch panel 1002 of the input / output device 105 and a finger 1010, with position 1020 being the touch-down position and position 1030 being the touch-up position. Fig. 10(c) shows a trajectory frame 1050.

[0046] As shown in FIG. 10(b), the user touches down on the touch panel at position 1020 with finger 1010, moves to position 1030 with a touch-move, and touches up at position 1030. The operation acquisition unit 260 generates a trajectory frame 1050 that surrounds the trajectory 1000 input by the user, as shown in FIG. 10(c). Note that the trajectory frame does not need to be a frame that surrounds the entire trajectory, and the coordinate position and size of the trajectory may be corrected taking into account the user's intention and input error. For example, the coordinates of the trajectory itself may be used. Alternatively, an area of ​​any shape surrounded by a touch-move, as described below, may be used as the trajectory.

[0047] FIG. 16 is a diagram illustrating another example of setting a trajectory frame based on a trajectory of a user input. In FIG. 16(a), an arrow trajectory 1600 imitates a trajectory input by a user with a touch-move, and a trajectory area 1605 is an area surrounded by the trajectory input by a touch-move. FIG. 16(b) shows a finger 1610, a touch-down position 1620, and a touch-up position 1630. As shown in FIG. 16(b), the user touches down on the touch panel with the finger 1610 at position 1620, moves to position 1630 with a touch-move, and touches up at position 1630. If the control device 106 determines that the trajectory of the touch-move is closed, the control device 106 performs processing similar to the processing for the trajectory frame described below on the inside of the closed area.

[0048] In S411, the selection unit 270 selects a combination of object candidates from the object candidates detected by the detection unit 220 and the information acquired by the operation acquisition unit 260. The combination of object candidates includes two or more object candidates. Fig. 6 is a detailed flowchart of the second detection frame selection (S411).

[0049] In S600, the selection unit 270 acquires, in accordance with a control signal from the control device 106, two or more detection frames (referred to as trajectory overlapping frames) that have overlapping portions with the trajectory acquired by the operation acquisition unit 260, from among the detection frames detected by the detection unit 220. For example, it determines whether the coordinates of the trajectory on the image overlap with the coordinates of the area of ​​the detection frames.

[0050] In S601, the selection unit 270 generates a combination of detection frames from one or more trajectory overlapping frames obtained in S600. All combinations are generated as the combinations. However, to speed up the process, frames with likelihoods equal to or less than a preset threshold may be excluded from the trajectory overlapping frames.

[0051] Fig. 11 shows an example of generating a combination of frames. Fig. 11(a) shows an example in which detection frames 821 to 826 and 828 are selected and combined from the detection frames detected in S401 shown in Fig. 8(b). Fig. 11(b) shows an example in which detection frames 820 to 826 are selected and combined.

[0052] In S412, the integration unit 280 integrates multiple detection frames into one frame based on the combination of trajectory overlapping frames selected by the selection unit 270. In this embodiment, a rectangular frame (called an integrated frame) that surrounds the entire area of ​​the trajectory overlapping frames is generated. The area within the integrated frame is called an integrated area.

[0053] Figure 12 is a diagram showing an example of generating an integrated frame. In Figure 12(a), frame 1201 is the integrated frame of the combination example shown in Figure 11(a). In Figure 12(b), frame 1210 is the integrated frame of the combination example shown in Figure 11(b). For example, the minimum x-coordinate, y-coordinate, and maximum x-coordinate, y-coordinate on the image coordinates of the trajectory overlapping frames included in the combination are calculated, and an integrated frame is generated based on the calculated coordinates.

[0054] In S413, the selection unit 285 selects one integrated frame from the integration results obtained by the integration unit 280. Fig. 7 is a detailed flowchart of the integrated frame selection (S413).

[0055] In S700, the selection unit 285 calculates the overlapping degree of each of the multiple integrated frames generated in S412 and the trajectory frame generated in S410. As the overlapping degree, for example, the ratio of the area of ​​the intersection (overlapped area) of two regions of interest to the area of ​​the union of the two regions (IoU: Intersection over Union) is calculated.

[0056] Figure 13 is a diagram explaining the overlapping degree of the integrated frame and the trajectory frame. In Figure 13(a), frame 1201 is the integrated frame of the combination example shown in Figure 12(a), and frame 1050 is the trajectory frame generated in Figure 10(c). In Figure 12(b), frame 1210 is the integrated frame of the combination example shown in Figure 12(b), and frame 1050 is the trajectory frame generated in Figure 10(c).

[0057] In S701, the selection unit 285 selects an integrated frame with the maximum IoU, which is the degree of overlap. For example, since the overlap between the integrated frame and the trajectory frame in FIG. 13(b) is greater than the overlap between the integrated frame and the trajectory frame in FIG. 13(a), the integrated frame in FIG. 13(b) is selected. Furthermore, a threshold may be set in advance, and if the IoU does not exceed the threshold, a frame not intended by the user may not be selected.

[0058] In S414, the superimposing unit 240 determines whether a selected integrated frame exists in accordance with a control signal from the control device 106. If it is determined that a selected integrated frame exists, the process proceeds to S415. If it is not determined that a selected integrated frame exists, the process skips S415 and proceeds to S416. In S415, the selecting unit 285 updates the selected integrated frame as a tracking template.

[0059] In S416, the image display unit 250 displays the superimposed image on the input / output device 105. Specifically, the image input by the image input unit 210 in S400, the detection frame selected by the selection unit 230 in S403, the integrated frame selected in S413, and the frame subjected to tracking processing in S418 (described later) are superimposed and displayed. Note that if frames are selected in both S403 and S413, it is preferable to give priority to displaying the integrated frame selected in S413. Also, if an integrated frame is selected in S413 and tracking processing has been performed on the previous frame in S418 (described later), it is preferable to give priority to displaying the integrated frame selected in S413.

[0060] In S417, the tracking unit 290 determines whether a tracking template is set in accordance with a control signal from the control device 106. If it is determined in S417 that the template is set as a tracking template, the process proceeds to S418. If it is not determined in S417 that the template is set as a tracking template, the process skips S418 and S419 and proceeds to S420.

[0061] In S418, the tracking unit 290 performs tracking processing based on the image and object candidates obtained by the image display unit 250. As a tracking processing method, template matching is applied to search for an area that has a high similarity to the template. For example, the method described in Japanese Patent Laid-Open No. 2020-21250 (Patent Document 2) can be used.

[0062] In S419, the tracking unit 290 performs AF processing on the area tracked by the tracking unit 290 in accordance with a control signal from the control device 106. As an AF processing method, for example, a phase difference detection AF method described in Patent Document 2 can be used.

[0063] In S420, tracking unit 290 determines whether to continue the tracking process in accordance with a control signal from control device 106. If it is determined that the tracking process should be continued, the process returns to S400.

[0064] In the above description, object candidate detection is performed using a rectangular frame and used as a tracking template, but object candidate detection may also be performed using an area of ​​any shape instead of a rectangular frame and used as a tracking template.

[0065] As described above, according to the first embodiment, a tracking template (an attention area corresponding to a tracking target object) is set by using the detection frame set by the detection unit and the trajectory of the user's touch operation. In particular, the user can modify the tracking template to a more appropriate size by simply performing a simple operation (touch operation) on the detection frame.

[0066] (Variation 1) In the first embodiment described above, in the second detection frame selection (S411), an integrated frame is generated for all frame combinations from among the trajectory overlapping frames. However, generating an integrated frame for all frame combinations incurs a large calculation cost, including the calculation of the degree of overlap with subsequent trajectory frames. Furthermore, similar integrated frames are often generated, making the process redundant.

[0067] Furthermore, when an object in front of the imaging device 101 and an object behind it are close to each other on the image (when they are far apart on the Z-axis (depth direction) coordinate but close to each other on the X- and Y-axis coordinates of the image), there is a risk that inappropriate frames will be combined. Therefore, in Modification 1, an example will be described in which the selection unit 270 generates a combination of frames using distance information on the Z-axis (depth direction) coordinate of the frame.

[0068] In Modification 1, a method of calculating distance information from parallax images is applied as a method of acquiring distance information. For example, the method described in JP 2019-126091 A (Patent Document 3) can be used. Parallax images acquired by the imaging device 101 are stored in RAM 102, and distance information (depth information) is calculated based on the parallax images by the control device 106 and used as information for generating frame combinations. Of course, distance information (depth information) may be acquired using other methods.

[0069] Fig. 14 is a flowchart illustrating the processing of the selection unit 270 in Modification 1. Note that reference numerals are assigned to correspond to the flowchart of the first embodiment in Fig. 6. Steps S1400 and S1401 are the same processing as steps S600 and S601 in Fig. 6.

[0070] In S1402, the selection unit 270 acquires distance information of the area within the trajectory overlap frame detected by the detection unit 220 from the RAM 102 in accordance with a control signal from the control device 106. For example, the distance information of the center position with the highest likelihood is used. Alternatively, the average value of the distance information within the area or the average value of the distance information multiplied by the likelihood may be used.

[0071] In S1403, selection unit 270 calculates the difference in distance information between each of the trajectory overlapping frames. For example, if detection frames 820, 821, and 831 are trajectory overlapping frames in Fig. 8, the distance between detection frame 820 and detection frame 821 is short, and the distance between detection frame 820 and detection frame 831 and the distance between detection frame 820 and detection frame 831 are long.

[0072] In S1404, the selection unit 270 excludes from the combination any frame whose difference from the average value of the distance between the respective trajectory overlapping frames is equal to or greater than a predetermined threshold. Here, detection frame 831 is excluded from the combination generated in S1401. This is because it can be assumed that the distance information of each trajectory overlapping frame corresponding to a tracking target object (one object) is similar, and that trajectory overlapping frames whose distance information is not similar can be assumed to represent a different object. Note that, instead of the average value of the distance information of each trajectory overlapping frame, the distance information of the trajectory overlapping frame with the highest likelihood can be considered as reference distance information, and frames whose difference from that reference distance information is equal to or greater than a predetermined threshold can be excluded from the combination.

[0073] As described above, in Modification 1, trajectory overlapping frames whose difference from the average distance is equal to or greater than a preset threshold are excluded from among multiple trajectory overlapping frames. This reduces the number of frame combinations and makes it possible to reduce calculation costs.

[0074] (Variation 2) In the first embodiment, first, a tracking template based on one detection frame selected from near the center of the image is set as an initial template (S402 to S405). After that, when a tracking template based on an integrated frame generated based on user input is set (S406 to S415), the tracking template based on the integrated frame is used. However, a configuration may be adopted in which setting of a tracking template is not performed in S402 to S405. In other words, a tracking template based on an integrated frame generated based on user input may be set as the initial template.

[0075] (Other Examples) The present invention can also be realized by supplying a program that realizes one or more functions of the above-described embodiments to a system or device via a network or a storage medium, and having one or more processors in the computer of the system or device read and execute the program.The present invention can also be realized by a circuit (e.g., ASIC) that realizes one or more functions.

[0076] The invention is not limited to the above-described embodiments, and various changes and modifications can be made without departing from the spirit and scope of the invention. Accordingly, the following claims are appended to apprise the public of the scope of the invention. [Explanation of symbols]

[0077] 210 image input unit; 220 detection unit; 230 selection unit; 240 superimposition unit; 250 image display unit; 260 operation acquisition unit; 270 selection unit; 280 integration unit; 285 selection unit; 290 tracking unit

Claims

1. image input means for inputting an image; a detection means for detecting object candidates from the image; a receiving means for receiving an input of a trajectory for the image; a selection means for selecting, based on a trajectory region determined by the trajectory, two or more object candidates that are included in the plurality of object candidates detected by the detection means and that are parts of the same object; an integration means for generating an integrated region based on two or more regions in the image corresponding to two or more object candidates selected by the selection means, and setting the integrated region as a region of interest corresponding to the tracking target object in the image; An image processing device comprising:

2. The receiving means is configured as a touch display that displays the image and receives input of the trajectory on the image by a touch operation.

2. The image processing device according to claim 1, wherein:

3. an initial setting means for selecting one object candidate from the plurality of object candidates detected by the detection means and setting it as an initial region of interest in the image; an update means for updating the initial region of interest with the integrated region when the integrated region is generated by the integration means; 3. The image processing device according to claim 1, further comprising:

4. The initial setting means selects, from among the object candidates detected by the detection means, an object candidate whose distance from the center of the image is equal to or less than a preset threshold and whose likelihood is the greatest.

4. The image processing device according to claim 3.

5. The initial setting means selects one detection frame from among the object candidates detected by the detection means based on a selection instruction from a user.

4. The image processing device according to claim 3.

6. The selection means determining a plurality of combinations including two or more object candidates each having an overlapping portion with the trajectory region from among the plurality of object candidates detected by the detection means; selecting one combination from the plurality of combinations that maximizes the ratio of the area of ​​the intersection of an integrated region obtained by integrating two or more regions corresponding to two or more object candidates included in the combination and the trajectory region to the area of ​​the union of the integrated region and the trajectory region; Select two or more object candidates included in the selected combination.

6. The image processing device according to claim 1, wherein the image processing device is a computer.

7. The selecting means further excludes from the plurality of combinations one or more combinations including an object candidate whose difference from an average value of distance information corresponding to two or more object candidates included in the combination is equal to or greater than a predetermined threshold.

7. The image processing device according to claim 6,

8. The locus area is a rectangular area that encompasses the locus.

8. The image processing device according to claim 1, wherein the image processing device is a computer.

9. The locus area is an area of ​​any shape surrounded by the locus.

8. The image processing device according to claim 1, wherein the image processing device is a computer.

10. an imaging means for capturing images and generating moving images; An image processing device according to any one of claims 1 to 9; a tracking means for tracking an object candidate included in the video image according to a region of interest set by the image processing device; An imaging device comprising:

11. A control method for an image processing device that sets an attention area corresponding to a tracking target object in an image, comprising: an image input step of inputting an image; a detection step of detecting object candidates from the image; a receiving step of receiving an input of a trajectory for the image; a selection step of selecting, based on a trajectory region determined by the trajectory, two or more object candidates that are included in the plurality of object candidates detected by the detection step and are parts of the same object; an integration step of generating an integrated region based on two or more regions in the image corresponding to the two or more object candidates selected by the selection step, and setting the integrated region as the region of interest; A control method comprising:

12. A program for causing a computer to function as each of the means of the image processing device according to any one of claims 1 to 9.

Citation Information

Patent Citations

  • Video recording method and device, electronic equipment and readable storage medium

    CN112714253A

  • Method and apparatus for object size adjustment on a screen

    EP2631778A2

  • JP1973049163A

  • washer

    JP1988097454A

  • Information processing device, information processing method, and program

    WO2021145071A1