Image processing apparatus, image processing method, and program
The image processing apparatus addresses the challenge of detecting intended subject regions by integrating object region candidates using likelihood maps and region tensors, enhancing autofocus accuracy and handling multiple subjects effectively.
Patent Information
- Application Number
- JP2022037599
- Authority / Receiving Office
- JP · JP
- Patent Type
- Patents
- Current Assignee / Owner
- Filing Date
- 2022-03-10
- Publication Date
- 2025-06-30
- Estimated Expiration
- 2042-03-10
AI Technical Summary
Existing image processing technologies struggle to accurately detect a subject region intended by a user, especially when multiple subjects of the same category are present, leading to difficulties in autofocus and feature extraction.
An image processing apparatus that includes an image acquisition unit, an instruction reception unit, a likelihood map acquisition unit, and a determination unit, which uses region tensors and likelihood maps to determine the object region corresponding to the user's instruction by integrating object region candidates.
The apparatus effectively detects the subject region intended by the user, improving autofocus accuracy and overcoming challenges posed by multiple subjects of the same category.
Smart Images

Figure 0007700068000013 
Figure 0007700068000014 
Figure 0007700068000015
Abstract
Description
Technical Field
[0001] The present invention relates to an image processing technology for detecting a subject.
Background Art
[0002] Object detection is one of the fields of computer vision research and has been widely studied so far. Computer vision is a technology that understands an image input to a computer and recognizes various characteristics of the image. Among them, object detection is a task of estimating the position and type of an object existing in a natural image. In Non-Patent Document 1, a likelihood map indicating the center of an object is obtained by using a multi-layer neural network, and the center position of the object is detected by extracting the peak point of the likelihood map. In addition, by inferring the offset amount corresponding to the center position and the object size, the frame of the object to be detected can be obtained.
[0003] Object detection can be applied to the autofocus function of an imaging device. In Patent Document 1, by receiving the designated coordinates of the user and inputting them together with an image into a multi-layer neural network, the main subject according to the user's intention is detected, and the autofocus function is realized. In Patent Document 1, in addition to the likelihood map in the multi-layer neural network, a position map is generated based on a two-dimensional Gaussian spreading from the designated coordinates. Further, by integrating the position map and the likelihood map in the multi-layer neural network, the main subject is detected. When there is a peak of the likelihood map near the designated coordinates, the contribution degree of the position map in the integration process is increased, and when not, it is decreased. In Patent Document 1, furthermore, using imaging information such as the electronic zoom ratio and the amount of camera shake, the spread of the Gaussian when generating the position map from the designated coordinates is adjusted. For example, when the amount of camera shake is acquired as imaging information, it is considered that it is difficult to indicate the coordinates of the subject when the amount of camera shake is large, so the spread of the position map is increased.
Prior Art Documents
Patent Documents
[0004]
Patent Document 1
Non-Patent Document
[0005]
Non-Patent Document 1
Summary of the Invention
Problems to be Solved by the Invention
[0006] As described above, Patent Document 1 discloses a method for detecting a main subject based on a user's intention. However, in Patent Document 1, when there are multiple subjects of the same category, it is difficult to autofocus on the subject as intended by the user. For example, consider the case where subjects of the same category are located in the front and back. At this time, assume that the rear subject partially overlaps the front subject and is partially hidden. If the user has specified the rear subject in this case, the features of the rear subject cannot be well extracted, the reliability of the position map becomes low, and there is a risk that the rear subject cannot be correctly detected. Or, the reliability of the position map for the front subject from which features can be well extracted becomes high, and there is a risk that the front subject becomes the main subject. Also, in Patent Document 1, when the features of the main subject include a dog's face and a dog's body, it is necessary to prepare a main subject detection unit that responds to each of them. There are innumerable objects that can be the main subject, and it is difficult to prepare a main subject detection unit for all of them in advance.
[0007] As described above, consider applying Non-Patent Document 1 to the case where subjects of the same category are located in the front and back, and using the detection result closest to the specified coordinates. It is difficult to separate and infer the likelihood map indicating the center for the front subject and the rear subject, and there is a high possibility that a peak of the likelihood map will appear in the front subject.
[0008] The present invention has been made in view of the above problems, and an object thereof is to provide an image processing apparatus capable of detecting a subject region intended by a user. Another object is to provide a method and a program therefor.
Means for Solving the Problems
[0009] The image processing apparatus according to the present invention has the following configuration.
[0010] That is, an image acquisition unit that acquires a captured image, an instruction reception unit that receives an instruction for the captured image acquired by the image acquisition unit, and the captured image At each position of a likelihood map acquisition unit that represents the presence likelihood of an object, At each position of the captured image, region acquisition means for acquiring a region tensor indicating the position of the object center and the object size based on each position; the position of the instruction received by the instruction reception unit Obtained using the region tensor corresponding to coordinates located within a predetermined range from an image processing apparatus including a determination unit that determines an object region corresponding to the instruction using one or more object region candidates and the likelihood map.
Effects of the Invention
[0011] According to the present invention having the above configuration, it is possible to provide an image processing apparatus capable of detecting a subject region intended by a user.
Brief Description of the Drawings
[0012]
Figure 1
Figure 2
Figure 3
Figure 4
Figure 5
Figure 6
Figure 7
Figure 8
Figure 9
Figure 10
Figure 11
Figure 12
Mode for Carrying Out the Invention
[0013] Hereinafter, with reference to the accompanying drawings, the present invention will be described in detail based on its preferred embodiments. Note that the configurations shown in the following embodiments are merely examples, and the present invention is not limited to the illustrated configurations.
[0014] <Embodiment 1> FIG. 1 shows a schematic configuration diagram of the image processing apparatus according to the present embodiment. Using FIG. 1, the configuration of the present embodiment will be described. Here, only an overview will be described, and details will be described later.
[0015] The imaging device 110 is configured using an optical system, an image sensor, etc., captures an image, and outputs it to the image acquisition unit 101. For example, it is conceivable to use an imaging device such as a digital camera or a surveillance camera. Further, the imaging device 110 has an interface for receiving an input from the user, and outputs the input from the user to the instruction reception unit 104. For example, it is conceivable to mount a touch display as the interface and output the touch operation result from the user to the instruction reception unit 104.
[0016] The image processing apparatus 100 includes an image acquisition unit 101, a likelihood map acquisition unit 102, an estimation unit 103, an instruction reception unit 104, an object region candidate selection unit 105, and an object region candidate integration unit 106. The image acquisition unit 101 receives an image from the imaging device 110. The image is input to the likelihood map acquisition unit 102 and the estimation unit 103, and a likelihood map and an object region candidate are acquired respectively. The instruction reception unit 104 acquires a user instruction from the imaging device 110. Specifically, one instruction coordinate input by the user through a touch operation is acquired. The object region candidate selection unit 105 selects one or more object region candidates to be integrated from among the object region candidates based on the likelihood map and the instruction coordinates. The object region candidate integration unit 106 integrates the selected object region candidates to obtain one object region. A part of the functional configuration in the image processing apparatus 100 (for example, the likelihood map acquisition unit 102, the estimation unit 103, the object region candidate selection unit 105, the object region candidate integration unit 106) can also be an image processing system provided by an image processing apparatus on the network. Further, the image processing apparatus 100 is composed of hardware such as a CPU, a ROM, a RAM, and various interfaces shown in the attached drawings.
[0017] The result output unit 120 outputs the object region, and the output object region is applied to, for example, the autofocus function of the camera. For example, several distance measurement points can be sampled from the object region and used for phase difference AF. If the object region intended by the user can be accurately detected, the autofocus accuracy will also be improved.
[0018] FIG. 2 shows a flowchart of the processing in the image processing apparatus 100 of the present embodiment. Hereinafter, it is assumed that the flowchart is realized by the CPU executing a control program. Note that only the outline is described here, and the details will be described later.
[0019] In S200, the imaging device 110 starts shooting. In S201, the image acquisition unit 101 of the image processing device 100 converts the captured image to a predetermined resolution. Based on the converted image, in S202, the likelihood map acquisition unit 102 acquires a likelihood map, and in S203, the estimation unit 103 acquires object region candidates. In S204, the instruction reception unit 104 receives an instruction of coordinates from the user, determines whether there is an instruction in S205, and if there is no instruction, proceeds to S210 and outputs that there is no object region. If there is an instruction, it proceeds to S206. In S206, the received instruction coordinates are converted to correspond to the resolution converted in S201. Based on the converted instruction coordinates, in S207, the object region candidate selection unit 105 selects object region candidates. The detailed flow of S207 will be described later with reference to FIG. 5. As a result of S207, it is determined in S208 whether an object region candidate has been selected. If not selected, it proceeds to S210 and outputs that there is no object region. If selected, in S209, the object region candidate integration unit 106 integrates the one or more selected object region candidates, and in S210, the result output unit 120 outputs one object region. In FIG. 2, S207 to S209 are described separately, but S207 to S209 are a series of steps for determining an object region.
[0020] <Image Conversion> The image conversion shown in S201 of FIG. 2 will be described. In the present embodiment, first, shooting is started at S200 to acquire a captured image. The captured image is, for example, an RGB image with a width of 6000 pixels and a height of 4000 pixels. In S201, the captured image is converted to a predetermined size to match the input format of a multi-layer neural network for acquiring a likelihood map and object region candidates. In the present embodiment, the input size of these multi-layer neural networks is an RGB image with a width of 500 pixels and a height of 400 pixels. In the present embodiment, 500 pixels on the left and right of the captured image are cropped and reduced to one-tenth, but black images with a height of 400 pixels may be padded above and below the captured image and reduced to one-twelfth, or a region with a width of 500 pixels and a height of 400 pixels may be directly cropped from the captured image. The converted image has the origin at the pixel in the upper left corner, and its coordinates are set to (0, 0). The coordinates (i, j) indicate the coordinates of the j-th row and i-th column, and the pixel in the lower right corner has the coordinates (499, 399). Hereinafter, the coordinate system of the converted image is referred to as the image coordinate system.
[0021] <Acquisition of Likelihood Map> The acquisition of the likelihood map shown in 102 of FIG. 1 and S202 of FIG. 2 will be described with reference to FIG. 3. In the present embodiment, as in Non-Patent Document 1, the likelihood map is acquired by a multi-layer neural network. The input of the multi-layer neural network is an image with the resolution converted, which is an RGB image with a width of 500 pixels, a height of 400 pixels, and 3 channels. The output of the multi-layer neural network is a tensor (matrix) with 10 columns, 8 rows, and 1 channel. The acquired tensor (matrix) is referred to as the likelihood map. The likelihood map (the first channel thereof) has the origin at the pixel in the upper left corner, and its coordinates are set to (0, 0). The coordinates (i, j) indicate the coordinates of the j-th row and i-th column, and the pixel in the lower right corner has the coordinates (9, 7). Hereinafter, the coordinate system of the likelihood map is referred to as the map coordinate system. The likelihood map can also be acquired by operations within the image processing apparatus 100, or the likelihood map can be calculated outside the image processing apparatus 100 and acquired by the likelihood map acquisition unit 102 within the image processing apparatus 100.
[0022] The multi-layer neural network that acquires the likelihood map is pre-trained using a large number of training data (pairs of images and likelihood maps). For details, refer to Non-Patent Document 1. In this embodiment, the likelihood map assumes a saliency map that responds to all objects, but it may also respond only to specific objects. A saliency map is a map that responds to parts where a person is likely to gaze.
[0023] An example of the captured image converted into an image is shown in Fig. 3(a), and an example of the corresponding likelihood map is shown in Fig. 3(b). The captured image 300 converted into an image shows two subjects, a rear subject 301 and a front subject 302. Each element of the likelihood map indicates the presence likelihood of the object at the location corresponding to that element. The presence likelihood takes a value from 0 to 255, and the larger the value, the higher the presence likelihood. Each element in Fig. 3(b) is shown whiter as the likelihood is lower and blacker as the likelihood increases. The likelihood map 304 is obtained with a particularly high likelihood for the front subject 302 and has a maximum likelihood of 204 at (6, 4).
[0024] <Acquisition of object region candidates> The estimation of the object region candidates shown in 103 of FIG. 1 and S203 of FIG. 2 will be described with reference to FIG. 4. Similar to the likelihood map, the object region candidates are obtained using a multi-layer neural network. The input to the multi-layer neural network is an image with resolution conversion, similar to the acquisition of the likelihood map, which is an RGB image with a width of 500 pixels, a height of 400 pixels, and 3 channels. The output of the multi-layer neural network is a tensor with 10 columns, 8 rows, and 4 channels. The first channel of the tensor indicates the offset amount in the x direction from the element to the object center, and the second channel similarly indicates the offset amount in the y direction. The third channel of the tensor indicates the width of the object indicated by the element, and the fourth channel similarly indicates the height. From the above four-channel information, the center position and size of the object can be obtained. This tensor is called the object region candidate tensor. In this embodiment, each channel of the object region candidate tensor has the same number of rows and columns as the likelihood map, and their coordinate systems are also the map coordinate systems. The number of columns and rows of each channel of the object region candidate tensor may be different from those of the likelihood map. If they are different, the number of rows and columns may be made to match by interpolation (for example, bilinear interpolation). The object region candidates can be obtained by operations within the image processing apparatus 100, or the object region candidates can be calculated outside the image processing apparatus 100 and obtained by the estimation unit 103 within the image processing apparatus 100.
[0025] The multi-layer neural network for obtaining the object region candidate tensor is pre-trained in advance using a large number of training data (pairs of images, offset amounts, and widths and heights), similar to the acquisition of the likelihood map. In this embodiment, a multi-layer neural network that outputs four-channel information simultaneously is used, but four multi-layer neural networks that output one channel at a time may be prepared and their results combined.
[0026] Fig. 4(a) shows an example of a captured image that has been image-converted. Figs. 4(b) to (d) show the offset amount map in the x direction to the object center, the offset amount map in the y direction to the object center, the width map of the object, and the height map of the object, respectively. The numerical values of the elements shown in Figs. 4(b) to (d) are in pixels as the unit. The x-direction offset amount is positive to the right, and the y-direction offset amount is positive downward.
[0027] Here, pay attention to the coordinates (6, 4) where the likelihood is maximized in Fig. 3(b). The elements that are inverted in black and white in Figs. 4(b) to (d) are the points of interest. In Fig. 4(a), the point shown as point 401 is the location on the image corresponding to this point of interest. The x-direction offset amount at the point of interest is -3, and the y-direction offset amount is -2. That is, it is obtained that the object center is in the upper left direction from the point 401, which is the point of interest, at the location of point 402.
[0028] The map coordinates can be converted to image coordinates by Equation 1 below.
[0029]
Equation
[0030] Here, I w , I h represent the width and height of the image-converted captured image, respectively, and M w , M h represent the width and height of the map, respectively. (I x , I y ) represents a point in the image coordinate system, and (M x , M y ) represents a point in the map coordinate system. According to Equation 1, the point (6, 4) in the map coordinates is (325, 225) when converted to image coordinates. That is, the image coordinates of point 401 in Fig. 4 are (325, 225), and the image coordinates of point 402, which is the result of adding the offset amount to this, are (322, 222).
[0031] Also, the width and height of the object at the point of interest are 166 and 348 respectively from FIGS. 4(d) and (e). From the above, the object region candidate 400 at the point of interest is represented as a rectangle centered at (322, 222) with a width of 166 and a height of 348 in the image coordinate system.
[0032] In this embodiment, the object region candidate is the amount of offset in the vertical and horizontal directions and the width and height of the object, but it may be, for example, the distances to the left and right ends and the upper and lower ends.
[0033] <Object Region Candidate Selection> The selection of the object region candidate shown in 105 of FIG. 1 and S207 of FIG. 2 will be described using the detailed flowchart of FIG. 5. First, in the preparation step S500, each variable is initialized. The variables are, respectively, n and m as counters, N as the number of object region candidates to be selected, T as the likelihood threshold, D as the distance threshold, L ij as the likelihood, S ij as the object region candidate, and (P x , P y ) is the coordinate of the coordinate (instruction coordinate) acquired by the instruction reception unit. The instruction coordinate (P x , P y ) is the instruction coordinate given in the image coordinate system converted to the map coordinate system based on Equation 1 and is a two-dimensional real vector. In S501, the map coordinate (u, v) that is the m-th closest to the instruction coordinate (P x , P y ) is selected from all the map coordinates. (u, v) is a two-dimensional positive integer vector. In S502, if (u, v) exists and the distance between (P x , P y ) and (u, v) is acquired, and if the distance is equal to or less than the threshold D, the process proceeds to the next step S503; otherwise, the process ends. In this embodiment, the Euclidean distance (Equation 2) is used as the distance function, but other distance functions may be used.
[0034]
Equation
[0035] In S503, the likelihood map L corresponding to the map coordinate (u, v) uvExtract it. In S504, L uv is compared with the likelihood threshold T, and if L uv is greater than or equal to T, proceed to the next step S505. If L uv is less than T, do not select it as an object region candidate, proceed to S508, increment m by 1, and return to S501. In S505, extract the object region candidate S uv corresponding to the map coordinates (u, v). In S506, save the current likelihood and the object region candidate as L n , S n respectively. In S507, compare n and N. If n is greater than or equal to N, end the process. If n is less than N, in S508, increment n, in S509, increment m by 1, and return to S501. In this embodiment, a predetermined number N of object region candidates are selected, but the method for determining the number of object region candidates to be selected is not limited to this. For example, object region candidates may be selected such that the sum of the likelihoods L n is greater than or equal to a predetermined value.
[0036] <Object Region Candidate Integration> The integration of the object region candidates shown in 106 of FIG. 1 and S209 of FIG. 2 will be described with reference to FIG. 6. Note that in this embodiment, it is assumed that the user selects the rear subject 301. However, even when the front subject 302 is selected, the same processing is performed except that the designated coordinates are different. When the user wants to select the rear subject 301, the user designates a location 600 corresponding to the rear subject 301 as shown in FIG. 6(a). The designated coordinates are (235, 245) in the image coordinate system, and when converted to the map coordinate system by Equation 1, they are (4.2, 4.4). When the above-described object region candidate selection process is performed, the hatched portions shown in 601, that is, the locations corresponding to the closest (4, 4), the second closest (4, 5), and the third closest (5, 4) in the map coordinate system, are selected. Therefore, the likelihoods L1, L2, L3 are respectively L 44 , L 45 , L 54 and the object region candidates S1, S2, S3 are respectively S 44 , S 45 , S 54 are substituted. Regions 602, 603, 604 are respectively S 44 , S 45 , S 54It shows the object region candidates corresponding thereto. 605 to 609 in FIG. 6(b) represent the likelihood, x offset amount, y offset amount, width, and height corresponding to 601 respectively. FIG. 6(c) shows the result of object region candidate integration. 610 is the center position of the integrated object region candidate, and 611 is the object region candidate with width and height added thereto.
[0037] Regarding object region candidate integration, it will be described using a specific calculation example. First, calculate the center position of the object region candidate in the image coordinate system. From FIG. 6(b), the x offset amount of the object region candidate S 44 is -1, and the y offset amount is -41. From Equation 1, the image coordinates corresponding to (4, 4) in the map coordinates are (225, 225). Adding the offset amounts of the object region candidate S 44 (602) to this, the center position (224, 184) of the object region candidate in the image coordinate system is obtained. Similarly, the center position (224, 172) of S 45 (603) and the center position (276, 223) of S 54 (604) can be obtained. The weighted average of likelihoods is used for the integration of object region candidates. The weighted average of likelihoods can be calculated using the following Equation 3.
[0038]
Equation
[0039] x n is the value for which the weighted average is taken, and x is the result of taking the weighted average. For example, when obtaining the x coordinate of the center position of the integrated object region, substitute the x coordinate of the center position of the object region candidate corresponding to S n into x n . Similarly, by substituting the y coordinate, width, and height of the center position into Equation 3, the center position, width, and height of the integrated object region can be obtained.
[0040] Likelihood L nBy assigning 0 to the initial value, even when the number of object region candidates exceeding the likelihood threshold T within the range of the distance threshold D is less than the predetermined number N, the integrated object region can be obtained using Equation 3. Also, when all likelihoods L n are 0, it is considered that there is no object region.
[0041] The above embodiments are summarized. First, image conversion is performed on the captured image so that a likelihood map can be obtained and object region candidates can be obtained. Obtaining the likelihood map and object region candidates is realized by a multi-layer neural network. The object region candidate selection unit selects three object region candidates located near the user's instructed coordinates obtained by the instruction reception unit. For the selected object region candidates, the weighted average using the likelihood as the weight is calculated by the object region candidate integration unit and integrated into a single object region. As a result, even when the likelihood map strongly reacts to the front subject as shown in FIG. 3, the object region 611 for the rear subject as intended by the user is output as shown in FIG. 6(c).
[0042] <Modification 1> In Embodiment 1, when integrating object region candidates, the value of the likelihood map was used, but the distance between the coordinates (instructed coordinates) obtained by the instruction reception unit and the object region candidates may also be used. In Modification 1, the shorter the distance between the instructed coordinates and the object region candidates, the larger the weight used for the weighted average when integrating the object region candidates. Specifically, as shown in Equation 4 below, the reciprocal of the distance between the instructed coordinates (P x , P y ) and the object region candidates is used to calculate the weight of the weighted average.
[0043]
Equation
[0044] Here, (u, v) n is the map coordinate corresponding to the likelihood L n . The Euclidean distance in the map coordinate system used in Equation 2 is used for the distance calculation. Instead of the likelihood L n in Equation 3, the weight W nBy calculating the weighted average using [the method], it becomes possible to integrate the object region candidates considering the distance from the specified position.
[0045] In the above description, in the step of integrating the object region candidates, the weights of the weighted average were recalculated, but a likelihood map considering the distance from the indicated coordinates may be calculated in advance. The likelihood map considering the distance from the indicated coordinates is called the corrected likelihood map K ij is calculated by the following Equation 5.
[0046]
Equation
[0047] In Equation 5, for all elements of the likelihood map L ij a calculation of division by the distance between the indicated coordinates and the element is performed and substituted into the corrected likelihood map K ij .
[0048] According to Modification Example 1, it is possible to integrate the object region candidates while emphasizing the object region candidates closer to the specified coordinates.
[0049] <Modification Example 2> In Modification Example 2, an example of expanding the method of selecting the object region candidates by interpolating each channel of the likelihood map and the object region candidate tensor is shown.
[0050] First, interpolation will be described with reference to FIG. 7. FIG. 7 is an excerpt of a width map (the third channel of the aforementioned object region candidate tensor) of the object region candidate shown in FIG. 4(d). Here, an example of applying bilinear interpolation is shown in the width map 700 excerpted from map coordinates (4, 4) to (5, 5) when the coordinate 701 is given. The interpolation method is not limited to bilinear interpolation, and other interpolation methods represented by nearest-neighbor interpolation or bicubic interpolation may be used. Also, interpolation processing can be similarly applied to ranges other than map coordinates (4, 4) to (5, 5). The values in parentheses shown in each element of the width map 700 in FIG. 7 indicate the map coordinates, and the value shown to the right of the colon indicates the width of the object region candidate at the corresponding coordinate.
[0051] As shown in FIG. 7, let the distances in the x-direction and y-direction between the point 701 and each map coordinate be x1, x2, y1, and y2, respectively. For example, if the map coordinate of the point 701 is (4.2, 4.4), then x1 = 0.2, x2 = 0.8, y1 = 0.4, and y2 = 0.6. Bilinear interpolation is realized as shown in Equation 6 below using the distances in the x-direction and y-direction between the point 701 and each map coordinate.
[0052]
Equation
[0053] S ij is the value at the map coordinate (i, j) to be interpolated, and S is the interpolation result. In the example of FIG. 7, S ij is the width of the object region candidate at the map coordinate (i, j). The interpolation value can also be calculated by performing the same calculation for the height of the object region candidate. When interpolating the offset amount to the object candidate region, it is necessary to convert it to the center position of the object region candidate in advance. For the method of converting the offset amount to the center position, refer to <Object Region Candidate Integration> in Embodiment 1. Also, the interpolation value can be calculated for the likelihood map by the same procedure.
[0054] By using interpolation, likelihoods and object region candidates for all coordinate positions can be obtained. Here, the selection of object region candidates using interpolation will be described with reference to FIG. 8. Consider a case where an imaging image 800 is given and a user wants to select a subject 802, and it is assumed that designated coordinates 801 are given. Consider a plurality of concentric circles (e.g., 803) having different radii that expand around 801. Points (e.g., 804) obtained by dividing each concentric circle by a predetermined number are called a neighborhood point group. In the present embodiment, the object region candidate selection unit selects values obtained by interpolating object region candidates at the designated coordinates 801 and in the neighborhood point group. However, points in the neighborhood point group located outside the range of the map coordinates are not selected.
[0055] Let the number of concentric circles be Nr, the difference in radius between adjacent concentric circles be dr, and the difference in the number of divisions between adjacent concentric circles be dq. These values are set in advance. For example, in FIG. 8, Nr = 3, dr = 0.5, and dq = 4 are set. Considering the nr-th concentric circle from the designated coordinates 801, its radius r nr is given by Equation 7, and the number of divisions q nr is calculated from Equation 8.
[0056]
Equation
[0057] In Equations 7 and 8, nr = 0 indicates the designated coordinates 801, at which time the radius r0 = 0 and the number of divisions q0 = 1. 803 in FIG. 8 is the third concentric circle from the center, with a radius of r3 = 3 and a number of divisions of q3 = 12.
[0058] The likelihoods of the object at the designated coordinates 801 and in the neighborhood point group are obtained by interpolating the likelihood map. Here, an index representing a neighborhood point is defined. Consider the neighborhood point group on the nr-th concentric circle. The neighborhood point located upward is numbered 0, and the numbers are assigned clockwise. The index of the q-th neighborhood point is (nr, q). The index of the neighborhood point located one to the left of the 0-th neighborhood point is (nr, q nr - 1). 804 in FIG. 8 is represented by the index (3, 3).
[0059] By transforming Equation 3 using the above index, it is possible to calculate the weighted average based on the likelihood of the object region candidate in the present embodiment (Equation 9).
[0060]
Equation
[0061] Here, L (nr,q) is the interpolated value of the likelihood at the neighboring point expressed by the index (nr, q). x (nr,q) is the value for which the weighted average is taken at the neighboring point expressed by the index (nr, q), and x is the result of taking the weighted average. For example, when obtaining the x - coordinate of the center position of the integrated object region, the interpolated value of the x - coordinate of the center position of the object region candidate corresponding to the neighboring point expressed by the index (nr, q) may be substituted into x (nr,q) . Similarly, by substituting the y - coordinate, width, and height of the center position into Equation 3, the center position, width, and height of the integrated object region can be obtained.
[0062] By using the radius r nr of the concentric circles, as shown in Modification 1, the distance between the indicated coordinates and the object region candidate can be considered.
[0063]
Equation
[0064] r nr is the radius of the nr - th concentric circle shown in Equation 7. By replacing L (nr,q) in Equation 9 with W (nr,q) in Equation 10, it becomes possible to integrate the object region candidates considering the distance between the indicated coordinates and the object region candidates.
[0065] When the circuit mounted on the imaging device restricts the multi-layer neural network to be lightweight, the resolution of the map coordinates, which is the output, decreases. When the resolution of the map coordinates is low, even if map coordinates near the instruction coordinates are selected, the deviation between the instruction coordinates and the map coordinates increases. In Modification 2, by interpolating each channel of the likelihood map and the object region candidate tensor, it is possible to achieve integration of object region candidates that is independent of the resolution of the map coordinates.
[0066] <Modification 3> The position of the object region determined in the object region candidate integration S209 of Embodiment 1 is calculated from the position of the object candidate region determined by the object region candidate selection. However, depending on the accuracy of the likelihood map, the object region obtained by the object region candidate integration S209 may be different from the region of the object intended by the user.
[0067] In Modification 3, the object region obtained as a result of the object region candidate integration S209 of Embodiment 1 is corrected based on the coordinates (instruction coordinates) acquired by the instruction reception unit and the values of the likelihood map.
[0068] A specific correction method will be described with reference to FIG. 10.
[0069] First, a method of acquiring and correcting only the likelihood regarding the object region will be described.
[0070] An object region likelihood representing the existence likelihood of the object region 1001 obtained by the object region candidate integration S209 is acquired. The position 1002 (C x , C y ) corresponding to the center in the image coordinate system of the object region 1001 and each point (M x , M y ) of the likelihood map are converted into points 1003 (I x , I y ) in the image coordinate system, and the Euclidean distance is calculated using a distance function similar to Equation 2. And (C x , C yThe object region likelihood is obtained using the values of the likelihood map corresponding to one or more points with a small Euclidean distance from [[ID=]]. The object region likelihood acquisition method may be, for example, the value of the likelihood map corresponding to the point with the closest Euclidean distance, or the average of the values of the likelihood maps corresponding to a plurality of points with a close Euclidean distance. When the object region likelihood thus obtained is below a certain value, it is considered to be an object region where the probability of the object's existence is low. Therefore, there may be a possibility that an object region 1001 having a center position 1002 different from the position indicated by the instruction coordinates 1004 is estimated. By moving the center position 1002 of the object region in the direction of the instruction coordinates 1004, there is a high possibility that it can be directly corrected to the region intended by the user. As a method of moving the object region in the direction of the instruction coordinates, a method of replacing the center position 1002 of the object region with the instruction coordinates 1004 can be considered. The object region is also shifted in accordance with the center position. In addition, a method of determining the amount of movement according to the object region likelihood L о is also conceivable. As an example, the maximum value L output in the likelihood map max and the object region likelihood L о , using the vector D1007 from the center position 1002 of the object region to the specified coordinates, each component V of the movement vector V1008 from the center position of the object region x、 V y is obtained by Equations 11 and 12. By applying the vector V1008 to the center position 1002 of the object region, the position of the object region can be corrected.
[0071] [Number]
[0072] Also, a method of correcting the object region according to the values of the likelihood map near the instruction coordinates can be considered. First, the object likelihood regarding the vicinity of the instruction coordinates (instruction coordinate object likelihood) is obtained (S902). The instruction coordinates 1004 (S x , S y ) and each point 1003 of the likelihood map converted into the image coordinate system (I x , I y) Calculate the Euclidean distance and obtain the likelihood of the object at the indicated coordinates from one or more likelihood map points (I x , I y ) that are close to the indicated coordinates 1004. The acquisition of the likelihood of the object at the indicated coordinates may be, for example, the value of the likelihood map corresponding to the closest point, or the average of the likelihood map values corresponding to the N closest points to the indicated coordinates. When the likelihood of the object at the indicated coordinates obtained in this way is high, it can be said that the probability of the presence of an object near the specified coordinates is estimated to be high. Similar to the correction based on the likelihood of the object region for the center position 1002 of the object region, by moving in the direction of the indicated coordinates, it is possible to correct to an object region close to the position intended by the user.
[0073] In addition, a correction method using both the likelihood of the object region and the likelihood of the object at the indicated coordinates can also be considered. According to the flowcharts of FIGS. 9 and 10, the flow of the correction method using both the likelihood of the object region and the likelihood of the object at the indicated coordinates will be described.
[0074] First, compare the likelihood of the object region and the likelihood of the object at the indicated coordinates (S903). When the likelihood of the object at the indicated coordinates is higher than the likelihood of the object region, perform a correction process based on the position 1002 of the object region obtained by object region candidate integration S209 and the indicated coordinates 1004 (S904). On the other hand, when the likelihood of the object at the indicated coordinates is less than or equal to the likelihood of the object region, output the object region obtained by object region integration as it is (S905).
[0075] The method of correction will be described with the example of FIG. 10. When the object region is a rectangle 1001, the width of the rectangle 、 The height of the rectangle is used as the size of the object, and the estimated center 1002 (C x , C y ) of the rectangle is used as the indicated coordinates 1004 (S x , S y) and the new object region 1005 thus replaced is output as the corrected object region (S905). Additionally, there is also a method of moving the object region in the direction where it is more likely that an object exists, according to the object region likelihood of the object region 1001 and the value of the indicated coordinate object likelihood at the indicated coordinates 1004. Let the vector from the indicated coordinates 1004 to the center position 1002 of the object region be D1007, and let the indicated coordinate object likelihood be L s , and let the object region likelihood be L о . Using Equation 14 and Equation 15, each component V x、 V y of the vector V for moving the object region according to each likelihood is obtained. By setting the position obtained by applying the vector V thus obtained to the center position 1002 of the object region as the center position of the new object region, the object region 1001 can also be corrected in the direction with a higher likelihood.
[0076]
Equation
[0077] By using the object region correction unit described in Modification 3, it is possible to correct to an object region closer to the object intended by the user than the object region output by the object region integration S209 described in Embodiment 1.
[0078] When realizing the likelihood map estimation unit using limited computing resources, the accuracy of the likelihood map and the object region candidates is limited, and there is a possibility that the position of the object cannot be captured as in the object region 1001 in FIG. 10. Even when the indicated coordinates 1004 accurately indicate the position of the object, the output object region may be estimated as an object region not intended by the user. In such a case, in Modification 3, it is possible to correct to the object region 1005 that more accurately captures the subject 1006 intended by the user using the indicated coordinates 1004 and the likelihood map.
[0079] <Modification 4> In the image processing apparatus of the above-described embodiment, the likelihood of an object included in the captured image was obtained using one likelihood map acquisition unit. Therefore, the accuracy of the likelihood map acquisition unit directly affects the accuracy of the object region obtained by the object region candidate integration unit. Further, in order to improve the accuracy of the object region, a second likelihood map acquisition unit is introduced in this embodiment.
[0080] The configuration of the image processing apparatus of this embodiment will be described with reference to FIG. 11.
[0081] The image processing apparatus 1100 includes the same configuration as that of the first embodiment, and includes a second likelihood map acquisition unit 1101 and an object region correction unit 1102. The configuration until the object region candidate integration unit 106 obtains one object region based on the captured image and the instruction coordinates obtained by the image acquisition unit 101 and the instruction reception unit 104 is the same as that of the first embodiment. The second likelihood map acquisition unit 1101 receives the captured image acquired by the image acquisition unit 101 and outputs a second likelihood map. The object region correction unit 1102 receives the object region, the instruction coordinates, the likelihood map, and the second likelihood map, and outputs one object region correction result by the result output unit 120.
[0082] Next, the specific processing flow will be described with reference to the flowchart of FIG. 12.
[0083] First, the processing from the start of shooting S200 to the acquisition of the likelihood map S202 is the same as that of the first embodiment. Next, a second likelihood map is acquired in the second likelihood map acquisition S1201. The second likelihood map acquisition S1201 outputs a likelihood map different from the likelihood map of the first embodiment. For example, it may be a likelihood map acquisition unit using a color histogram or edge density, or a likelihood map acquisition unit using a multi-layer neural network learned by a learning method different from that of the first embodiment. Following the second likelihood map acquisition S1201, the flow from the object region candidate acquisition S203 to the object region candidate integration S209 is the same as that of the first embodiment. Correction is performed on the one object region integrated by the object region candidate integration S209 based on the instruction coordinates, the likelihood map obtained in the likelihood map acquisition S202, and the second likelihood map obtained by the second likelihood map acquisition unit.
[0084] The correction method is, for example, first to obtain a vector V1007 for correcting the object region by performing the same processing as the object correction unit described in Modification 3 using the likelihood map obtained in the instruction coordinate and likelihood map acquisition S202. Next, the distance between the point obtained by converting each point of the second likelihood map into the image coordinate system and the point obtained by converting the instruction coordinate into the image coordinate system is obtained, and the second likelihood map value corresponding to the second likelihood map coordinate closest to that point is set as the second instruction coordinate object likelihood Ls2. Also, the second object region likelihood Lo2 is obtained using the second likelihood map value corresponding to one or more points whose distance from the position (C x ,C y ) corresponding to the center of the object region and the point obtained by converting each point of the second likelihood map into the image coordinate system is close. The second object region likelihood acquisition method may be the second likelihood map value closest to (C x ,C y ), or may be the average of the second likelihood map values corresponding to a plurality of points whose distance to (C x ,C y ) is close.
[0085] Using the second object region likelihood Lo 2、 thus obtained, the second specified coordinate object likelihood Ls 2、 and the vector D1007 from the center position of the object region to the instruction coordinate, each component of the second vector W for correcting the position of the object region is obtained by the following formulas 16 and 17.
[0086]
Equation
[0087] A method of correction can be considered by applying the average vector of the vector V in Modification 3 and the second vector W thus obtained to the center position of the object region output by the object candidate integration unit.
[0088] One object region corrected by the object region correction unit of the fifth embodiment is output by the result output S210.
[0089] In addition to the likelihood map acquisition unit in the above-described embodiment, by using a second likelihood map acquisition unit, it is possible to output a more plausible object region as the output object region.
[0090] By the second likelihood map acquisition unit acquiring a likelihood map from a color histogram or edge density, the validity of the object region output in the above-described embodiment can be determined by a method different from that of the likelihood map acquisition unit. Additionally, by making the second likelihood map acquisition unit a multi-layer neural network trained with learning data different from that of the likelihood map acquisition unit in Embodiment 1, the accuracy of the object region can also be supplemented. For example, a method of training the second likelihood map acquisition unit to react to a specific object that is easily misdetected and correcting it by the object region correction unit based on that information is also conceivable. Thereby, it becomes possible to perform correction while comprehensively considering the reliability of the likelihood map referred to when performing object region correction.
[0091] The present invention is also realized by executing the following processing. That is, software (program) that realizes the functions of the above-described embodiment is supplied to a system or device via a network or various storage media, and a computer (or CPU, MPU, etc.) of the system or device reads and executes the program.
Explanation of Signs
[0092] 100 Image processing apparatus 101 Image acquisition unit 102 Likelihood map acquisition unit 103 Estimation unit 104 Instruction reception unit 105 Object region candidate selection unit 106 Object region candidate integration unit 110 Imaging device 120 Result output unit
Claims
1. image acquisition means for acquiring a captured image instruction reception means for receiving an instruction for the captured image acquired by the image acquisition means likelihood map acquisition means for acquiring a likelihood map representing the presence likelihood of an object at each position of the captured image region acquisition means for acquiring a region tensor indicating the position of the object center and the object size based on each position of the captured image determination means for determining an object region corresponding to the instruction using one or more object region candidates obtained using the region tensor corresponding to coordinates located within a predetermined range from the position of the instruction received by the instruction reception means and the likelihood of the likelihood map An image processing apparatus, characterized by comprising the above
2. The determination means is characterized in that it sequentially selects a predetermined number of candidates from the object regions located closest to the position of the instruction acquired by the instruction reception means The image processing apparatus according to claim 1
3. The determination means is characterized in that it does not select an object region estimated to have a low presence probability of an object as a candidate The image processing apparatus according to claim 1 or 2
4. The determination means is characterized in that it integrates object region candidates using a weighted average of the values of the likelihood map The image processing apparatus according to any one of claims 1 to 3
5. The determination means is characterized in that it integrates object region candidates based on the distance between the position of the instruction acquired by the instruction reception means and the object region candidates The image processing apparatus according to any one of claims 1 to 4
6. The image processing apparatus further comprises correction means for correcting the object region based on the position of the instruction acquired by the instruction reception means The image processing apparatus according to any one of claims 1 to 4
7. The determination means is characterized in that it selects object region candidates corresponding to one or more coordinates located concentrically from the position of the instruction acquired by the instruction reception means. The image processing apparatus according to any one of claims 1 to 6
8. The likelihood map acquisition means is characterized in that it acquires a likelihood map using a neural network having a plurality of layers The image processing apparatus according to any one of claims 1 to 7
9. image acquisition means for acquiring a captured image instruction reception means for receiving an instruction for the captured image acquired by the image acquisition means likelihood map acquisition means for acquiring a likelihood map representing the presence likelihood of an object at each position of the captured image Region acquisition means for acquiring, at each position of the captured image, a region tensor indicating the position of the object center and the object size based on each position; An image processing system comprising: determination means for determining an object region corresponding to the instruction using one or more object region candidates obtained using the region tensor corresponding to coordinates located within a predetermined range from the position of the instruction received by the instruction reception means and the likelihood of the likelihood map.
10. An image acquisition step of acquiring a captured image; An instruction reception step of receiving an instruction for the captured image acquired in the image acquisition step; A likelihood map acquisition step of acquiring a likelihood map representing the presence likelihood of an object at each position of the captured image; A region acquisition step of acquiring, at each position of the captured image, a region tensor indicating the position of the object center and the object size based on each position; An image processing method comprising: a determination step of determining an object region corresponding to the instruction using one or more object region candidates obtained using the region tensor corresponding to coordinates located within a predetermined range from the position of the instruction received in the instruction reception step and the likelihood of the likelihood map.
11. A computer, Image acquisition means for acquiring a captured image; Instruction reception means for receiving an instruction for the captured image acquired by the image acquisition means; Likelihood map acquisition means for acquiring a likelihood map representing the presence likelihood of an object at each position of the captured image; Region acquisition means for acquiring, at each position of the captured image, a region tensor indicating the position of the object center and the object size based on each position; A program for causing the computer to function as determination means for determining an object region corresponding to the instruction using one or more object region candidates obtained using the region tensor corresponding to coordinates located within a predetermined range from the position of the instruction received by the instruction reception means and the likelihood map.
Citation Information
Patent Citations
Image processing apparatus, and image processing method
JP2019032773A
Information processing method and information processing system
JP2020042765A
Image processing device
JP2020173678A