Object position estimation device

The object position estimation device simplifies the estimation of subject position in images by converting relative to absolute distances using a conversion coefficient, reducing computational load and enhancing processing efficiency.

JP7839698B2Active Publication Date: 2026-04-02NTT DOCOMO INC +1
View PDF 4 Cites 0 Cited by

Patent Information

Authority / Receiving Office
JP · JP
Patent Type
Patents
Current Assignee / Owner
Filing Date
2022-06-24
Publication Date
2026-04-02

AI Technical Summary

Technical Problem

Existing methods for estimating the position of a subject in an image captured by a monocular camera require significant computational resources due to the application of multiple blur correction filters, leading to a high processing load.

Method used

An object position estimation device that utilizes an image acquisition unit, absolute and relative distance acquisition units, an object region estimation unit, a coefficient calculation unit, and a distance calculation unit to convert relative distance information into absolute distance information using a conversion coefficient, thereby simplifying the estimation process.

Benefits of technology

Enables accurate estimation of the position and orientation of a subject in an image with reduced computational complexity by leveraging pre-acquired absolute distance information and machine learning models for relative distance inference.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure 0007839698000001
    Figure 0007839698000001
  • Figure 0007839698000002
    Figure 0007839698000002
  • Figure 0007839698000003
    Figure 0007839698000003
Patent Text Reader

Abstract

To estimate the position of a subject included in a photographed image by simpler processing.SOLUTION: An object position estimation device 10 comprises: an absolute distance acquisition section 12 that acquires absolute distance information on each of a plurality of pixels included in a set region, which is at least a part of a region of a first image; a relative distance acquisition section 13 that acquires relative distance information on each pixel included in a second image where a person is present in a photographic range; an object region estimation section 14 that estimates an object region where the person is present in the second image; a coefficient calculation section 15 that calculates a conversion coefficient for converting the relative distance information into the absolute distance information on the basis of a correspondence between the absolute distance information and the relative distance information for each pixel in an invariable region; and a distance calculation section 16 that calculates an absolute distance from an imaging section 20 to the person by converting the relative distance information on the pixel included in the object region of the second image by the conversion coefficient.SELECTED DRAWING: Figure 1
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] One aspect of the present invention relates to an object position estimation device.

Background Art

[0002] Patent Document 1 discloses a method of estimating the distance to a subject shown in a target image by applying a plurality of blur correction filters to the target image captured by a monocular camera and obtaining a blur correction filter that, when added to the target image, results in a higher correlation with a reference image.

Prior Art Documents

Patent Documents

[0003]

Patent Document 1

Summary of the Invention

Problems to be Solved by the Invention

[0004] According to the method described in Patent Document 1 above, the distance to a subject (i.e., the position of the subject) can be estimated based on an image captured by a monocular camera. However, since it is necessary to apply a plurality of blur correction filters to the image to obtain the blur correction filter with the highest correlation, the processing load (computation amount) is relatively large.

[0005] Therefore, an object of one aspect of the present invention is to provide an object position estimation device that can estimate the position of a subject included in a captured image by simpler processing.

Means for Solving the Problems

[0006] An object position estimation device according to one aspect of the present invention comprises: an image acquisition unit that acquires an image of a predetermined shooting range captured by an imaging unit; an absolute distance acquisition unit that acquires absolute distance information indicating the absolute distance from the imaging unit for each of a plurality of pixels included in a set region, which is at least a part of the first image acquired by the image acquisition unit; a relative distance acquisition unit that acquires relative distance information indicating the relative distance between pixels for each pixel included in a second image in which a specific object exists within the shooting range acquired by the image acquisition unit; an object region estimation unit that estimates the object region in which the specific object exists in the second image; a coefficient calculation unit that calculates a conversion coefficient for converting relative distance information into absolute distance information based on the correspondence between the absolute distance information and relative distance information of each pixel included in an invariant region within the set region of the second image that has not changed from the first image; and a distance calculation unit that calculates the absolute distance from the imaging unit to the specific object by converting the relative distance information of pixels included in the object region of the second image using the conversion coefficient.

[0007] In an object position estimation device according to one aspect of the present invention, a conversion coefficient is calculated for converting the relative distance information of the second image into absolute distance information, based on absolute distance information acquired in advance based on the first image and relative distance information acquired for each pixel of the second image in which the specific object is captured. By using this conversion coefficient to convert the relative distance information of pixels included in the region of interest in the second image (i.e., the object region in which the specific object is captured), the absolute distance from the imaging unit to the specific object (i.e., the absolute position of the specific object in real space) can be calculated (estimated). Therefore, according to the above object position estimation device, the position of a subject (specific object) included in the captured image (second image) can be estimated with simpler processing. [Effects of the Invention]

[0008] According to one aspect of the present invention, it is possible to provide an object position estimation device that can estimate the position of a subject included in an captured image through simpler processing. [Brief explanation of the drawing]

[0009] [Figure 1] This figure shows an example of the functional configuration of an object position estimation device according to one embodiment. [Figure 2] This is a diagram showing an example of the first image. [Figure 3] This is a diagram showing an example of the second image. [Figure 4] This figure shows an example of a depth image that visually represents relative distance information. [Figure 5] This figure shows an example of how to calculate the conversion coefficient. [Figure 6] This figure shows an example of the correction process in the correction unit. [Figure 7] This is a flowchart illustrating an example of how an object position estimation device works. [Figure 8] This figure shows an example of the hardware configuration of an object position estimation device. [Modes for carrying out the invention]

[0010] Hereinafter, an embodiment of the present invention will be described in detail with reference to the attached drawings. In the description of the drawings, the same or equivalent elements will be denoted by the same reference numerals, and redundant descriptions will be omitted.

[0011] Figure 1 shows an example of the functional configuration of an object position estimation device 10 according to one embodiment. The object position estimation device 10 is a device that detects a predetermined specific object from images (video) sequentially acquired by an imaging unit 20 which is installed (fixed) in a predetermined location (e.g., a monitoring area) and has a fixed shooting direction and field of view (shooting range), and estimates the absolute position of the specific object. The imaging unit 20 is, for example, a monocular camera.

[0012] A specific object is a predetermined type of object that may appear as the subject of an image captured by the imaging unit 20. For example, a specific object is an object having a plurality of predetermined feature points. In this embodiment, the specific object is a person. However, the specific object may be a living organism other than a person, such as a dog or a cat, or a non-living thing like a robot.

[0013] Feature points are distinctive parts that constitute a portion of a specific object. In this embodiment, feature points are parts detected in human skeletal estimation, such as joints like the shoulders, elbows, wrists, and ankles, as well as ears, noses, and eyes.

[0014] The absolute position is positional information linked to the real space in which the imaging unit 20 (camera) exists. For example, the absolute position can be determined by a three-dimensional coordinate system with the position of the imaging unit 20 as the reference point (origin). In this embodiment, the object position estimation device 10 estimates the absolute position of a person in an image by estimating the absolute distance from the imaging unit 20 to each of the multiple feature points of the person captured in the image by the imaging unit 20. That is, the direction in which the person is located relative to the imaging unit 20 is determined from the position of the person captured in the image, and the absolute distance from the imaging unit 20 to the person is estimated, thereby estimating the position coordinates (absolute position) of the person in real space. Furthermore, in this embodiment, the posture of the person (i.e., the posture determined based on the position of each feature point) is also estimated by estimating the absolute position of each of the multiple feature points.

[0015] As shown in Figure 1, the object position estimation device 10 includes an image acquisition unit 11, an absolute distance acquisition unit 12, a relative distance acquisition unit 13, an object region estimation unit 14, a coefficient calculation unit 15, a distance calculation unit 16, and a correction unit 17.

[0016] The image acquisition unit 11 acquires images of a predetermined shooting range captured by the imaging unit 20. The imaging unit 20 is, for example, a monocular camera installed (fixed) in a predetermined location (e.g., a monitoring area). The image acquisition unit 11 sequentially acquires images (e.g., each frame that makes up a video) at predetermined intervals. The imaging unit 20 may be provided within the object position estimation device 10 itself, or it may be an external device configured to communicate data with the object position estimation device 10.

[0017] The absolute distance acquisition unit 12 acquires absolute distance information indicating the absolute distance (depth) from the imaging unit 20 for each of a plurality of pixels included in a setting region which is at least a partial region of the first image acquired by the image acquisition unit 11. FIG. 2 is a diagram showing an example of the first image IM1. In the present embodiment, as an example, the imaging unit 20 is fixed at a certain location in the room so as to photograph the scenery in the room. Note that the first image IM1 may be an image including a person as a subject, but from the viewpoint of more surely and sufficiently securing an invariant region (that is, the number of pixels available for calculating the conversion coefficient) described later, it is preferable that the first image IM1 includes a large amount of a background portion composed of fixed objects such as walls and ceilings that do not change over time. From such a viewpoint, in the present embodiment, the absolute distance acquisition unit 12 uses, as the first image IM1, an image captured in a state where no person exists in the shooting range.

[0018] The absolute distance acquisition unit 12 acquires the absolute distance information of a plurality of pixels included in the setting region by using an arbitrary known method. For example, the absolute distance acquisition unit 12 may acquire the absolute distance information of all pixels included in the first image IM1 by setting the entire region of the first image IM1 as the setting region by executing the following processing, for example.

[0019] First, before installing (fixing) the imaging unit 20, the absolute distance acquisition unit 12 acquires multi-viewpoint images obtained by photographing the above shooting range from a plurality of different directions with the imaging unit 20. The absolute distance acquisition unit 12 estimates the camera parameters (external parameters and internal parameters) of the imaging unit 20, which is a monocular camera, by applying SfM (Shape from Motion) to the multi-viewpoint images. The external parameters are, for example, information indicating the three-dimensional position and orientation of the imaging unit 20. The internal parameters are, for example, information indicating the lens distortion of the imaging unit 20.

[0020] Next, the absolute distance acquisition unit 12 generates dense 3D point cloud data of the target space (the space included in the shooting range) by applying MVS (Multi-View Stereo) to the multi-view images and camera parameters. Each point in the 3D point cloud data is associated with an absolute position (coordinate information) in real space. Subsequently, the absolute distance acquisition unit 12 projects the 3D point cloud data onto the first image IM1 using the camera parameters. By projecting the 3D point cloud data onto the first image IM1 in this way and associating each point in the 3D point cloud data with each pixel, absolute distance information indicating the absolute distance to each pixel in the entire area of ​​the first image IM1 can be obtained.

[0021] The method for acquiring absolute distance information is not limited to the method described above. In this embodiment, the entire area of ​​the first image IM1 is used as the setting area, but the setting area may be a part of the first image IM1. For example, the absolute distance acquisition unit 12 may measure the distance from the imaging unit 20 to a fixed object included in the shooting range (for example, a wall, furniture, or a sign prepared as a landmark). The absolute distance acquisition unit 12 may then set the area including the fixed object as the setting area and acquire absolute distance information for each pixel within the setting area by associating the measured distance with the pixels included in the setting area (i.e., pixels in the first image IM1 that show the fixed object).

[0022] The relative distance acquisition unit 13 acquires relative distance information indicating the relative distance between pixels for each pixel included in the second image, which is captured by the image acquisition unit 11 and shows a person in the shooting range. Figure 3 shows an example of the second image IM2. In the example in Figure 3, two people H1 and H2 are included in the shooting range.

[0023] Relative distance information is a numerical value that indicates the distance from the imaging unit 20 to the pixels included in the second image IM2 (i.e., the objects captured in the pixels). However, the numerical value of the relative distance information does not necessarily coincide with the absolute distance (i.e., the actual distance from the imaging unit 20 to the objects) mentioned above. For example, consider a case where the second image IM2 includes objects a, which is "1m" away from the imaging unit 20, and objects b, which is "2m" away from the imaging unit 20. In this case, the relative distance information (numerical value) for the pixel corresponding to object a may be expressed as "30", and the relative distance information for the pixel corresponding to object b may be expressed as "60". Thus, while the relative distance information allows us to understand the relative distance between pixels (in the above example, that object b is twice as far from the imaging unit 20 as object a), it does not allow us to understand the actual distance (absolute distance) from the imaging unit 20 to each object a and b.

[0024] The relative distance acquisition unit 13 acquires relative distance information for each pixel, for example, as follows. First, a known machine learning model for depth inference is prepared. Such a machine learning model (trained model) can be obtained, for example, by performing machine learning such as deep learning using training data consisting of training images and depth images (ground truth labels) corresponding to those images. An example of such a machine learning model is a model using ViT (Vision Transformer). The relative distance acquisition unit 13 inputs the second image IM2 to such a machine learning model and acquires the output result (inference result) of the machine learning model as relative distance information for each pixel of the second image IM2.

[0025] Figure 4 shows an example of a depth image IM3 that visually represents relative distance information. The depth image IM3 is obtained by applying a color and intensity corresponding to the relative distance information (numerical value) acquired by the relative distance acquisition unit 13 as described above to each pixel of the second image IM2. In the example of the depth image IM3 shown in Figure 4, the areas closer to the imaging unit 20 are made darker.

[0026] The object region estimation unit 14 estimates the object region in the second image IM2 in which a person exists. In this embodiment, the object region estimation unit 14 further estimates the positions of multiple feature points of the person included in the second image IM2. In the example in Figure 3, the object regions A1 and A2 of the two people H1 and H2 included in the second image IM2, estimated by the object region estimation unit 14, are displayed on the second image IM2. In addition, the positions of multiple feature points F1 and F2 of people H1 and H2, estimated by the object region estimation unit 14, are also displayed on the second image IM2. In this embodiment, line segments connecting adjacent feature points F1 and F2 are also displayed together with the feature points F1 and F2, so that the estimated skeletons of people H1 and H2 are displayed on the second image IM2. Thus, the object region estimation unit 14 may superimpose the estimated object region, feature points, line segments, etc., on the second image IM2.

[0027] The object region estimation unit 14 estimates the object region and the position of feature points, for example, as follows. First, a known machine learning model for skeletal position estimation (estimation of the object region and skeleton (multiple feature points)) is prepared. Such a machine learning model (trained model) can be obtained, for example, by performing machine learning such as deep learning using training data consisting of training images (images containing people) and corresponding skeletal positions (positions of the object region and feature points) (ground truth labels). Examples of such machine learning models include models using MMPose, OpenPose, etc. The object region estimation unit 14 inputs the second image IM2 into such a machine learning model and obtains the output result of the machine learning model as the estimated results of object regions A1, A2 and multiple feature points F1, F2 contained in the second image IM2. The object region estimation unit 14 may also estimate object regions A1, A2 using other known methods (for example, background subtraction method).

[0028] The coefficient calculation unit 15 calculates a conversion coefficient for converting relative distance information to absolute distance information based on the correspondence between absolute distance information and relative distance information of each pixel included in the invariant region of the setting area of ​​the second image IM2 (in this embodiment, the entire area of ​​the second image IM2) that has not changed from the first image IM1. In this embodiment, if the first image IM1 is an image that does not include a person as a subject (i.e., an image of only the background), the region of the setting area of ​​the second image IM2 that does not overlap with the object regions A1 and A2 (i.e., the region in which the same background is shown in both the first image IM1 and the second image IM2) can be used as the invariant region. Such an invariant region can be extracted (estimated), for example, by using the background subtraction method described above.

[0029] Figure 5 shows an example of the calculation of conversion coefficients. In the graph in Figure 5, the horizontal axis represents relative distance information (arbitrary unit au), and the vertical axis represents absolute distance information (here, as an example, the unit is m). The coefficient calculation unit 15 calculates the conversion coefficients as follows, for example. That is, as shown in Figure 5, the coefficient calculation unit 15 plots points P on the graph for each pixel included in the invariant region, associating the relative distance information (i.e., the value obtained by the relative distance acquisition unit 13) and the absolute distance information (i.e., the value obtained by the absolute distance acquisition unit 12) corresponding to the same pixel. Then, the coefficient calculation unit 15 finds the regression line R of the multiple points P. This regression line R (regression equation) is expressed by the following equation (1), where X is the explanatory variable (variable corresponding to relative distance information), Y is the dependent variable (variable corresponding to absolute distance information), A is the slope, and B is the intercept. In this case, the slope A and intercept B are obtained as conversion coefficients (parameters) that convert relative distance information to absolute distance information. In other words, by substituting the relative distance information x1 of any pixel into X in equation 1 above, it becomes possible to obtain the absolute distance information y1 (= A × x1 + B) of that pixel. Formula 1: Y=A×X+B

[0030] Furthermore, the coefficient calculation unit 15 does not necessarily need to use information from all pixels included in the invariant region when calculating the above-mentioned conversion coefficients. For example, the coefficient calculation unit 15 may calculate the regression line R (i.e., the conversion coefficients, namely the slope A and intercept B) based only on data (points P) corresponding to a portion of pixels randomly selected from a plurality of pixels included in the invariant region.

[0031] The distance calculation unit 16 calculates the absolute distance from the imaging unit 20 to people H1 and H2 by converting the relative distance information of pixels included in object regions A1 and A2 of the second image IM2 using conversion coefficients. In this embodiment, the distance calculation unit 16 calculates the absolute distance from the imaging unit 20 to each of the multiple feature points F1 and F2 by converting the relative distance information of pixels corresponding to each of the multiple feature points F1 and F2 in the object regions A1 and A2 using conversion coefficients. By calculating the absolute distance of the multiple feature points F1 that constitute the skeleton of person H1 in this way, it becomes possible to understand where each feature point F1 is located in real space (3D space), and thus it becomes possible to understand the position and orientation of person H1 in real space. The same applies to person H2.

[0032] The correction unit 17 corrects the absolute distance between feature points calculated by the distance calculation unit 16 based on predetermined constraints regarding the positional relationship between multiple feature points.

[0033] Figure 6 is a schematic diagram illustrating an example of the correction process of the correction unit 17. Figure 6(A) shows the absolute positions of feature points Fa, Fb, and Fc, which are identified based on the absolute distance between the three feature points Fa, Fb, and Fc calculated by the distance calculation unit 16. Here, feature points Fa, Fb, and Fc are adjacent to each other. Line segment Bab is the line segment connecting feature point Fa and feature point Fb, line segment Bac is the line segment connecting feature point Fa and feature point Fc, and line segment Bbc is the line segment connecting feature point Fb and feature point Fc.

[0034] Regarding the positional relationship between feature points Fa, Fb, and Fc, it is assumed that the appropriate ratio of line segments Bab, Bac, and Bbc is defined as a constraint condition, for example, from the viewpoint of the appropriate dimensional ratio of the human skeleton. Furthermore, in light of this constraint condition, the absolute positions of each feature point Fa, Fb, and Fc shown in Figure 6(A) are assumed to be inappropriate. Specifically, the ratio of the lengths of line segments Bab and Bac to the length of line segment Bbc is assumed to be greater than the appropriate ratio. In this case, as shown in Figure 6(B), the correction unit 17 corrects the absolute distance of feature point Fa calculated by the distance calculation unit 16 to be shorter so as to satisfy the above constraint condition.

[0035] Next, an example of the operation of the object position estimation device 10 will be described with reference to Figure 7.

[0036] In step S1, the image acquisition unit 11 acquires the first image IM1 captured by the imaging unit 20. In this embodiment, as an example, the image acquisition unit 11 acquires the first image IM1 as an image captured when there are no people in the shooting range.

[0037] In step S2, the absolute distance acquisition unit 12 acquires absolute distance information indicating the absolute distance (depth) from the imaging unit 20 for each of the multiple pixels included in the setting region, which is at least a part of the first image IM1. As described above, in this embodiment, as an example, the setting region is the entire region of the first image IM1.

[0038] In step S3, the image acquisition unit 11 acquires the second image IM2 captured by the imaging unit 20 while a person is present in the shooting range.

[0039] In step S4, the relative distance acquisition unit 13 acquires relative distance information for each pixel included in the second image IM2. For example, the relative distance acquisition unit 13 acquires a depth image IM3 (see Figure 4) corresponding to the second image IM2.

[0040] In step S5, the object region estimation unit 14 estimates the positions of object regions A1 and A2 in which people H1 and H2 exist, and the positions of multiple feature points F1 and F2 of people H1 and H2 in the second image IM2.

[0041] In step S6, the coefficient calculation unit 15 calculates conversion coefficients (in this embodiment, the slope A and intercept B of the regression line R in Figure 5) based on the correspondence between the absolute distance information and the relative distance information of each pixel included in the invariant region of the setting area of ​​the second image IM2 (in this embodiment, the entire area of ​​the second image IM2) that has not changed from the first image IM1.

[0042] In step S7, the distance calculation unit 16 calculates the absolute distance from the imaging unit 20 to people H1 and H2 (multiple feature points F1 and F2) by converting the relative distance information of the pixels corresponding to each of the multiple feature points F1 and F2 in the object regions A1 and A2 using a conversion coefficient.

[0043] In step S8, the correction unit 17 corrects the absolute distance of the feature points calculated by the distance calculation unit 16 based on predetermined constraints regarding the positional relationship between the multiple feature points (see Figure 6).

[0044] Of the processes described above, steps S1 and S2 only need to be executed once. On the other hand, steps S3 to S8 are executed for each image (frame) acquired by the imaging unit 20. This is because the conversion coefficient calculated in step S6 depends on the inference result of the relative distance information (depth image IM3) acquired in step S4. In other words, if the second image IM2 changes and the above inference result changes, the conversion coefficient also changes accordingly. Furthermore, the correction process in step S8 may be omitted if it is not necessary to correct the absolute distance of each feature point calculated in step S7.

[0045] In the object position estimation device 10 described above, a conversion coefficient (in this embodiment, the slope A and intercept B in Equation 1 above) is calculated for converting the relative distance information of the second image IM2 into absolute distance information, based on absolute distance information acquired in advance based on the first image IM1 and relative distance information acquired for each pixel of the second image IM2 in which a specific object (in this embodiment, "person") is captured. By using this conversion coefficient to convert the relative distance information of pixels included in the region of interest in the second image IM2 (i.e., the object region A1, A2 in which the specific object is captured), the absolute distance from the imaging unit 20 to the specific object (person H1, H2) (i.e., the absolute position of the specific object in real space) can be calculated (estimated). Therefore, the object position estimation device 10 allows the position of a subject (specific object) included in the captured image (second image IM2) to be estimated with simpler processing.

[0046] In this embodiment, the specific object is an object having a plurality of predetermined feature points F1 and F2 (in this embodiment, people H1 and H2), and the object region estimation unit 14 estimates the positions of the plurality of feature points F1 and F2 of the specific object included in the second image IM2. The distance calculation unit 16 calculates the absolute distance from the imaging unit 20 to each of the plurality of feature points F1 and F2 by converting the relative distance information of the pixels corresponding to each of the plurality of feature points F1 and F2 included in the second image IM2 using a conversion coefficient. With the above configuration, since the absolute position of each of the plurality of feature points F1 and F2 constituting the specific object is estimated individually, not only the position but also the orientation of the specific object can be estimated.

[0047] In this embodiment, the correction unit 17 corrects the absolute distance of the feature points calculated by the distance calculation unit 16 based on predetermined constraints regarding the positional relationship between multiple feature points. With the above configuration, as shown in the example in Figure 6, the absolute distance (absolute position) of the feature points Fa, Fb, and Fc can be estimated with greater accuracy by correcting based on constraints regarding the positional relationship between the feature points Fa, Fb, and Fc (in the example in Figure 6, the correction shortens the absolute distance of feature point Fa).

[0048] Furthermore, the block diagrams used in the description of the above embodiments show functional units. These functional blocks (components) are realized by any combination of at least one of hardware and software. Moreover, the method of realizing each functional block is not particularly limited. That is, each functional block may be realized using one device that is physically or logically coupled, or it may be realized using two or more physically or logically separated devices that are directly or indirectly connected (for example, using wired or wireless connections). A functional block may be realized by combining the above one device or the above multiple devices with software.

[0049] Functions include, but are not limited to, judgment, decision, judgment, calculation, calculation, processing, derivation, investigation, exploration, confirmation, reception, transmission, output, access, resolution, selection, selection, establishment, comparison, assumption, expectation, deem, broadcasting, notifying, communicating, forwarding, configuring, reconfiguring, allocating (mapping), and assigning.

[0050] For example, the object position estimation device 10 in one embodiment of the present disclosure may function as a computer that performs the object detection method of the present disclosure. Figure 8 is a diagram showing an example of the hardware configuration of the object position estimation device 10 according to one embodiment of the present disclosure. The object position estimation device 10 may be physically configured as a computer device including a processor 1001, memory 1002, storage 1003, communication device 1004, input device 1005, output device 1006, bus 1007, etc.

[0051] In the following explanation, the term "device" can be replaced with "circuit," "device," "unit," etc. The hardware configuration of the object position estimation device 10 may include one or more of the devices shown in Figure 8, or it may be configured without some of the devices.

[0052] Each function in the object position estimation device 10 is realized by loading predetermined software (programs) onto hardware such as the processor 1001 and memory 1002, which allows the processor 1001 to perform calculations, control communication by the communication device 1004, and control at least one of data reading and writing in the memory 1002 and storage 1003.

[0053] The processor 1001 controls the entire computer, for example, by running an operating system. The processor 1001 may consist of a central processing unit (CPU) that includes interfaces with peripheral devices, control units, arithmetic units, registers, and so on.

[0054] Furthermore, the processor 1001 reads programs (program code), software modules, data, etc., from at least one of the storage 1003 and the communication device 1004 into the memory 1002 and executes various processes accordingly. The program used is one that causes the computer to execute at least a part of the operations described in the above embodiment. For example, each functional part of the object position estimation device 10 (e.g., the coefficient calculation unit 15) may be stored in the memory 1002 and implemented by a control program that runs on the processor 1001, and other functional blocks may be implemented similarly. The above-described various processes have been explained as being executed by one processor 1001, but they may be executed simultaneously or sequentially by two or more processors 1001. The processor 1001 may be implemented by one or more chips. The program may be transmitted from the network via a telecommunications line.

[0055] Memory 1002 is a computer-readable recording medium and may consist of at least one of the following: ROM (Read Only Memory), EPROM (Erasable Programmable ROM), EEPROM (Electrically Erasable Programmable ROM), RAM (Random Access Memory), etc. Memory 1002 may also be called a register, cache, main memory, etc. Memory 1002 can store executable programs (program code), software modules, etc., for carrying out the object position estimation method according to one embodiment of the present disclosure.

[0056] Storage 1003 is a computer-readable recording medium and may consist of at least one of the following: an optical disc such as a CD-ROM (Compact Disc ROM), a hard disk drive, a flexible disk, a magneto-optical disk (e.g., a compact disc, a digital multipurpose disc, a Blu-ray® disc), a smart card, flash memory (e.g., a card, a stick, a key drive), a floppy® disk, a magnetic strip, etc. Storage 1003 may also be called an auxiliary storage device. The above-mentioned storage medium may be, for example, a database, server, or other suitable medium including at least one of memory 1002 and storage 1003.

[0057] The communication device 1004 is hardware (transceiver / receiver device) for communicating between computers via at least one of a wired network and a wireless network, and is also called a network device, network controller, network card, communication module, etc.

[0058] The input device 1005 is an input device that accepts input from an external source (e.g., a keyboard, mouse, microphone, switch, button, sensor, etc.). The output device 1006 is an output device that outputs to an external source (e.g., a display, speaker, LED lamp, etc.). The input device 1005 and the output device 1006 may be configured as an integrated unit (e.g., a touch panel).

[0059] Furthermore, each device, such as the processor 1001 and memory 1002, is connected by a bus 1007 for communicating information. The bus 1007 may be configured using a single bus, or different buses may be configured for each device.

[0060] Furthermore, the object position estimation device 10 may be configured to include hardware such as a microprocessor, a digital signal processor (DSP), an ASIC (Application Specific Integrated Circuit), a PLD (Programmable Logic Device), and an FPGA (Field Programmable Gate Array), and some or all of each functional block may be realized by such hardware. For example, the processor 1001 may be implemented using at least one of these hardware components.

[0061] Although this embodiment has been described in detail above, it will be clear to those skilled in the art that this embodiment is not limited to the embodiments described herein. This embodiment can be implemented as a modified and altered form without departing from the spirit and scope of the invention as defined by the claims. Therefore, the description herein is for illustrative purposes only and is not intended to be restrictive in any way to this embodiment.

[0062] The processing procedures, sequences, flowcharts, etc., of each aspect / embodiment described herein may be reordered, provided they are consistent with each other. For example, the methods described herein present various step elements in an exemplary order and are not limited to that specific order.

[0063] Input and output information may be stored in a specific location (e.g., memory) or managed using a management table. Input and output information may be overwritten, updated, or appended to. Output information may be deleted. Input information may be transmitted to other devices.

[0064] The determination may be made by a value represented by 1 bit (0 or 1), by a boolean value (true or false), or by a numerical comparison (for example, a comparison with a predetermined value).

[0065] Each aspect / embodiment described herein may be used individually, in combination, or switched between as needed during implementation. Furthermore, notification of specific information (e.g., notification that "X is") is not limited to explicit notification, but may also be implicit (e.g., by not providing such notification).

[0066] Software should be broadly interpreted to mean instructions, instruction sets, code, code segments, program code, programs, subprograms, software modules, applications, software applications, software packages, routines, subroutines, objects, executable files, execution threads, procedures, functions, and so on, whether they are called software, firmware, middleware, microcode, hardware description languages, or by any other name.

[0067] Furthermore, software, instructions, information, etc., may be transmitted and received via a transmission medium. For example, if software is transmitted from a website, server, or other remote source using at least one of wired technology (such as coaxial cable, fiber optic cable, twisted pair, or digital subscriber line (DSL)) and wireless technology (such as infrared or microwave), then at least one of these wired and wireless technologies is included in the definition of a transmission medium.

[0068] The information, signals, etc. described in this disclosure may be represented using any of the various different techniques. For example, the data, instructions, commands, information, signals, bits, symbols, chips, etc. that may be referred to throughout the above description may be represented by voltage, current, electromagnetic waves, magnetic fields or magnetic particles, optical fields or photons, or any combination thereof.

[0069] Furthermore, the information, parameters, etc., described in this disclosure may be expressed using absolute values, relative values ​​from a predetermined value, or corresponding other information.

[0070] The names used for the parameters described above are not restrictive in any way. Furthermore, the formulas and other expressions using these parameters may differ from those expressly disclosed in this disclosure. Since various information elements can be identified by any suitable name, the various names assigned to these various information elements are not restrictive in any way.

[0071] In this disclosure, the phrase "based on" does not mean "based solely on" unless otherwise specified. In other words, the phrase "based on" means both "based solely on" and "based at least on."

[0072] Any reference to elements using the designations “first,” “second,” etc., as used in this disclosure does not generally limit the quantity or order of those elements. These designations may be used in this disclosure as a convenient way to distinguish between two or more elements. Accordingly, references to the first and second elements do not imply that only two elements may be employed, or that the first element must precede the second element in any way.

[0073] Where the terms “include,” “including,” and variations thereof are used in this disclosure, these terms are intended to be inclusive, as is the term “comprising.” Furthermore, the term “or” as used in this disclosure is not intended to mean exclusive OR.

[0074] In this disclosure, if articles are added through translation, such as a, an, and the in English, this disclosure may include the fact that the noun following these articles is plural.

[0075] In this disclosure, the term "A and B are different" may mean "A and B are different from each other." The term may also mean "A and B are each different from C." Terms such as "separate" and "combine" may be interpreted similarly to "different." [Explanation of Symbols]

[0076] 10...Object position estimation device, 11...Image acquisition unit, 12...Absolute distance acquisition unit, 13...Relative distance acquisition unit, 14...Object region estimation unit, 15...Coefficient calculation unit, 16...Distance calculation unit, 17...Correction unit, 20...Imaging unit, A1, A2...Object region, F1, F2, Fa, Fb, Fc...Feature points, H1, H2...Person (specific object), IM1...First image, IM2...Second image, R...Regression line.

Claims

1. An image acquisition unit acquires an image of a predetermined shooting range captured by the imaging unit, An absolute distance acquisition unit acquires absolute distance information indicating the absolute distance from the imaging unit for each of a plurality of pixels included in a set region, which is at least a part of the first image acquired by the image acquisition unit. A relative distance acquisition unit acquires relative distance information indicating the relative distance between pixels for each pixel included in the second image obtained by the image acquisition unit, in which a specific object is present in the shooting range. An object region estimation unit that estimates the object region in which the specific object exists in the second image, A coefficient calculation unit calculates a conversion coefficient for converting relative distance information to absolute distance information based on the correspondence between the absolute distance information and the relative distance information of each pixel included in the invariant region of the setting region of the second image that has not changed from the first image, A distance calculation unit calculates the absolute distance from the imaging unit to the specific object by converting the relative distance information of pixels included in the object region of the second image using the conversion coefficient, An object position estimation device equipped with the following features.

2. The absolute distance acquisition unit uses an image captured when the specific object is not present in the shooting range as the first image. The object position estimation device according to claim 1.

3. The aforementioned specific object is an object having a predetermined number of characteristic points, The object region estimation unit further estimates the positions of the plurality of feature points of the specific object included in the second image, The distance calculation unit calculates the absolute distance from the imaging unit to each of the multiple feature points by converting the relative distance information of the pixels corresponding to each of the multiple feature points included in the second image using the conversion coefficient. The object position estimation device according to claim 1 or 2.

4. The system further includes a correction unit that corrects the absolute distance of the feature points calculated by the distance calculation unit based on predetermined constraints regarding the positional relationship between the plurality of feature points. The object position estimation device according to claim 3.

Citation Information

Patent Citations

  • Image processing device and imaging device

    JP2020024563A

  • Method and apparatus for acquiring joint position, and method and apparatus for acquiring motion

    JP2020042476A

  • Image processing method and device, electronic apparatus, storage medium and computer program

    JP2023027227A

  • Extracting depth information from video from a single camera

    US20130063556A1