Information processing device and information processing method
The system addresses low-resolution surveillance cameras by using a virtual space model and high-resolution images to accurately identify objects of interest, enhancing customer behavior analysis in commercial facilities.
Patent Information
- Application Number
- PCT/JP2024/010826
- Authority / Receiving Office
- WO · WO
- Patent Type
- Applications
- Current Assignee / Owner
- Filing Date
- 2024-03-19
- Publication Date
- 2025-09-25
AI Technical Summary
Existing surveillance cameras in commercial facilities have low resolution, making it difficult to estimate customer attention and identify objects of interest accurately, and updating 3D models of store interiors to reflect product changes is time-consuming.
A system that uses a first camera with low resolution to estimate spatial coordinates in a virtual space model, combined with a second camera with higher resolution to calculate transformation parameters, allowing precise identification of objects of interest using a virtual space model and high-resolution images.
Enables easy identification of objects at positions of interest in real space by leveraging high-resolution images and transformation parameters, improving accuracy and efficiency in tracking customer behavior.
Smart Images

Figure JP2024010826_25092025_PF_FP_ABST
Abstract
Description
Information processing device and information processing method
[0001] The present invention relates to an information processing device and an information processing method.
[0002] In commercial facilities, stores, etc., it is required to estimate the target of a customer's attention as customer behavior by determining the position where the customer touches a displayed commodity, the direction of the customer's gaze, etc., using images from a surveillance camera, etc. For example, Patent Document 1 listed below discloses a system that identifies a person's attention position based on image information.
[0003] Japanese Patent Application Laid-Open No. 2007-299207
[0004] Because the resolution of images captured by surveillance cameras permanently installed in stores and the like is often low, it has been difficult to estimate an object of interest from the images captured by the surveillance cameras. Furthermore, by preparing a 3D model that reflects the interior of the store and products, etc., based on images captured inside the store, it is possible to identify an object located at a position of interest identified from the camera image based on the 3D model. However, it has been difficult to prepare a 3D model that precisely reflects the products, etc., inside the store. Furthermore, it has been time-consuming and difficult to constantly update the 3D model to the latest version as products, etc., are replaced.
[0005] The present invention has been made in consideration of the above problems, and has as its object to easily identify an object that exists at a position of interest in space.
[0006] In order to solve the above problem, an information processing device according to one aspect of the present disclosure includes: a spatial coordinate acquisition unit that acquires focus position spatial coordinates, which are the three-dimensional coordinates of the focus position in virtual space, estimated by reflecting the focus position in real space represented in a first image, which is an image of real space captured by a first camera, in a virtual space model that represents a virtual space corresponding to the real space; a second image acquisition unit that acquires a second image, which is an image of real space captured by a second camera and includes the focus position; a parameter calculation unit that calculates conversion parameters in a predetermined conversion formula that expresses, using predetermined conversion parameters, the relationship between the three-dimensional coordinates of a specific point, which is a specific point in real space, and the two-dimensional coordinates of the specific point in the second image captured of real space; and an image coordinate acquisition unit that acquires focus position image coordinates, which are the two-dimensional coordinates of the focus position in the second image, by applying the focus position spatial coordinates to the conversion formula.
[0007] According to the above aspect, transformation parameters of a transformation formula expressing the relationship between the three-dimensional coordinates of a specific point in real space and the two-dimensional coordinates of the specific point in the second image are calculated, and information for identifying the focus position in the second image is obtained by applying the focus position spatial coordinates estimated by reflecting the focus position acquired based on the first image in the virtual space model to the transformation formula. Therefore, by using a camera capable of capturing a second image with a resolution high enough to identify an object represented in the image as the second camera, it is possible to identify the object represented at the focus position.
[0008] It becomes possible to easily identify an object that exists at a position of interest in real space that is displayed in an image.
[0009] 6( a) is a block diagram showing the functional configuration of an information processing system and an information processing device according to this embodiment. FIG. 6( b) is a sub-block diagram showing the functional configuration of a parameter calculation unit. FIG. 6( a) is a diagram showing an example of a real space and an example of a first image captured by a surveillance camera. FIG. 6( b) is a diagram showing an example of a virtual space model representing a virtual space corresponding to the real space. FIG. 6( a) is a diagram showing a state in which a second image is captured in the real space and an example of the captured second image. FIG. 6( a) is a diagram showing a first example of installation of a virtual camera in a virtual space. FIG. 6( b) is a diagram showing a second example of installation of a virtual camera in a virtual space. FIG. 6( b) is a diagram showing an example of a virtual space image, an example of a real space image (second image), and an example of feature point extraction and matching processing. FIG. 6( b) is a diagram showing an example of calculation processing of reference space coordinates, which are the coordinates of reference feature points in the virtual space. FIG. 6( a) is a diagram showing an example of reinstallation of a virtual camera. FIG. 6( b) is a diagram showing an example of how to superimpose an image capture area between a virtual space image captured by a specific virtual camera and a virtual space image captured by a reinstalled virtual camera. FIG. 6( b) is a flowchart showing the processing content of an information processing method in an information processing device. FIG. 6( a) is a diagram showing the configuration of an information processing program. FIG. 6( b) is a hardware block diagram of an information processing device.
[0010] An embodiment of an information processing system according to the present invention will be described with reference to the drawings. Whenever possible, the same components are designated by the same reference numerals, and redundant description will be omitted.
[0011] 1 is a diagram showing the functional configuration of an information processing system and an information processing device according to this embodiment. The information processing system 1 of this embodiment is a system that obtains information that identifies a position of interest in an image in order to identify an object that exists at a position of interest in real space represented in the image, and is configured, for example, by an information processing device 10.
[0012] 1, the information processing device 10 functionally comprises a first image acquisition unit 11, a spatial coordinate acquisition unit 12, a second image acquisition unit 13, a parameter calculation unit 14, an image coordinate acquisition unit 15, an output unit 16, an extraction unit 17, and an identification unit 18. These functional units 11 to 18 may be configured in one device as exemplified in FIG. 1, or may be configured as being distributed across multiple devices.
[0013] Each of the functional units 11 to 18 of the information processing device 10 is configured to be able to access a storage means (storage) such as a virtual space model storage unit 31. The virtual space model storage unit 31 may be provided in the information processing device 10, or may be configured in another device that is configured to be able to access the information processing device 10, as exemplified in FIG.
[0014] The information processing device 10 is configured to acquire a first image, which is an image of real space captured by a first camera c1 installed in a target real space. The first camera c1 may be, for example, a camera permanently installed in real space, such as a surveillance camera in a store. The resolution of the first image captured by the first camera c1, which is a surveillance camera, is often low, making it difficult to directly estimate the target object based on the first image.
[0015] The information processing device 10 is also configured to acquire a second image, which is an image captured by a second camera c2 capable of capturing an image of a target real space. In this embodiment, the second camera c2 may be, for example, a smartphone camera operated by a user. In this embodiment, the second camera c2 has a resolution sufficient to identify an object depicted in the second image.
[0016] Fig. 2 is a sub-block diagram showing the functional configuration of the parameter calculation unit 14. As shown in Fig. 2, the parameter calculation unit 14 includes a setting unit 21, a virtual space image acquisition unit 22, a feature point extraction unit 23, a feature point matching unit 24, a calculation unit 25, and a resetting unit 26. The functions of these functional units 21 to 26 will be described in detail later.
[0017] Referring again to Fig. 1, the functional units of the information processing device 10 will be described. The first image acquisition unit 11 acquires a first image, which is an image of the target real space captured by the first camera c1. Fig. 3 is a diagram schematically illustrating an example of the real space and an example of the first image captured by the surveillance camera.
[0018] 3, a first camera c1, which is a surveillance camera, is installed in a real space rs, which may be a store, for example. The first image acquisition unit 11 acquires a first image rp1 captured from the real space rs.
[0019] The spatial coordinate acquisition unit 12 acquires focus position spatial coordinates, which are the three-dimensional coordinates of the focus position in the virtual space, estimated by reflecting the focus position in the real space rs shown in the first image rp1 in a virtual space model representing the virtual space corresponding to the real space rs.
[0020] The focus position can be set as a position where various events occur in the first image rp1 representing the real space rs. The focus position may be, for example, the position of an object touched by a person in the real space with their hand. Alternatively, the focus position may be the position of an object in the person's line of sight.
[0021] A virtual space model for estimating the spatial coordinates of the focus position may be generated in advance. Fig. 4 is a diagram schematically illustrating an example of a virtual space model representing a virtual space corresponding to the real space rs. The virtual space model vm may be generated based on an image captured of the real space rs by any known method. The virtual space model vm may also be generated based on the first image rp1.
[0022] The virtual space model vm may be generated based on the distance to an object in an image of real space measured by LiDAR (Light Detection and Ranging), for example. Alternatively, the virtual space model vm may be generated by estimating the depth of an object represented by each pixel in an image of two-dimensional real space using known technology, converting the image into point cloud data based on the depth and color value of each pixel, and generating a three-dimensional virtual space by a three-dimensional display based on the point cloud data. Alternatively, the virtual space model vm may be generated using known photogrammetry technology based on multiple images captured of the real space.
[0023] The generated virtual space model vm may be stored in the virtual space model storage unit 31. The virtual space model storage unit 31 is a storage unit that stores the virtual space model vm that has been generated in advance.
[0024] As an example, if an object touched by a person's hand is taken as the target of attention and the position of that object is taken as the target position, the spatial coordinate acquisition unit 12 extracts the two-dimensional position of the target of attention, person h1, based on the first image rp1 and further estimates the depth of the extracted person h1. This defines the position of person h1 in three-dimensional space. Then, the spatial coordinate acquisition unit 12 transfers person h1 to the virtual space model vm based on the estimated position of person h1 in three-dimensional space.
[0025] The spatial coordinate acquisition unit 12 detects the position of an object touched by the person h1 by determining whether the person h2 transferred to the virtual space model vm has collided with an object included in the virtual space model vm. In the example shown in Figure 4, the spatial coordinate acquisition unit 12 detects a focus position ap by determining whether the person h2 has collided with a product object, and acquires the three-dimensional coordinates of the focus position ap in the space represented by the virtual space model vm as focus position spatial coordinates. Note that known techniques can be applied to the processing at each stage for detecting the focus position ap.
[0026] The spatial coordinate acquisition unit 12 may estimate the gaze direction of the persons h1 and h2 using known technology, detect the attention position ap by determining whether the estimated gaze direction collides with an object in the virtual space model vm, and acquire the three-dimensional coordinates of the detected attention position ap as the attention position spatial coordinates.
[0027] The second image acquisition unit 13 acquires a second image captured by the second camera c2 of the real space rs, the second image including the position of interest ap. Fig. 5 is a diagram illustrating an example of how the second image is captured in the real space rs and the captured second image.
[0028] The second camera c2 may be a camera of a smartphone operated by the user, as exemplified in Fig. 5. The second camera c2 has a resolution sufficient to identify an object depicted in the second image. The user captures an image of real space using the second camera c2, including the focus position ap. The second image acquisition unit 13 then acquires a second image rp2 including the focus position ap.
[0029] The parameter calculation unit 14 calculates transformation parameters in a predetermined transformation formula that expresses the relationship between the three-dimensional coordinates of a specific point in the real space rs and the two-dimensional coordinates of the specific point in the second image rp2 captured from the real space rs using predetermined transformation parameters.
[0030] The calculation process of the transformation parameters by the parameter calculation unit 14 will be described with reference to Fig. 2. As shown in Fig. 2, the parameter calculation unit 14 includes a setting unit 21, a virtual space image acquisition unit 22, a feature point extraction unit 23, a feature point matching unit 24, a calculation unit 25, and a resetting unit 26.
[0031] The setting unit 21 sets up a virtual camera in the virtual space represented by the virtual space model vm. Specifically, the setting unit 21 sets the position of the virtual camera in the virtual space represented by the virtual space model vm. The position of the virtual camera is a viewpoint position for capturing a virtual space image, which is an image projected from the virtual space model vm.
[0032] 6A and 6B are diagrams showing examples of setting up virtual cameras in a virtual space, and Fig. 6A and Fig. 6B respectively show first and second examples of setting positions of virtual cameras in the virtual space as viewed from above the virtual space vs. As shown in Fig. 6A and Fig. 6B, the setting unit 21 may set up multiple virtual cameras v in the virtual space vs.
[0033] In a first example shown in Fig. 6(a), the setting unit 21 may install multiple virtual cameras vc along the periphery of the virtual space vs, with the normal direction of the periphery as the imaging direction. Alternatively, as shown in Fig. 6(b), the setting unit 21 may install multiple virtual cameras vc within the virtual space vs (for example, around the center of the virtual space vs), with the periphery direction of the virtual space vs as the imaging direction. By installing multiple virtual cameras vc in this way, it is possible to comprehensively capture images of the virtual space corresponding to the space represented in the real-space image rp.
[0034] The virtual space image acquisition unit 22 acquires at least one virtual space image captured by the virtual camera vc of the virtual space vs. The virtual space image acquisition unit 22 may acquire the virtual space image by a known method based on the virtual space vs represented by the virtual space model vm.
[0035] 7 is a diagram showing an example of a virtual space image, an example of a real space image, and an example of feature point extraction and matching processing. Specifically, the virtual space image acquisition unit 22 acquires a virtual space image vp by projecting the virtual space vs onto a virtual screen, using the position of a virtual camera vc set in the virtual space vs represented based on the virtual space model vm as the viewpoint position.
[0036] The feature point extraction unit 23 extracts feature points from a second image rp2, which is a real space image captured of the real space rs, and from at least one virtual space image vp2. The feature point matching unit 24 matches the feature points of the second image rp2 with the feature points of the virtual space image vp2. The feature point extraction unit 23 and the feature point matching unit 24 may extract and match the feature points using known methods, such as SIFT and AKAZE.
[0037] 7 , the feature point extraction unit 23 extracts feature points vfp1 and vfp2 from the virtual space image vp2. The feature point extraction unit 23 also extracts feature points rfp1 and rfp2 from the second image rp2. The feature points are extracted, for example, based on the difference in pixel values between adjacent pixels in the image. For example, the endpoints and corners of an object represented in the image are extracted as feature points.
[0038] Furthermore, the feature point matching unit 24 matches each of the feature points vfp1 and vfp2 of the virtual space image vp2 with each of the feature points rfp1 and rfp2 of the second image rp2 based on the feature amounts of each feature point, as shown by the symbols fm1 and fm2.
[0039] The calculation unit 25 calculates transformation parameters related to imaging of the real space by substituting, into a predetermined transformation formula, reference space coordinates, which are the three-dimensional coordinates in the virtual space vs of the reference feature point, which is the matched feature point, and reference image coordinates, which are the coordinates of the reference feature point in the second image rp2.
[0040] The transformation formula is a formula that expresses, using predetermined transformation parameters, the relationship between the three-dimensional coordinates of a specific point in a three-dimensional space and the two-dimensional coordinates of the specific point in a captured image of the three-dimensional space. If the three-dimensional coordinates in the three-dimensional space are (X, Y, Z) and the two-dimensional coordinates in the two-dimensional image are (u, v), an example of the predetermined transformation formula is expressed by the following transformation formula (1): The above conversion formula (1) is based on the camera external parameters R (3 degrees of freedom), t (3 degrees of freedom), and the camera internal parameter f x , f y , c x , c y , and a parameter k related to the lens distortion of the camera 1 , k 2 , k 3 , p 1 , p 2 as a transformation parameter.
[0041] Then, assuming that the three-dimensional reference space coordinate is t, the virtual reference image coordinate, which is the two-dimensional coordinate of the reference feature point in the virtual space image vp, is d', and the transformation matrix that is expressed by the transformation parameters of transformation equation (1) and transforms the reference space coordinate t into the virtual reference image coordinate d' is A, the relationship between these coordinates is expressed by the following equation (2).
[0042] d' = tA (2) Furthermore, if the two-dimensional reference image coordinate is d and the transformation matrix that transforms the reference space coordinate t into the reference image coordinate d, expressed by the transformation parameters of transformation equation (1), is B, the relationship between these coordinates is expressed as in the following equation (3).
[0043] d=tB (3) As shown in equation (2), the transformation matrix A is a transformation matrix (transformation parameter) that indicates the relationship between three-dimensional coordinates indicating a position in the virtual space vs and two-dimensional coordinates of the position in an image captured by the virtual camera v placed in the virtual space, and is therefore known. Therefore, the calculation unit 25 can calculate the reference space coordinate t based on the transformation matrix A and the virtual reference image coordinate d'.
[0044] Specifically, the calculation unit 25 may calculate the reference space coordinate t using a technique known as hit determination. Fig. 8 is a diagram schematically illustrating an example of a calculation process of the reference space coordinate by hit determination.
[0045] 8 , since the transformation matrix A is known and therefore the position of the virtual camera v is also known, the calculation unit 25 can calculate a three-dimensional straight line ln that passes through the virtual reference image coordinate d′ of the reference feature point fp in the virtual space image vp. Then, the calculation unit 25 calculates the point where the straight line ln intersects with the plane (mesh) that configures the virtual space model vm as the reference space coordinate t of the reference feature point fp.
[0046] The calculation unit 25 calculates the transformation matrix B in equation (3) as a transformation parameter related to imaging of the real space, based on the reference space coordinate t and the reference image coordinate d of the reference feature point fp.
[0047] 6(a) and 6(b), the setting unit 21 may set up multiple virtual cameras vc in the virtual space vs. Then, the virtual space image acquisition unit 22 acquires multiple virtual space images vp acquired by each of the multiple virtual cameras vc.
[0048] The calculation unit 25 may calculate transformation parameters related to imaging of real space based on at least one virtual space image vp selected from the plurality of virtual space images vp based on the number of reference feature points fp.
[0049] Since the number of transformation parameters in transformation formula (1) constituting transformation matrix B is 15, the calculation unit 25 constructs transformation formula (1) using the reference space coordinates t and reference image coordinates d of at least eight reference feature points fp, and calculates the transformation parameters by solving the constructed multiple transformation formulas as equations. Therefore, the calculation unit 25 may calculate transformation parameters related to imaging of real space based on a virtual space image vp having a number of reference feature points fp equal to or greater than a predetermined threshold.
[0050] Specifically, the calculation unit 25 may calculate the transformation parameters related to imaging of the real space using a virtual space image vp having eight or more reference feature points fp. Alternatively, the calculation unit 25 may calculate the transformation parameters related to imaging of the real space using a plurality of virtual space images vp in which the total number of reference feature points fp is equal to or greater than a predetermined threshold.
[0051] Furthermore, the calculation unit 25 may calculate transformation parameters related to capturing real space based on each of the multiple virtual space images vp, thereby calculating transformation parameters associated with each virtual space image vp.
[0052] The calculation unit 25 may use the calculated transformation parameters (corresponding to transformation matrix B) related to capturing the real space as camera information that essentially represents the position in the virtual space of the second camera c2 that captured the real space image.
[0053] The calculation unit 25 may also use the transformation parameters obtained by processing the calculated transformation parameters using a predetermined statistical method as the camera information. Specifically, the calculation unit 25 may use the transformation parameters obtained by processing the calculated transformation parameters using a least squares method or the like as the camera information. In this way, the transformation parameters calculated based on the virtual space images vp acquired by each of the multiple virtual cameras vc are statistically processed to obtain the final output transformation parameters, thereby improving the accuracy of the transformation parameters used as the camera information.
[0054] The resetting unit 26 resets the virtual camera vc in the virtual space vs, to which the transformation parameters for capturing images of the real space calculated by the calculation unit 25 have been applied as transformation parameters for capturing images of the virtual space vs.
[0055] When setting up the initial virtual camera vc and acquiring the virtual space image vp, the transformation parameters applied to the virtual camera vc may be set arbitrarily. After the transformation parameters for capturing the real space are calculated based on the initial virtual space image vp, the calculated transformation parameters can be applied to the virtual camera vc. Once the calculated transformation parameters are applied, the virtual camera vc is likely to have a positional relationship and characteristics similar to those of the camera that captured the second image rp2.
[0056] The virtual space image acquisition unit 22 acquires a virtual space image vp captured by a virtual camera vc reinstalled in the virtual space vs. The feature point extraction unit 23 extracts feature points from the virtual space image vp captured by the reinstalled virtual camera vc, and the feature point matching unit 24 matches the feature points. The calculation unit 25 then calculates transformation parameters based on the virtual space image vp captured by the reinstalled virtual camera vc. In this way, by recalculating the transformation parameters based on the virtual space image vp captured by the virtual camera vc to which the calculated transformation parameters have been applied, it is possible to improve the accuracy of the transformation parameters.
[0057] The resetting unit 26 may reset at least one virtual camera vc within a predetermined range around a specific virtual camera vc, which is the virtual camera vc that captured the virtual space image vp that has the most reference feature points fp, among the virtual cameras vc set by the setting unit 21.
[0058] 9 is a diagram showing an example of virtual camera resetting. In the example shown in Fig. 9, the resetting unit 26 extracts a specific virtual camera vc1, which is the virtual camera that captured the virtual space image vp that has the most reference feature points fp, from among the virtual cameras vc installed in the virtual space vs in the calculation of the initial or previous transformation parameters. The specific virtual camera vc1 is the virtual camera that captured the virtual space image vp that has many feature points that match the feature points in the second image rp2, and therefore is highly likely to be a camera installed in a position in the virtual space vs that corresponds to the position of the camera rc (second camera c2) that captured the second image rp2.
[0059] The resetter 26 resets the virtual camera vc within a predetermined range around the specific virtual camera vc1. For example, the resetter 26 may reset the virtual camera vc to a position within a predetermined distance from the position of the specific virtual camera vc1. As illustrated in Fig. 9, the resetter 26 may reset the virtual cameras vc2 and vc3 to positions adjacent to the specific virtual camera vc1 along the outer periphery of the virtual space vs.
[0060] In this way, by recalculating the transformation parameters based on the virtual space image vp based on the virtual camera vc that has been re-installed within a predetermined range around the specific virtual camera vc1, it is possible to further improve the accuracy.
[0061] Furthermore, the resetting unit 26 may reset at least one virtual camera vc so that a virtual space image vc is captured in which a partial area including the reference feature point fp overlaps with the virtual space image vc captured by the specific virtual camera vc1. Fig. 10 is a diagram showing an example of how to superimpose the captured area between the virtual space image captured by the specific virtual camera and the virtual space image captured by the reset virtual camera.
[0062] 10 , when a virtual space image vp having an imaging area ts0 is captured by a specific virtual camera vc1, the resetter 26 resets the virtual camera vc so as to capture an imaging area that is superimposed on the imaging area ts0 and a partial area that includes the reference feature point fp. Specifically, the resetter 26 resets the virtual camera vc to a position such that imaging areas ts1, ts2, ts3, and ts4, including an area that includes the reference feature point fp located at an edge of the imaging area ts0, are captured as the virtual space image vp. By acquiring the virtual space image vp using the virtual camera vc reset in this way, it is possible to improve the accuracy of the transformation parameters related to lens distortion.
[0063] Referring back to FIG. 1, the image coordinate acquisition unit 15 applies the focus position spatial coordinates to a conversion formula to acquire focus position image coordinates, which are the two-dimensional coordinates of the focus position in the second image.
[0064] Specifically, the image coordinate acquisition unit 15 calculates the focus position image coordinates, which are the two-dimensional coordinates of the focus position ap in the second image rp2, by substituting the focus position spatial coordinates, which are the three-dimensional coordinates of the focus position ap acquired by the spatial coordinate acquisition unit 12, into the transformation formula (formula (1)) to which the transformation parameters calculated by the parameter calculation unit 14 are applied.
[0065] The output unit 16 outputs the image coordinates of the target position acquired by the image coordinate acquisition unit 15 as identification information for identifying the item present at the target position. By outputting the identification information, the position where the target of interest is represented in the second image rp2 can be identified.
[0066] The extraction unit 17 extracts a specific image, which is an image of an area including the position of the focus position image coordinates, from the second image rp2. Specifically, the extraction unit 17 may extract an image of a predetermined range including the position specified by the focus position image coordinates in the second image rp2 as the specific image.
[0067] In addition, the extraction unit 17 may extract an object present at a position specified by the focus position image coordinates in the second image rp2 using a known image processing technique, and generate an image of the extracted object as a specific image.
[0068] The output unit 16 outputs the specific image extracted by the extraction unit 17. As a result, the object present at the target position is depicted in the specific image, and therefore the target object can be identified based on the specific image.
[0069] The identification unit 18 acquires item identification information identified by reference based on a specific image of item information that associates item identification information that identifies an item with appearance information that represents the appearance of the item.
[0070] Specifically, the identification unit 18 may refer to item information consisting of a table that associates item identification information, such as item name and product name, with an appearance image of the item, extract an appearance image that is similar to the appearance of the item represented by the identified image to a certain degree or more, and obtain item identification information associated with the extracted appearance image.
[0071] In addition, the identification unit 18 may use image recognition technology to recognize objects depicted in an image and obtain item identification information based on the identified image using a service that provides information related to the object (e.g., Google Lens (registered trademark)).
[0072] The output unit 16 outputs the item identification information acquired by the identification unit 18. By outputting the item identification information, the item present at the target position can be identified.
[0073] 11 is a flowchart showing the processing content of the information processing method in the information processing system 1. In step S1, the first image acquisition unit 11 acquires a first image rp1 that is an image of the real space of the target captured by the first camera c1.
[0074] In step S2, the spatial coordinate acquisition unit 12 acquires the attention position ap in the real space rs represented in the first image rp1. Then, in step S3, the spatial coordinate acquisition unit 12 acquires attention position spatial coordinates, which are the three-dimensional coordinates of the attention position ap in the virtual space vs, estimated by reflecting the attention position ap in a virtual space model vm representing the virtual space vs corresponding to the real space rs.
[0075] In step S4, the second image acquisition unit 13 acquires a second image rp2 that is an image of the real space rs captured by the second camera c2 and includes the attention position ap.
[0076] In step S5, the parameter calculation unit 14 calculates transformation parameters in a predetermined transformation formula (formula (1)) that expresses the relationship between the three-dimensional coordinates of a specific point, which is a specific point in the real space rs, and the two-dimensional coordinates of the specific point in the second image rp2 that captures the real space rs, using predetermined transformation parameters.
[0077] In step S6, the image coordinate acquisition unit 15 applies the focus position spatial coordinates to a conversion formula to acquire focus position image coordinates, which are the two-dimensional coordinates of the focus position ap in the second image rp2.
[0078] In step S7, the output unit 16 outputs the image coordinates of the target position acquired by the image coordinate acquisition unit 15 as identification information for identifying the article present at the target position.
[0079] Next, an information processing program for causing a computer to function as the information processing device 10 of this embodiment will be described with reference to Fig. 12. Fig. 12 is a diagram showing the configuration of the information processing program. The information processing program P1 is configured to include a main module m10 that controls information processing in the information processing device 10 in an overall manner, a first image acquisition module m11, a spatial coordinate acquisition module m12, a second image acquisition module m13, a parameter calculation module m14, an image coordinate acquisition module m15, an output module m16, an extraction module m17, and an identification module m18. Each of the modules m11 to m18 realizes a function for each of the functional units 11 to 18.
[0080] The information processing program P1 may be transmitted via a transmission medium such as a communication line, or may be stored in a recording medium M1 as shown in FIG.
[0081] According to the information processing system, information processing device 10, information processing method, and information processing program P1 of the present embodiment described above, transformation parameters of a transformation formula expressing the relationship between the three-dimensional coordinates of a specific point in real space and the two-dimensional coordinates of the specific point in the second image are calculated, and information indicating the focus position in the second image is obtained by applying the focus position space coordinates estimated by reflecting the focus position acquired based on the first image in the virtual space model to the transformation formula. Therefore, by using a camera capable of capturing a second image with high resolution enough to identify an object represented in the image as the second camera, it is possible to identify the object represented at the focus position.
[0082] The information processing device and information processing method according to the present disclosure may have the following configurations: The actions and effects of each configuration will be described as follows.
[0083] An information processing device according to one aspect of the present disclosure includes a spatial coordinate acquisition unit that acquires focus position spatial coordinates, which are the three-dimensional coordinates of the focus position in virtual space, estimated by reflecting the focus position in real space shown in a first image, which is an image of real space captured by a first camera, in a virtual space model that represents a virtual space corresponding to the real space; a second image acquisition unit that acquires a second image, which is an image of real space captured by a second camera and includes the focus position; a parameter calculation unit that calculates conversion parameters in a predetermined conversion formula that expresses, using predetermined conversion parameters, the relationship between the three-dimensional coordinates of a specific point, which is a specific point in real space, and the two-dimensional coordinates of the specific point in the second image captured of real space; and an image coordinate acquisition unit that acquires focus position image coordinates, which are the two-dimensional coordinates of the focus position in the second image, by applying the focus position spatial coordinates to the conversion formula.
[0084] An information processing method according to one aspect of the present disclosure includes a spatial coordinate acquisition step executed by a processor to acquire focus position spatial coordinates, which are three-dimensional coordinates of the focus position in virtual space, estimated by reflecting the focus position in real space shown in a first image, which is an image of real space captured by a first camera, in a virtual space model that represents a virtual space corresponding to the real space; a second image acquisition step to acquire a second image, which is an image of real space captured by a second camera and includes the focus position; a parameter calculation step to calculate transformation parameters in a predetermined transformation formula that expresses, using predetermined transformation parameters, the relationship between the three-dimensional coordinates of a specific point, which is a specific point in real space, and the two-dimensional coordinates of the specific point in the second image captured of real space; and an image coordinate acquisition step to acquire focus position image coordinates, which are two-dimensional coordinates of the focus position in the second image, by applying the focus position spatial coordinates to the transformation formula.
[0085] According to the above aspect, transformation parameters of a transformation formula expressing the relationship between the three-dimensional coordinates of a specific point in real space and the two-dimensional coordinates of the specific point in the second image are calculated, and information indicating the focus position in the second image is obtained by applying the focus position space coordinates estimated by reflecting the focus position acquired based on the first image in the virtual space model to the transformation formula. Therefore, by using a camera capable of capturing a second image with a resolution high enough to identify an object represented in the image as the second camera, it is possible to identify the object represented at the focus position.
[0086] Furthermore, an information processing device according to another aspect may further include an extraction unit that extracts, from the second image, a specific image that is an image of an area including the position of the focus position image coordinates.
[0087] According to the above aspect, since the object present at the target position is depicted in the specific image, the target object can be identified based on the specific image.
[0088] In addition, an information processing device according to another aspect may further include an identification unit that acquires item identification information identified by reference to a specific image of item information that associates item identification information that identifies an item with appearance information that represents the appearance of the item.
[0089] According to the above aspect, the item identification information is identified based on the appearance of the item depicted in the specific image, so that the object present at the target position can be identified.
[0090] In addition, in an information processing device according to another aspect, the parameter calculation unit may include: a setting unit that sets a virtual camera as a viewpoint position for capturing an image in which a virtual space model is projected in the virtual space; a feature point extraction unit that extracts feature points from each of at least one virtual space image captured of the virtual space by the virtual camera and a second image; a feature point matching unit that matches feature points of the second image with feature points of the virtual space image; and a calculation unit that calculates transformation parameters for capturing the second image by substituting reference space coordinates in the virtual space of a reference feature point that is the matched feature point and reference image coordinates that are coordinates of the reference feature point in the second image into a transformation formula, wherein the reference space coordinates are calculated based on the transformation parameters for capturing the virtual space by the virtual camera and virtual reference image coordinates that are coordinates of the reference feature point in the virtual space image.
[0091] According to the above aspect, feature points extracted from each of the real space image and the virtual space image are matched as reference feature points, and the two-dimensional reference image coordinates of the reference feature points in the real space image are associated with the three-dimensional reference space coordinates of the reference feature points in the virtual space, thereby enabling calculation of transformation parameters in the transformation formula.
[0092] In another aspect of the information processing device, the transformation parameters include six predetermined external camera parameters related to the projection of the three-dimensional space onto the two-dimensional image, four internal camera parameters, and five parameters related to the lens distortion of the camera, and the calculation unit calculates the following transformation formula representing the relationship between three-dimensional coordinates (X, Y, Z) in the three-dimensional space and two-dimensional coordinates (u, v) in the two-dimensional image: In this case, by substituting the three-dimensional reference space coordinates into the three-dimensional coordinates (X, Y, Z) and the two-dimensional reference image coordinates into the two-dimensional coordinates (u, v), the external camera parameters R and t and the internal camera parameters f each have three degrees of freedom. x , fy , c x , c y , and a parameter k related to the lens distortion of the camera 1 , k 2 , k 3 , p 1 , p 2 , may be calculated.
[0093] According to the above aspect, it is possible to calculate transformation parameters consisting of 15 variables as camera information.
[0094] In an information processing device according to another aspect, the calculation unit may calculate the transformation parameters based on a virtual space image having a number of reference specific points equal to or greater than a threshold value related to the reference feature points.
[0095] According to the above aspect, the threshold is set according to the number of unknown transformation parameters in the transformation formula, so that the transformation parameters can be reliably found.
[0096] In an information processing device according to another aspect, the focus position may be a collision position between a part of a person depicted in the first image and an object.
[0097] According to the above aspect, a collision between a part of a person and an object indicates that the person has touched the object, and therefore the thing that the person has touched can be identified as the target of attention.
[0098] In an information processing device according to another aspect, the focus position may be a collision position between the line of sight of a person depicted in the first image and an object.
[0099] According to the above aspect, an object that a person focuses on can be identified as an object of attention.
[0100] The block diagram shown in FIG. 1 shows functional blocks. These functional blocks (components) are realized by any combination of hardware and / or software. Furthermore, the method for realizing each functional block is not particularly limited. That is, each functional block may be realized using a single device that is physically or logically coupled, or may be realized using two or more physically or logically separated devices that are directly or indirectly connected (e.g., wired, wireless, etc.) and these multiple devices. The functional block may also be realized by combining software with the single device or multiple devices.
[0101] Functions include, but are not limited to, judgment, determination, assessment, calculation, computation, processing, derivation, investigation, search, confirmation, reception, transmission, output, access, resolution, selection, selection, establishment, comparison, assumption, expectation, consideration, broadcasting, notifying, communicating, forwarding, configuring, reconfiguring, allocating, mapping, and assignment. For example, a functional block (component) that performs transmission is called a transmitting unit or transmitter. As mentioned above, there are no particular limitations on how these functions are implemented.
[0102] For example, the information processing device 10 according to an embodiment of the present invention may function as a computer. Fig. 13 is a diagram showing an example of the hardware configuration of the information processing device 10 according to this embodiment. The information processing device 10 may be physically configured as a computer device including a processor 1001, a memory 1002, a storage 1003, a communication device 1004, an input device 1005, an output device 1006, a bus 1007, etc.
[0103] In the following description, the term "apparatus" can be interpreted as a circuit, a device, a unit, etc. The hardware configuration of the information processing device 10 may be configured to include one or more of the apparatuses shown in Fig. 13, or may be configured to exclude some of the apparatuses.
[0104] Each function of the information processing device 10 is realized by loading specified software (programs) onto hardware such as the processor 1001 and memory 1002, causing the processor 1001 to perform calculations and control communication via the communication device 1004 and the reading and / or writing of data in the memory 1002 and storage 1003.
[0105] The processor 1001 controls the entire computer by running, for example, an operating system. The processor 1001 may be configured as a central processing unit (CPU) including an interface with peripheral devices, a control device, an arithmetic unit, a register, etc. For example, the functional units 11 to 18 shown in FIG. 1 may be realized by the processor 1001.
[0106] The processor 1001 also reads programs (program codes), software modules, and data from the storage 1003 and / or the communication device 1004 into the memory 1002 and executes various processes in accordance with these. The programs used are those that cause a computer to execute at least some of the operations described in the above-described embodiments. For example, the functional units 11 to 18 of the information processing device 10 may be implemented by a control program stored in the memory 1002 and running on the processor 1001. While the above-described various processes have been described as being executed by one processor 1001, they may also be executed simultaneously or sequentially by two or more processors 1001. The processor 1001 may be implemented on one or more chips. The programs may also be transmitted from a network via a telecommunications line.
[0107] The memory 1002 is a computer-readable recording medium and may be composed of at least one of, for example, a read-only memory (ROM), an erasable programmable ROM (EPROM), an electrically erasable programmable ROM (EEPROM), and a random access memory (RAM). The memory 1002 may also be called a register, a cache, a main memory (primary storage device), or the like. The memory 1002 can store executable programs (program codes), software modules, and the like for implementing an information processing method according to one embodiment of the present invention.
[0108] Storage 1003 is a computer-readable recording medium, and may be, for example, at least one of an optical disk such as a CD-ROM (Compact Disc ROM), a hard disk drive, a flexible disk, a magneto-optical disk (e.g., a compact disk, a digital versatile disk, a Blu-ray (registered trademark) disk), a smart card, a flash memory (e.g., a card, a stick, a key drive), a floppy (registered trademark) disk, a magnetic strip, etc. Storage 1003 may also be referred to as an auxiliary storage device. The above-mentioned storage medium may be, for example, a database, a server, or other appropriate medium including memory 1002 and / or storage 1003.
[0109] The communication device 1004 is hardware (transmission / reception device) for communicating between computers via a wired and / or wireless network, and is also called, for example, a network device, a network controller, a network card, or a communication module.
[0110] The input device 1005 is an input device (e.g., a keyboard, a mouse, a microphone, a switch, a button, a sensor, etc.) that receives input from the outside. The output device 1006 is an output device (e.g., a display, a speaker, an LED lamp, etc.) that outputs to the outside. The input device 1005 and the output device 1006 may be integrated into one device (e.g., a touch panel).
[0111] Furthermore, each device such as the processor 1001 and the memory 1002 is connected to a bus 1007 for communicating information. The bus 1007 may be configured as a single bus, or may be configured as different buses between the devices.
[0112] The information processing device 10 may also be configured to include hardware such as a microprocessor, a digital signal processor (DSP), an application specific integrated circuit (ASIC), a programmable logic device (PLD), or a field programmable gate array (FPGA), and some or all of the functional blocks may be realized by the hardware. For example, the processor 1001 may be implemented by at least one of these pieces of hardware.
[0113] The notification of information is not limited to the aspects / embodiments described in the present disclosure and may be performed using other methods. For example, the notification of information may be performed by physical layer signaling (e.g., Downlink Control Information (DCI) and Uplink Control Information (UCI)), higher layer signaling (e.g., Radio Resource Control (RRC) signaling, Medium Access Control (MAC) signaling, broadcast information (Master Information Block (MIB) and System Information Block (SIB))), other signals, or a combination thereof. Furthermore, the RRC signaling may be referred to as an RRC message, and may be, for example, an RRC Connection Setup message, an RRC Connection Reconfiguration message, or the like.
[0114] Each aspect / embodiment described in the present disclosure may be applied to at least one of systems using LTE (Long Term Evolution), LTE-Advanced (LTE-A), SUPER 3G, IMT-Advanced, 4G (4th generation mobile communication system), 5G (5th generation mobile communication system), FRA (Future Radio Access), NR (New Radio), W-CDMA (registered trademark), GSM (registered trademark), CDMA2000, UMB (Ultra Mobile Broadband), IEEE 802.11 (Wi-Fi (registered trademark)), IEEE 802.16 (WiMAX (registered trademark)), IEEE 802.20, UWB (Ultra-Wide Band), Bluetooth (registered trademark), or other suitable systems, and next-generation systems enhanced based on these. Furthermore, a combination of multiple systems (e.g., a combination of at least one of LTE and LTE-A with 5G, etc.) may also be applied.
[0115] The order of the procedures, sequences, flowcharts, etc. of each aspect / embodiment described in this disclosure may be changed unless it is consistent. For example, the methods described in this disclosure present elements of various steps using an example order, and are not limited to the particular order presented.
[0116] In the present disclosure, a specific operation described as being performed by a base station may be performed by its upper node in some cases. In a network consisting of one or more network nodes having a base station, it is clear that various operations performed for communication with a terminal may be performed by at least one of the base station and another network node other than the base station (for example, an MME or an S-GW, etc., but are not limited to these). Although the above example illustrates a case where there is one other network node other than the base station, a combination of multiple other network nodes (for example, an MME and an S-GW) may also be used.
[0117] Information etc. may be output from a higher layer (or a lower layer) to a lower layer (or a higher layer), or may be input / output via multiple network nodes.
[0118] Input and output information may be stored in a specific location (for example, memory) or managed in a management table. Input and output information may be overwritten, updated, or added to. Output information may be deleted. Input information may be sent to another device.
[0119] The determination may be made based on a value represented by one bit (0 or 1), a Boolean value (true or false), or a numerical comparison (e.g., comparison with a predetermined value).
[0120] The aspects / embodiments described in this disclosure may be used alone, in combination, or switched depending on the implementation. Notification of predetermined information (e.g., notification that "X is true") is not limited to explicit notification, but may be implicit (e.g., not notifying the predetermined information).
[0121] Although the present disclosure has been described in detail above, it is clear to those skilled in the art that the present disclosure is not limited to the embodiments described herein. The present disclosure can be implemented in modified and altered forms without departing from the spirit and scope of the present disclosure as defined by the claims. Therefore, the description of the present disclosure is intended to be illustrative and does not have any limiting meaning on the present disclosure.
[0122] Software shall be construed broadly to mean instructions, instruction sets, code, code segments, program code, programs, subprograms, software modules, applications, software applications, software packages, routines, subroutines, objects, executable files, threads of execution, procedures, functions, etc., whether referred to as software, firmware, middleware, microcode, hardware description language, or otherwise.
[0123] Software, instructions, etc. may also be transmitted or received over a transmission medium. For example, if the software is transmitted from a website, server, or other remote source using wired technologies such as coaxial cable, fiber optic cable, twisted pair, and Digital Subscriber Line (DSL), and / or wireless technologies such as infrared, radio, and microwave, these wired and / or wireless technologies are included within the definition of transmission media.
[0124] The information, signals, etc. described in this disclosure may be represented using any of a variety of different technologies. For example, data, instructions, commands, information, signals, bits, symbols, chips, etc. that may be referred to throughout the above description may be represented by voltages, currents, electromagnetic waves, magnetic fields or magnetic particles, optical fields or photons, or any combination thereof.
[0125] It should be noted that terms explained in this disclosure and / or terms necessary for understanding this specification may be replaced with terms having the same or similar meanings.
[0126] As used in this disclosure, the terms "system" and "network" are used interchangeably.
[0127] Furthermore, the information, parameters, etc. described in the present disclosure may be expressed as absolute values, relative values from a predetermined value, or other corresponding information. For example, a radio resource may be indicated by an index.
[0128] The names used for the above-described parameters are not intended to be limiting in any way. Furthermore, the mathematical expressions using these parameters may differ from those explicitly disclosed in this disclosure. The various channels (e.g., PUCCH, PDCCH, etc.) and information elements may be identified by any suitable names, and therefore the various names assigned to these various channels and information elements are not intended to be limiting in any way.
[0129] As used in this disclosure, the terms "determining" and "determining" may encompass a wide variety of actions. "Determining" and "determining" may include, for example, judging, calculating, computing, processing, deriving, investigating, looking up, searching, inquiring (e.g., searching in a table, database, or other data structure), ascertaining, and the like. "Determining" and "determining" may also include receiving (e.g., receiving information), transmitting (e.g., sending information), input, output, accessing (e.g., accessing data in memory), and the like. Furthermore, "judgment" and "decision" can include regarding resolving, selecting, choosing, establishing, comparing, etc. as having been "judged" or "decided." In other words, "judgment" and "decision" can include regarding some action as having been "judged" or "decided." Furthermore, "judgment (decision)" can be interpreted as "assuming," "expecting," "considering," etc.
[0130] As used in this disclosure, the phrase "based on" does not mean "based only on," unless expressly specified otherwise. In other words, the phrase "based on" means both "based only on" and "based at least on."
[0131] When designations such as "first," "second," etc. are used in this disclosure, any reference to an element does not generally limit the quantity or order of those elements. These designations may be used herein as a convenient method of distinguishing between two or more elements. Thus, a reference to a first and a second element does not imply that only two elements may be employed therein or that the first element must precede the second element in some way.
[0132] The "means" in the configuration of each of the above devices may be replaced with "part," "circuit," "device," etc.
[0133] To the extent that the terms "include," "including," and variations thereof are used herein or in the claims, these terms are intended to be inclusive, similar to the term "comprising." Furthermore, the term "or," as used herein or in the claims, is not intended to be an exclusive or.
[0134] In this disclosure, where articles are added by translation, such as a, an, and the in English, the disclosure may include that the nouns following these articles are in the plural form.
[0135] In the present disclosure, the term "A and B are different" may mean "A and B are different from each other." The term may also mean "A and B are each different from C." Terms such as "separate" and "coupled" may also be interpreted in the same way as "different."
[0136] An information processing system 1 of the present disclosure may have the following configuration: [1] An information processing device comprising: a space coordinate acquisition unit that acquires focus position space coordinates, which are three-dimensional coordinates of the focus position in virtual space, estimated by reflecting a focus position in a first image, which is an image of real space captured by a first camera, in a virtual space model that represents a virtual space corresponding to the real space; a second image acquisition unit that acquires a second image, which is an image of the real space captured by a second camera, and includes the focus position; a parameter calculation unit that calculates transformation parameters in a predetermined transformation formula that expresses, using the transformation parameters, a relationship between the three-dimensional coordinates of a specific point, which is a specific point in the real space, and the two-dimensional coordinates of the specific point in the second image captured of the real space; and an image coordinate acquisition unit that acquires focus position image coordinates, which are two-dimensional coordinates of the focus position in the second image, by applying the focus position space coordinates to the transformation formula. [2] The information processing device described in [1], further comprising: an extraction unit that extracts a specific image, which is an image of an area including the focus position image coordinates, from the second image. [3] The information processing device according to [2], further comprising an identification unit that acquires identified item identification information by reference based on the specific image of item information that associates item identification information that identifies an item with appearance information that represents the appearance of the item.[4] The information processing device according to any one of [1] to [3], wherein the parameter calculation unit includes: a setting unit that sets a virtual camera as a viewpoint position for capturing an image in which the virtual space model is projected in the virtual space; a feature point extraction unit that extracts feature points from each of at least one virtual space image captured of the virtual space by the virtual camera and the second image; a feature point matching unit that matches feature points of the second image with feature points of the virtual space image; and a calculation unit that calculates the transformation parameters related to capturing the second image by substituting reference space coordinates in the virtual space of reference feature points that are the matched feature points and reference image coordinates that are coordinates of the reference feature points in the second image into the transformation formula, wherein the reference space coordinates are calculated based on the transformation parameters related to capturing the virtual space by the virtual camera and virtual reference image coordinates that are coordinates of the reference feature points in the virtual space image. [5] The transformation parameters include six predetermined external camera parameters, four internal camera parameters, and five parameters related to the lens distortion of the camera, related to the projection of the three-dimensional space onto the two-dimensional image, and the calculation unit uses the following transformation formula that represents the relationship between three-dimensional coordinates (X, Y, Z) in the three-dimensional space and two-dimensional coordinates (u, v) in the two-dimensional image: In the above, by substituting the three-dimensional reference space coordinates into the three-dimensional coordinates (X, Y, Z) and the two-dimensional reference image coordinates into the two-dimensional coordinates (u, v), the camera external parameters R, t and the camera internal parameters f each having three degrees of freedom are obtained. x , f y , c x , c y , and a parameter k relating to the lens distortion of the camera 1 , k 2 , k 3 , p 1 , p 2, and calculates the transformation parameters based on the virtual space image having a number of reference specific points equal to or greater than a threshold value for the reference feature points. [6] The information processing device according to [4] or [5], wherein the calculation unit calculates the transformation parameters based on the virtual space image having a number of reference specific points equal to or greater than a threshold value for the reference feature points. [7] The information processing device according to any one of [1] to [6], wherein the focus position is a collision position between a part of a person depicted in the first image and an object. [8] The information processing device according to any one of [1] to [6], wherein the focus position is a collision position between a line of sight of the person depicted in the first image and an object. [9] An information processing method executed by a processor, comprising: a spatial coordinate acquisition step of acquiring focus position spatial coordinates, which are three-dimensional coordinates of the focus position in the virtual space, estimated by reflecting a focus position in the real space shown in a first image, which is an image of the real space captured by a first camera, in a virtual space model that represents a virtual space corresponding to the real space; a second image acquisition step of acquiring a second image, which is an image of the real space captured by a second camera and includes the focus position; a parameter calculation step of calculating a transformation parameter in a predetermined transformation equation that expresses, using the transformation parameter, the relationship between the three-dimensional coordinates of a specific point, which is a specific point in the real space, and the two-dimensional coordinates of the specific point in the second image captured of the real space; and an image coordinate acquisition step of acquiring focus position image coordinates, which are two-dimensional coordinates of the focus position in the second image, by applying the focus position spatial coordinates to the transformation equation.
[0137] 1...information processing system, 10...information processing device, 11...first image acquisition unit, 12...spatial coordinate acquisition unit, 13...second image acquisition unit, 14...parameter calculation unit, 15...image coordinate acquisition unit, 16...output unit, 17...extraction unit, 18...identification unit, 21...setting unit, 22...virtual space image acquisition unit, 23...feature point extraction unit, 24...feature point matching unit, 25...calculation unit, 26...resetting unit, 31...virtual space model storage unit, c1...first camera, c2...second camera, ln...straight line, M1...recording medium, m11...first image acquisition module, m12...spatial coordinate acquisition module, m13...second image acquisition module, m14...parameter calculation module, m15...image coordinate acquisition module, m16...output module, m17...extraction module, m18...identification module, P1...information processing program, rp...real space image, rs...real space, vm...virtual space model.
Claims
1. An information processing device comprising: a spatial coordinate acquisition unit that acquires focus position spatial coordinates, which are the three-dimensional coordinates of the focus position in virtual space, estimated by reflecting a focus position in a first image, which is an image of real space captured by a first camera, onto a virtual space model that represents a virtual space corresponding to the real space; a second image acquisition unit that acquires a second image, which is an image of real space captured by a second camera and includes the focus position; a parameter calculation unit that calculates a transformation parameter in a predetermined transformation equation that expresses, using the transformation parameter, the relationship between the three-dimensional coordinates of a specific point, which is a specific point in the real space, and the two-dimensional coordinates of the specific point in the second image captured of the real space; and an image coordinate acquisition unit that acquires focus position image coordinates, which are the two-dimensional coordinates of the focus position in the second image, by applying the focus position spatial coordinates to the transformation equation.
2. The information processing device according to claim 1, further comprising an extraction unit that extracts a specific image, which is an image of an area including the position of the image coordinates of the focus position, from the second image.
3. An information processing device as described in claim 2, further comprising an identification unit that acquires item identification information identified by reference based on the specific image of item information that associates item identification information that identifies an item with appearance information that represents the appearance of the item.
4. The information processing device of claim 1, wherein the parameter calculation unit includes: a setting unit that sets a virtual camera as a viewpoint position for capturing an image in which the virtual space model is projected into the virtual space; a feature point extraction unit that extracts feature points from at least one virtual space image captured of the virtual space by the virtual camera and from the second image; a feature point matching unit that matches feature points of the second image with feature points of the virtual space image; and a calculation unit that calculates the transformation parameters for capturing the second image by substituting reference space coordinates in the virtual space of a reference feature point that is the matched feature point and reference image coordinates that are the coordinates of the reference feature point in the second image into the transformation formula, wherein the reference space coordinates are calculated based on the transformation parameters for capturing the virtual space by the virtual camera and virtual reference image coordinates that are the coordinates of the reference feature point in the virtual space image.
5. The transformation parameters include six predetermined external camera parameters, four internal camera parameters, and five parameters related to the lens distortion of the camera, which are related to the projection of the three-dimensional space onto the two-dimensional image, and the calculation unit calculates the following transformation formula, which represents the relationship between three-dimensional coordinates (X, Y, Z) in the three-dimensional space and two-dimensional coordinates (u, v) in the two-dimensional image: In the above, by substituting the three-dimensional reference space coordinates into the three-dimensional coordinates (X, Y, Z) and the two-dimensional reference image coordinates into the two-dimensional coordinates (u, v), the camera external parameters R, t and the camera internal parameters f each having three degrees of freedom are obtained. x , f y , c x , c y , and a parameter k relating to the lens distortion of the camera 1 , k 2 , k 3 , p 1 , p 2 The information processing apparatus according to claim 4 , wherein:
6. The information processing device according to claim 4, wherein the calculation unit calculates the transformation parameters based on the virtual space image having a number of reference specific points equal to or greater than a threshold value related to the reference feature points.
7. The information processing device according to claim 1, wherein the focus position is a collision position between a part of a person depicted in the first image and an object.
8. The information processing device according to claim 1, wherein the focus position is a collision position between the line of sight of a person depicted in the first image and an object.
9. An information processing method executed by a processor, comprising: a spatial coordinate acquisition step of acquiring focus position spatial coordinates, which are three-dimensional coordinates of the focus position in virtual space, estimated by reflecting a focus position in real space shown in a first image, which is an image of real space captured by a first camera, onto a virtual space model that represents a virtual space corresponding to the real space; a second image acquisition step of acquiring a second image, which is an image of the real space captured by a second camera and includes the focus position; a parameter calculation step of calculating transformation parameters in a predetermined transformation equation that expresses, using predetermined transformation parameters, the relationship between the three-dimensional coordinates of a specific point, which is a specific point in the real space, and the two-dimensional coordinates of the specific point in the second image captured of the real space; and an image coordinate acquisition step of acquiring focus position image coordinates, which are two-dimensional coordinates of the focus position in the second image, by applying the focus position spatial coordinates to the transformation equation.
Citation Information
Patent Citations
Crime prevention support system
JP2005347905A
Camera parameter set calculation method, camera parameter set calculation program and camera parameter set calculation device
JP2018191275A
Image processing device, image processing system, and program
JP2021174136A