Sensing device, sensing method, and sensing system

The sensing device addresses the challenge of accurately estimating object positions in complex environments by generating conversion parameters to convert image data into simulated point cloud data, ensuring precise three-dimensional positioning even with occlusions or uneven surfaces.

JP2025080816APending Publication Date: 2025-05-27HITACHI LTD
View PDF 1 Cites 0 Cited by

Patent Information

Application Number
JP2023194101
Authority / Receiving Office
JP · JP
Patent Type
Applications
Current Assignee / Owner
Filing Date
2023-11-15
Publication Date
2025-05-27

AI Technical Summary

Technical Problem

Existing sensing technologies struggle to accurately estimate the three-dimensional position of objects in environments with occlusions or uneven surfaces, such as benches or steps, using images from environmental cameras.

Method used

A sensing device that acquires image data from a camera and point cloud data from a sensor, generates conversion parameters to convert image-based information into simulated point cloud data, allowing for accurate estimation of object positions in three-dimensional space.

Benefits of technology

The proposed solution enables accurate estimation of object positions in three-dimensional space even when parts of objects are occluded or when objects are on uneven surfaces, improving the reliability of sensing systems in complex environments.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure 2025080816000001_ABST
    Figure 2025080816000001_ABST
Patent Text Reader

Abstract

To provide a sensing device capable of generating parameters for highly precisely estimating a position of an object by an image acquired from a camera installed in an environment.SOLUTION: A sensing device comprises: an image data reception unit for acquiring image data from a camera; a point cloud data reception unit for acquiring point cloud data from a sensor; a first information generation unit for generating first information related to an object included in the image data by using the image data; a second information generation unit for generating second information related to an object included in the point cloud data by using the point cloud data; and a conversion parameter generation unit for, by using the first information and the second information, generating a conversion parameter for converting the first information into pseudo second information related to the object included in the image data. The pseudo second information includes information of the same kind as the second information.SELECTED DRAWING: Figure 2A
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention relates to a sensing device, a sensing method, and a sensing system for estimating the position of an object in an image captured by a camera installed in an environment based on the image.

Background Art

[0002] In recent years, there has been an increasing need for sensing technologies that detect, track, and predict the path of a target object by analyzing data obtained from sensors for infrastructure such as cameras and LiDAR (Laser Imaging Detection And Ranging) installed in the environment. In particular, in the field of autonomous driving, infrastructure sensors are installed in areas where people and vehicles are mixed or in areas that are blind spots from the host vehicle, and based on the output of the infrastructure sensors, the presence or absence of objects (other vehicles, people, etc.) on the host vehicle's path and the predicted trajectory of the objects are estimated, and it is expected to realize safe and secure movement that avoids accidents by using these estimation results for autonomous driving.

[0003] In object recognition technology using a camera, a method that uses a machine learning-based detector such as a neural network to estimate the type and pixel area of an object reflected in an image has shown high performance. To apply such object recognition technology to the field of autonomous driving, it is necessary to estimate the position of the object in the real-world three-dimensional coordinate system based on the estimated information of the type and pixel area of the object in the image.

[0004] As a method for estimating the position of an object in a three-dimensional coordinate system from information on the type and pixel area of an object in an image, there is a method of converting the lower end position of the pixel area of the object in the image into a position in the three-dimensional coordinate system by assuming that the road surface is a plane and using external parameters such as the height and attitude of the camera with respect to the road surface and camera parameters such as the camera focal length.

[0005] For example, in the abstract of Patent Document 1, it is described as a problem to "provide an object position estimation device capable of calculating the position of an object in a three-dimensional space." In claim 1 of the same document, "an image acquisition means (101) for acquiring an image (113) in one direction in a vehicle (111), an object recognition means (3) for recognizing an object (115) in the image, a posture change amount detection means (3) for detecting a change amount (ΔP) with respect to a reference posture in the posture of the vehicle, a first correction means (3) for correcting a position (y) on the vertical axis of the object in the image to a position (y') when the posture of the vehicle is the reference posture based on the change amount, a undulation information acquisition means (3, 7, 105) for acquiring the degree of undulation of a road surface (123) where the object exists with respect to a virtual flat surface (121) including a road surface (119) where the vehicle exists, a second correction means (3) for correcting the position on the vertical axis of the object in the image after correction by the first correction means to a position (y'') when the object exists on the virtual flat surface based on the degree of undulation, a map (9) for defining the relationship between the position on the vertical axis in the image when the posture of the vehicle is the reference posture and the position on the virtual flat surface, a flat surface position calculation means for calculating the position of the object on the virtual flat surface from the position on the vertical axis of the object in the image after correction by the second correction means using the map, and a three-dimensional space position calculation means (3) for calculating the position of the object in a three-dimensional space from the position of the object on the virtual flat surface and the degree of undulation, are provided, and an object position estimation device (1) is characterized by this."

[0006] Thus, in Patent Document 1 that uses an in-vehicle camera, the position on the vertical axis of an object in an image is corrected from the change amount of the vehicle's posture and the degree of road inclination, the position of the object on a virtual plane is calculated using a map, and the position of the object in a three-dimensional coordinate system of the road is estimated using the road inclination information.

Prior Art Documents

Patent Documents

[0007] [Patent Document 1] Japanese Patent Application Laid-Open No. 2015-215299 [Summary of the Invention] [Problems to be Solved by the Invention]

[0008] According to the technology of the above Patent Document 1, based on an image captured by an in-vehicle camera, the three-dimensional position of an object existing on an inclined road surface can be estimated. However, when this technology is applied to an environmental camera installed as an infra sensor, the following problems may be considered.

[0009] For example, when estimating the three-dimensional position of an object (person) in an area with many structures such as benches, like in the outdoor area of a commercial facility or a square, as shown in Fig. 17A, due to occlusion such as a structure being reflected in front of the object (person), only the upper part of the object (person) may be reflected in the image. In this case, since the object position is estimated on the assumption that the lower end of the recognized object frame (near the abdomen of the person) is on the ground, there is a risk of estimating the three-dimensional position of the object (person) farther than the actual position.

[0010] Also, as shown in Fig. 17B, in a scene where a person is standing on a step or the like, since the object position is estimated on the assumption that the lower end of the recognized object (person) is on the ground, there is a risk of estimating the three-dimensional position of the object (person) farther than the actual position.

[0011] Therefore, an object of the present invention is to provide a sensing device, a sensing method, and a sensing system that can generate conversion parameters for accurately estimating the three-dimensional position of an object based on an image even when a part of the object in the image captured by the environmental camera is occluded by another structure or when the object in the image exists on a step. [Means for Solving the Problems]

[0012] In order to solve the above problems, the sensing device of the present invention includes an image data receiving unit that acquires image data from a camera, a point cloud data receiving unit that acquires point cloud data from a sensor, a first information generating unit that generates first information related to an object included in the image data using the image data, a second information generating unit that generates second information related to an object included in the point cloud data using the point cloud data, and a conversion parameter generating unit that generates a conversion parameter for converting the first information into pseudo-second information related to the object included in the image data using the first information and the second information, wherein the pseudo-second information includes information of the same type as the second information. Other aspects of the present invention will be described in the embodiments described below.

Advantages of the Invention

[0013] According to the present invention, even when a part of an object in an image captured by an environmental camera is blocked by other structures or when the object in the image is on a step, conversion parameters for accurately estimating the three-dimensional position of the object based on the image can be generated.

Brief Description of the Drawings

[0014]

Figure 1

Figure 2A

Figure 2B

Figure 2C

Figure 3

Figure 4

Figure 5

Figure 6A

Figure 6B

Figure 6C

Figure 7

Figure 8

Figure 9

Figure 10

Figure 11

Figure 12

Figure 13

Figure 14

Figure 15

Figure 16

Figure 17A

Figure 17B

Embodiments for Carrying Out the Invention

[0015] Hereinafter, embodiments of the present invention will be described with reference to the drawings.

Examples

[0016] FIG. 1 is a configuration diagram of the sensing system according to Embodiment 1. This sensing system is a system that generates conversion parameters for converting the output (image data) of camera 200 and the output (point cloud data) of sensor 300 into simulated data (simulated second information I2') of first information I1 based on the image data to second information I2 based on the point cloud data by inputting them to the sensing device 100.

[0017] The camera 200 is, for example, an environmental camera such as a monocular camera, a stereo camera, or an infrared camera installed to view objects around the crosswalk obliquely, but is not limited thereto.

[0018] The sensor 300 is, for example, an environmental sensor such as a scanning LiDAR or a distance sensor installed to measure objects around the crosswalk, but is not limited thereto.

[0019] (Sensing Device 100) The sensing device 100 is a device that generates conversion parameters based on image data and point cloud data. As shown in FIG. 1, it includes a processing unit 101, a storage unit 102, an input unit 103, a display unit 104, and a communication unit 105. The processing unit 101 is a central processing unit (CPU) that executes various programs stored in a RAM, an HDD, etc. The storage unit 102 is a RAM, an HDD, etc., and stores various programs and various data (calibration parameter P0, conversion parameter P1, machine learning parameter P2) for the sensing device 100 to execute processing. The input unit 103 is a device for inputting instructions to a computer such as a keyboard or a mouse, and inputs instructions such as program startup. The display unit 104 is a display or the like that displays the execution status and execution results of the processing by the sensing device 100. The communication unit 105 is a device that exchanges various data and commands with other devices via a network. Although the processing unit 101 realizes each functional unit described later by executing various programs, the description of such well-known technologies will be appropriately omitted below.

[0020] Figure 2A is a functional block diagram of the sensing device 100. As shown here, the sensing device 100 of this embodiment includes an image data reception unit 1, a point cloud data reception unit 2, a first information generation unit 3, a second information generation unit 4, a conversion parameter generation unit 5, a conversion unit 6, a determination unit 7, and an update instruction unit 8. Note that the calibration parameter P0 in the figure is a parameter such as the installation and relative position and orientation of the camera 200 and the sensor 300, and the conversion parameter P1 is the final product of the sensing device 100 of this embodiment. Hereinafter, the outline of each functional unit shown in Figure 2A will be described.

[0021] The image data reception unit 1 is a functional unit that acquires image data from the camera 200, and the point cloud data reception unit 2 is a functional unit that acquires point cloud data from the sensor 300. Note that the image data and the point cloud data acquired by both reception units are a data pair with synchronized shooting timing and measurement timing. As a method for synchronizing the image data and the point cloud data, for example, a method of checking the times of both timings and synchronizing those measured and shot at the same time can be mentioned.

[0022] The first information generation unit 3 is a functional unit that generates first information I1 related to an object in the image based on the image data. The first information I1 includes, for example, any one or more of rectangular information surrounding the surrounding area of the object on the image, type information indicating the type of the object, color information indicating the color of the object, and position information indicating the position of the object in the three-dimensional space estimated from the rectangular information, but is not limited thereto. Note that the details of the first information generation unit 3 will be described later.

[0023] The second information generation unit 4 is a functional unit that generates second information I2 related to an object in the point cloud based on the point cloud data. The second information I2 includes, for example, any one or more of object point cloud information that is a point cloud corresponding to the object region, position information indicating the position of the object, size information, type information, shape information representing the shape, reliability information, and reflectance information of the point cloud data of the object part, but is not limited thereto. Note that the details of the second information generation unit 4 will be described later.

[0024] The conversion parameter generation unit 5 is a functional unit that generates the conversion parameter P1. The details of the conversion parameter generation unit 5 will be described later.

[0025] The conversion unit 6 is a functional unit that converts the first information I1 into the simulated second information I2' using the conversion parameter P1 generated by the conversion parameter generation unit 5. The conversion process here is a process of generating one or more pieces of information included in the original second information I2 using one or more pieces of information included in the first information I1.

[0026] The determination unit 7 is a functional unit that compares a certain type of information included in the second information I2 generated by the second information generation unit 4 with the same type of information included in the simulated second information I2' converted by the conversion unit 6, and determines whether the difference d between the two is within the threshold value. The threshold value used here is appropriately set according to the information type to be determined.

[0027] The update instruction unit 8 is a functional unit that instructs each unit to regenerate the conversion parameter P1 when the difference d is greater than the threshold value.

[0028] (Generation process of the conversion parameter P1) Hereinafter, the outline of the generation process of the conversion parameter P1 by the sensing device 100 will be described with reference to FIG. 3.

[0029] In step S1, the image data reception unit 1 and the point cloud data reception unit 2 acquire synchronized image data and point cloud data from the camera 200 and the sensor 300.

[0030] In step S2, the first information generation unit 3 generates the first information I1 based on the image data.

[0031] In step S3, the second information generation unit 4 generates the second information I2 based on the point cloud data. Therefore, by steps S2 and S3, synchronized first information and second information are generated.

[0032] In step S4, the sensing device 100 determines whether sufficient data (for example, 100 data) has been acquired. If the amount of data is insufficient, it returns to steps S1 to S3 to replenish the first information I1 and the second information I2. On the other hand, if the amount of data is sufficient, it proceeds to step S5.

[0033] In step S5, the conversion parameter generation unit 5 generates a conversion parameter P1 for converting the first information I1 into the simulated second information I2'. Then, in order to confirm the quality of the generated conversion parameter P1, the following test processes (steps S6 to S11) are performed.

[0034] In step S6, the image data reception unit 1 and the point cloud data reception unit 2 acquire a plurality of synchronized image and point cloud data pairs as a test data group.

[0035] In step S7, the first information generation unit 3 generates the first information I1t based on the image data of the test data group.

[0036] In step S8, the conversion unit 6 converts the first information I1t into the simulated second information I2t' using the conversion parameter P1.

[0037] Also, in step S9 that runs in parallel with steps S7 and S8, the second information generation unit 4 generates the second information I2t based on the point cloud data of the test data group.

[0038] In step S10, the determination unit 7 compares the same type of information included in the simulated second information I2t' generated in step S8 and the second information I2t generated in step S9, and calculates the difference d between the two. The same type of information means, for example, the position information of the same object at the same timing.

[0039] In step S11, the determination unit 7 determines whether the difference d is within the threshold value. If the difference d is within the threshold value, the generation process of the conversion parameter P1 is terminated on the assumption that the quality of the current conversion parameter P1 is guaranteed. On the other hand, if the difference d is not within the threshold value, the process proceeds to step S12 in order to improve the quality of the conversion parameter P1.

[0040] In step S12, the update instruction unit 8 sends instructions necessary for updating the conversion parameter P1 to each unit. Specifically, the image data reception unit 1 and the point cloud data reception unit 2 are instructed to replenish data, and the first information generation unit 3, the second information generation unit 4, and the conversion parameter generation unit 5 are instructed to perform processes on the replenished data as well (steps S1 to S5).

[0041] For example, if the number of acquired data used for generating the current conversion parameter P1 is 100 data, an instruction is given to replenish 50 data and regenerate the conversion parameter P1 based on a total of 150 data. Then, the tests from step S6 to step S11 are executed again for the regenerated conversion parameter P1. By repeating these processes until the difference d is within the threshold value, it is possible to generate the conversion parameter P1 for converting the first information I1 into the high-quality simulated second information I2'. Also, if the difference d becomes large only at a specific position in the image data, the update instruction unit 8 may instruct to replenish the data in the vicinity of the specific position preferentially and update the conversion parameter P1.

[0042] Note that the order of step S2 and step S3 is arbitrary and they may be executed in parallel. Also, in step S4, as a method for determining whether data has been sufficiently acquired, the number of data in images or point clouds may be used, or the number of first information I1 and second information I2 of objects included in the image data or point cloud data may be used, or the result calculated using these numbers may be used. Further, in FIG. 3, although a form in which steps S7, S8, and step S9 are executed in parallel is shown, they may be executed in order on a single thread. Also, as a completion determination when acquiring a plurality of test data in step S6, it may be determined based on whether the number of acquired data or the acquisition time exceeds a threshold value, or by sequentially executing steps S7 to S9 for the test data group, it may be determined whether the number of acquired second information I2 is equal to or greater than the threshold value, or it may be determined from the variation or comprehensiveness of the values included in the second information I2.

[0043] Through the above processing, it can be guaranteed that the generated conversion parameter P1 can convert the first information I1 into the simulated second information I2' with sufficient performance.

[0044] Next, the details of each functional unit of the sensing device 100 of the present embodiment, which are prepared to implement each process of FIG. 3, will be described.

[0045] (First Information Generation Unit 3) As shown in FIG. 2B, the first information generation unit 3 includes a camera parameter acquisition unit 31, an object detection unit 32, and a position estimation unit 33, and by using these, first information I1 related to an object included in the image data is generated. Hereinafter, the details of each functional unit will be sequentially described.

[0046] <Camera Parameter Acquisition Unit 31> The camera parameter acquisition unit 31 is a functional unit that acquires camera parameters indicating optical characteristics such as the focal length and optical axis of the camera 200.

[0047] <Object Detection Unit 32> The object detection unit 32 is one of the functional units that generates the first information I1, and estimates the rectangular information and type information of the moving object included in the image data.

[0048] FIG. 4 shows an example in which the object detection unit 32 infers the rectangular information and type information of the moving object in the image data as the first information I1. When the illustrated image data is input to the object detection unit 32, the inference unit 32a of the object detection unit 32 infers the rectangular information and type (class) information using the machine learning parameter P2, which is a parameter such as the weight of the machine learning-based detector. At the time of this inference, a parameter regarding the plausibility of the estimation is generally also inferred, and only the rectangular information and type of the object with a higher value than the threshold can be finally used as the first information I1. Thereby, for example, false inferences in a scene with low reliability such as backlight or nighttime can be suppressed.

[0049] In the example of FIG. 4, the inference unit 32a is a detector that infers the first information I1 consisting of the rectangular information of the six moving objects shown in the image and their respective type information (for example, Person, Car, Truck). However, not limited to this example, for example, an object other than the moving object may be targeted, or a task of predicting the label of which object each pixel of the image is, a task of predicting the label for all objects in the image and assigning a unique ID, a task of predicting the class label for all pixels in the image and estimating a unique ID, etc., may be targeted for segmentation. Also, in this embodiment, the distortion of the camera image is corrected by the image data reception unit 1, but if it is not corrected, it may be corrected by the first information generation unit 3.

[0050] In FIG. 4, when only a part of the moving object is shown instead of the whole body in the input image data, the inference unit 32a may infer only that part as the rectangular information. For example, in the image of FIG. 4, there is a bench in front of the worker carrying a load at the lower left, and the lower body is shielded and only the upper body is shown in the image. Therefore, only the upper body part may be inferred as the person area with rectangular information, and inappropriate processing may occur in the subsequent position estimation unit 33.

[0051] <Position estimation unit 33> The position estimation unit 33 is one of the functional units that generate the first information I1, and is a functional unit that estimates the position information of the moving body detected by the object detection unit 32.

[0052] FIG. 5 shows an example in which the position estimation unit 33 estimates the position of the moving body in the image data as the first information I1. The position estimation unit 33 acquires the camera parameters acquired by the camera parameter acquisition unit 31, the object region inferred by the object detection unit 32, and the installation information of the camera 200 and the sensor 300 and the parameters of the relative position and orientation based on the calibration parameter P0, and estimates the position of the moving body in the following procedure.

[0053] First, with the point at the center of the lower end of the object region in the image data as the object position (p u , p v ) in the image coordinate system, the object position (p u , p v ) is converted into the object position (p x , p y ) in the normalized image coordinate system by Equation 1.

[0054]

Equation

[0055] Here, c u , c v , f u , f v , and δ are the horizontal and vertical image centers of the camera 200, the horizontal and vertical focal lengths, and the physical interval between pixels, respectively. Usually, c u , c v , f u , f v are known by camera calibration and are acquired from the camera parameter acquisition unit 31. On the other hand, δ is generally an unknown parameter, but even if it is unknown, δ can be estimated by the method described later.

[0056] Subsequently, a coordinate system (world coordinate system (X where the X and Y axes are parallel to the groundW , Y W , Z W )) is virtually defined, and from Equation 2, the object position (p x , p y ) in the normalized image coordinate system is converted to the object position (p x w , p y w , p z w ) in the world coordinate system.

[0057]

Equation

[0058] Here, R in Equation 2 is a 3×3 rotation matrix for converting from the camera coordinate system (X C , Y C , Z C ) to the world coordinate system. Also, t in Equation 2 is a 3×1 translation matrix, which is obtained from the calibration parameter P0 as the installation information of the camera with respect to the ground.

[0059] In the above Equations 1 and 2, the coordinates (p u , p v ) at the center of the lower end of the rectangle recognized by the object detection unit 32 are assumed to be at the position where p z w = 0 in the world coordinate system. By calculating δ from Equations 1 and 2 using this constraint, the object position in the world coordinate system can be obtained.

[0060] With the above first information generation unit 3, as the first information I1 related to the object included in the image, the rectangle information, type information, and position information of the moving body can be estimated. Note that the position information as the first information I1 estimated by the position estimation unit 33 may include a relatively large position error due to the mechanism described in FIGS. 17A and 17B. In view of this problem, in this embodiment, a conversion parameter P1 for estimating high-precision position information based on the first information I1 as the simulated second information I2' is generated.

[0061] (Second Information Generation Unit 4) Next, a method for generating second information I2 related to an object included in point cloud data will be described using the flowcharts of FIGS. 6A to 6C.

[0062] <Method for generating second information I2 according to the flowchart of FIG. 6A> First, a point cloud recognition method called Pillar-based, which is used when there are multiple sensors 300, will be described using the flowchart of FIG. 6A.

[0063] In step S21, the second information generation unit 4 integrates the point cloud data of the multiple sensors.

[0064] In step S22, the second information generation unit 4 removes the background point cloud from the integrated point cloud data by taking the difference between the integrated point cloud data and a previously prepared background point cloud, and extracts the point cloud data of the moving object.

[0065] In step S23, the second information generation unit 4 divides the three-dimensional region of the point cloud data of the moving object into a plurality of grids called Pillars, and then extracts features such as the number and variance of the point cloud data for each Pillar, and generates a 2D map overlooking this grid.

[0066] In step S24, the second information generation unit 4 determines the presence or absence of the moving object for each grid grid of the map.

[0067] In step S25, the second information generation unit 4 groups the moving object grids.

[0068] In step S26, the second information generation unit 4 estimates the object point cloud information, position information, size information, type information, shape information, and reliability. The reliability can be calculated with a lower reliability for, for example, black objects or mirror objects that sensors such as LiDAR are not good at, i.e., those with low or high reflected intensities of the point cloud data, or it can be calculated to lower the reliability when the size, shape, or position of the object deviates from a specified pattern or threshold.

[0069] By the above-described point cloud recognition method called the Pillar base, the second information generation unit 4 can generate various types of second information I2 described above.

[0070] <Method for generating second information I2 according to the flowchart of FIG. 6B> Next, a point cloud recognition method called the voxel base, which is used when there are a plurality of sensors 300, will be described with reference to the flowchart of FIG. 6B. Note that steps S21 and S22 are the same as those in FIG. 6A, so redundant explanations will be omitted.

[0071] In step S27 of FIG. 6B, the second information generation unit 4 converts the point cloud data in the space into voxels composed of cubes.

[0072] In step S28, the second information generation unit 4 determines the presence or absence of an object for each voxel.

[0073] In step S29, the second information generation unit 4 groups the object voxels.

[0074] In step S2a, the second information generation unit 4 estimates object point cloud information, position information, size information, type information, and shape information.

[0075] Also by the above-described point cloud recognition method called the voxel base, the second information generation unit 4 can generate various types of second information I2 described above.

[0076] <Method for generating second information I2 according to the flowchart of FIG. 6C> Next, a point cloud recognition method called the machine learning base, which is used when there are a plurality of sensors 300, will be described with reference to the flowchart of FIG. 6C. Note that steps S21 and S22 are the same as those in FIG. 6A, so redundant explanations will be omitted.

[0077] In step S2b of FIG. 6C, the second information generation unit 4 directly estimates object point cloud information, position information, size information, type information, and shape information by inputting the point cloud data into a detector such as a neural network.

[0078] Even by the point cloud recognition method called the above machine learning-based method, the second information generation unit 4 can generate various pieces of second information I2 described above.

[0079] Note that in step S22 of each figure, the point cloud data of the moving object is extracted by removing the background point cloud from the moving object, and the accuracy of the subsequent processing is improved. However, for example, the stationary object may be recognized without removing the background difference.

[0080] By the above second information generation unit 4, the second information I2 including the object point cloud information, the position information, the size information, the type information, and the shape information can be estimated.

[0081] (Conversion parameter generation unit 5) As shown in FIG. 2C, the conversion parameter generation unit 5 includes a projection unit 51, a comparison unit 52, a learning unit 53, and a map generation diagram 54. By using these, a conversion parameter P1 for converting the first information I1 into the simulated second information I2' is generated. Hereinafter, the details of each functional unit will be sequentially described.

[0082] <Projection unit 51> The projection unit 51 is a functional unit that projects the information of the object included in the first information I1 and the information of the object included in the second information I2 into the same coordinate system by using the calibration parameter P0 and the camera parameter.

[0083] FIG. 7 is a diagram showing an example of projecting the object point cloud information included in the second information I2 onto the image data in order to associate it with the rectangular information which is the first information I1. Using this figure, a method of associating the object point cloud estimated by the second information generation unit 4 with the pixels of the image will be described.

[0084] First, using Equation 3, the point cloud sP in the sensor coordinate system is converted into the point cloud cP in the camera coordinate system by using the relative position and orientation parameter cTs of the camera 200 and the sensor 300 included in the calibration parameter P0.

[0085]

Equation

[0086] Note that the point cloud cP in the camera coordinate system and the point cloud sP in the sensor coordinate system represented by Equation 4 are usually represented by three-dimensional points (x, y, z). However, for dimensional consistency in calculations, they are described in a format where 1 is added in the dimensional direction, such as (x, y, z, 1).

[0087]

Number

[0088] Also, the position and orientation parameter cTs shown in Equation 5 is a position and orientation parameter representing the conversion from the camera coordinate system to the sensor coordinate system, and is a 4×4 matrix composed of the 3×3 rotation matrix cRs representing the rotation shown in Equation 6 and the translation matrix cTs showing the translation shown in Equation 7.

[0089]

Number

[0090]

Number

[0091]

Number

[0092] Through the above equations, the point cloud sP in the sensor coordinate system is converted into the point cloud cP in the camera coordinate system.

[0093] Subsequently, using Equation 8, the point cloud cP in the camera coordinate system is converted into pixels (u, v) on the image using camera parameters. Note that in this embodiment, various processes are performed on the image with image distortion, but it is not limited to this.

[0094]

Number

[0095] Here, in Equation 8, fx and fy are the focal length parameters of camera 200, and cx and cy are the principal points of the image center, that is, the parameters of the optical center of the lens. The above camera parameters can be obtained in advance by a method called camera calibration. Note that for Xc, Yc, and Zc in Equation 8, the values obtained in Equation 3 are used. By solving the above Equation 8, the point cloud cP in the camera coordinate system can be converted into pixels (u, v) on the image using the camera parameters as shown in Equation 9 and Equation 10.

[0096]

Equation

[0097]

Equation

[0098] Note that in this embodiment, the point cloud is superimposed on the image. However, it is not limited to this example. For example, by projecting the object position information in the first information I1 and the object position information in the second information I2 onto a three-dimensional coordinate system such as a sensor coordinate system, they may be superimposed on the same coordinate system.

[0099] <Comparison unit 52> The comparison unit 52 is a functional unit that associates the same objects by comparing the information of a plurality of objects included in the first information I1 and the information of a plurality of objects included in the second information I2. In this comparison unit 52, when the object point cloud information for each object included in the second information I2 is projected onto the image by the projection unit 51, it is determined whether the projected point cloud exists inside or around any of the rectangular information in the first information I1, thereby determining whether they are the same object.

[0100] Alternatively, the comparison unit 52 may determine whether the objects are the same based on the difference in distance when the object position information in the first information I1 and the object position information in the second information I2 are projected onto a three-dimensional coordinate system such as a sensor coordinate system by the projection unit 51. Further, the object type information based on the first information I1 may be compared with the size information, object type information, and shape information of the object based on the second information I2. Additionally, methods using past information, methods using the tracking information of each moving object estimated by a Kalman filter or the like, or methods combining these may be used, and the present invention is not limited to these examples.

[0101] <Learning unit 53> The learning unit 53 is a functional unit that learns the parameters of machine learning for converting the first information I1 into the simulated second information I2'. For example, learning may be performed using a neural network or the like that inputs the rectangular information and type information included in the first information I1 and uses the position information included in the second information I2 as a teacher so that the position information can be inferred from the rectangular information and type information, or an image may be added to the input information.

[0102] <Map generation unit 54> The map generation unit 54 is a functional unit that generates a map for converting the first information I1 into the simulated second information I2'. Hereinafter, the position information included in the first information I1 is treated as having been converted into the position information in the sensor coordinate system using the calibration parameter P0.

[0103] FIG. 8 is a diagram illustrating a previous stage of a map generation procedure for converting the position information included in the first information I1 into the position information included in the simulated second information I2'. In the present embodiment, the map is obtained by dividing a plane parallel to the ground assumed by the position estimation unit 33 into a grid shape based on the sensor coordinate system and assigning a correction vector for correcting the position in the first information I1 to the position in the second information I2 to each grid. That is, in a scene as shown in FIG. 17A, a map is generated that converts a position with a large error generated based on the image data of the camera 200 to a position with a small error generated based on the point cloud data of the sensor 300 with accurate distance information.

[0104] FIG. 8 shows the trajectories of the position information Fc based on the first information I1 and the position information Fs based on the second information I2 for the objects determined to be the same by the comparison unit 52. The map generation unit 54 obtains a correction vector based on the difference vector of the position information at the same time and updates the value in the map. As an example, during the period of frames F(0) to F(x), since the error between the position information Fc and the position information Fs is smaller than the threshold value, the map is not updated. On the other hand, in frame F(x + 1), since the error between the position information Fc(x + 1) and the position information Fs(x + 1) is larger than the threshold value, a correction vector as shown in FIG. 9 is calculated based on the difference vector in FIG. 8, and the map is updated. Note that the threshold value used here may be a fixed value. Generally, since the position error increases as the distance from the camera increases, it is desirable to reduce the threshold value for objects or grids close to the camera and increase the threshold value for objects or grids far from the camera. The reason for the large position error in frame (x + 1) may be the situations exemplified in FIGS. 17A and 17B.

[0105] Next, the method for updating the correction vector and the map will be described. In the present embodiment, a correction map is created for each object type included in the first information I1. First, based on the position information (the position converted into the sensor coordinate system using the calibration parameter P0) based on the first information I1, the grid to be updated is selected. In the case of FIG. 8, in frame F(x + 1), the grid in the lower right region is the target. Subsequently, a correction vector is obtained. Although it is conceivable to use, for example, the value of the error vector in FIG. 8 as the correction vector as it is, the position information at a single time may include many instantaneous errors. Therefore, by using an algorithm such as taking the weighted sum of the error vector and the past correction vector, a correction vector with a small error is obtained, and this value is registered in the map as the correction vector.

[0106] In this embodiment, the position information included in the first information I1 is treated as being converted into the position information in the sensor coordinate system using the calibration parameter P0. However, it is only necessary to handle the positions in the same coordinate system, and this is not limited to this example. Also, the threshold for the error between the position information Fc and the position information Fs may be determined as a single value, or may be determined to change according to the position on the grid. Further, although a correction map is created for each object type included in the first information I1, it is also conceivable to improve the performance more by generating a correction map in consideration of not only the object type but also the size information and shape information of the object, and this is not limited to these examples. Also, as a method for obtaining the correction vector, in addition to taking the weighted sum of the difference vector and the past correction vector, it may be obtained by a method using machine learning or the like.

[0107] Also, for an object whose reliability is calculated to be low in the second information I2, even when it is associated with the first information I1, it can be excluded from the data used for generating the conversion parameter P1. For example, an object detected as a black or mirror object may have low reliability of the second information I2 by the sensor 300. Therefore, when the object included in the first information I1 is black or mirror, or for data calculated to have low reliability by the second information generation unit 4, by excluding it from the data used by the conversion parameter generation unit 5, the error of the conversion parameter P1 due to the second information I2 including errors can be reduced.

[0108] Also, as another scene that the sensor 300 is not good at, when rain, snow, fog, etc. occur, the error of the information included in the second information I2 may increase. Therefore, for example, the weather information at the installation location of the sensing device 100 is acquired from the Internet, and using that information, the reliability of the second information I2 is changed. When the reliability is low, by excluding it from the data used by the conversion parameter generation unit 5, the error of the conversion parameter P1 due to the second information I2 including errors can be reduced.

[0109] The above conversion parameter generation unit 5 can generate a conversion parameter P1 for converting the first information I1 into a high-quality simulated second information I2'. Also, by constructing a neural network with the rectangular information and type information included in the first information I1 as inputs and the position information included in the second information I2 as the teacher in the learning unit 53, it is possible to convert to different types of information, such as inferring the position information from the rectangular information and type information by the conversion unit 6. Further, the map generation unit 54 can generate a map for converting the first information I1 into the simulated second information I2'.

[0110] (Conversion unit 6) The conversion unit 6 is a functional unit that converts the first information I1 into the simulated second information I2' using the conversion parameter P1. When the parameters of the machine learning generated by the learning unit 53 are used as the conversion parameter P1, the first information I1 is input to the machine learning model to which these machine learning parameters are applied, and the simulated second information I2' is generated by inference. Also, when the map created by the map generation unit 54 is used as the conversion parameter P1, after associating the position information of the object included in the first information I1 with the grid of the map, the position information is corrected by applying the correction vector in the corresponding grid, thereby performing conversion to the position information of the object included in the simulated second information I2'.

[0111] By the above conversion unit 6, it is possible to generate information with improved accuracy, such as converting a position (first information I1) with a large error generated based on the image of the camera 200 to a position (simulated second information I2') that is approximately equal to a position with a small error generated based on the output of the sensor 300 with accurate distance information. Also, in this embodiment, an example of converting the object position information of the first information I1 to the object position information of the simulated second information I2' is shown, but the breakdown of the information to be converted is not limited to this.

[0112] (Determination unit 7) The determination unit 7 is a functional unit that compares the simulated second information I2' converted by the conversion unit 6 with the second information I2 generated by the second information generation unit 4 and determines whether the difference d is within the threshold value.

[0113] As a method for the determination unit 7 to compare the position information included in the simulated second information I2' converted from the first information I1 by the conversion unit 6 with the position information included in the second information I2 generated by the second information generation unit 4, similar to the map generation unit 54, there is a method of calculating the difference vector of positions on the map and determining whether the difference amount exceeds a threshold value.

[0114] FIG. 10 is a diagram showing an example in which the calculated difference amounts are represented by shades for each grid of the map by the determination unit 7. Dark colors indicate grids with large difference amounts, and light colors indicate grids with small difference amounts. As in the example of FIG. 8, for the colored grids determined to have a large difference amount of positions, the determination unit 7 determines to update the position correction vector.

[0115] With the above determination unit 7, it becomes possible to test the performance of the conversion parameter P1.

[0116] (Update instruction unit 8) When the difference d is larger than the threshold value, the update instruction unit 8 collects data from the image data reception unit 1 and the point cloud data reception unit 2 so as to regenerate the conversion parameter P1, and uses the first information I1 and the second information added by the first information generation unit 3 and the second information generation unit 4 to update the conversion parameter generation unit 5.

[0117] As shown in FIG. 10, when there are grids determined to have a large difference amount of positions, the update instruction unit 8 issues an instruction to execute a series of correction vector generation processes again in order to update the position correction vector. At this time, the correction vectors of all grids may be updated, or only the grids determined to have a large difference amount of positions may have their correction vectors updated. Also, if the values of past correction vectors are stored and the values of the correction vectors do not converge and have a large variation, the update instruction unit 8 may exclude the grid with low reliability of the correction vector from the update target, and then notify the control server or the vehicle that the position information as a result of correction by the correction vector of that grid has low reliability.

[0118] By repeating the correction vector generation process by the above update instruction unit 8 until the performance becomes sufficient in design, the quality of the conversion parameter P1 can be ensured. Further, by calculating the reliability of the correction parameter, the reliability of the generated simulated second information I2' is evaluated and notified to the control server or the vehicle, so that driving can be performed in consideration of the reliability of the sensor information during automatic driving.

[0119] In addition, by periodically executing the above determination unit 7 and update instruction unit 8, it is possible to cope with changes in external structures and aging changes.

[0120] With the above sensing device 100, it is possible to generate a high-quality conversion parameter P1 for accurately estimating the position of an object from the image acquired from the environmental camera.

Embodiment

[0121] Next, Embodiment 2 of the present invention will be described with reference to FIGS. 11 and 12. Note that duplicate descriptions of the common points with Embodiment 1 are omitted.

[0122] FIG. 11 is a functional block diagram of the sensing device 100 of this embodiment. As shown here, the sensing device 100 of this embodiment includes, in addition to the above-described image data reception unit 1, point cloud data reception unit 2, first information generation unit 3, second information generation unit 4, conversion unit 6, and conversion parameter P1, an information integration unit 9.

[0123] The information integration unit 9 is a functional unit that generates integrated information by integrating the first information I1 generated by the first information generation unit, the second information I2 generated by the second information generation unit, and the simulated second information I2' generated by the conversion unit 6.

[0124] Next, the details of the processing by the sensing device 100 of this embodiment will be described using the flowchart of FIG. 12.

[0125] In step S31, the image data reception unit 1 acquires image data from the camera 200.

[0126] In step S32, the first information generation unit 3 generates first information I1 based on the image data.

[0127] In step S33, the conversion unit 6 generates simulated second information I2' using the first information I1 generated in step S32 and the conversion parameter P1 generated in the first embodiment.

[0128] In step S34, the point cloud data reception unit 2 acquires point cloud data synchronized with the image data acquired in step S31 from the sensor 300.

[0129] In step S35, the second information generation unit 4 generates second information I2 based on the point cloud data.

[0130] In step S36, the information integration unit 9 integrates the first information I1, the second information I2, and the simulated second information I2' generated in the above steps. As a result, integrated information with higher accuracy than each individual piece of information can be obtained. At this time, since the quality (similarity to the second information I2) of the simulated second information I2' is guaranteed, the second information I2 and the simulated second information I2' can be easily associated with each other.

[0131] In step S37, the information integration unit 9 transmits the integrated information to an external system (such as an in-vehicle automatic driving system or a control system).

[0132] The external system that receives the integrated information from the sensing device 100 of the present embodiment described above can realize appropriate vehicle control and the like based on the high-precision object position information.

Embodiment

[0133] Next, Embodiment 3 of the present invention will be described with reference to FIGS. 13 to 15. Note that the common points with the above-described embodiments will not be described repeatedly.

[0134] As described above, when the simulated second information I2' is generated using the conversion parameter P1 generated in Example 1, the quality (similarity to the second information I2) of the simulated second information I2' is sufficiently ensured. Therefore, when the system is in operation, the use of the second information I2 based on the point cloud data can be omitted. Therefore, the sensing system of this embodiment shown in FIG. 13 omits the sensor 300 necessary for generating the second information I2, and is configured by the sensing device 100 and the camera 200.

[0135] FIG. 14 is a functional block diagram of the sensing device 100 of this embodiment. As shown here, the sensing device 100 of this embodiment includes the above-described image data reception unit 1, first information generation unit 3, conversion unit 6, conversion parameter P1, and information integration unit 9.

[0136] Next, the details of the processing by the sensing device 100 of this embodiment will be described using the flowchart of FIG. 15.

[0137] In step S41, the image data reception unit 1 acquires image data from the camera 200.

[0138] In step S42, the first information generation unit 3 generates the first information I1 based on the image data.

[0139] In step S43, the conversion unit 6 generates the simulated second information I2' using the first information I1 generated in step S42 and the conversion parameter P1 generated in Example 1.

[0140] In step S44, the information integration unit 9 integrates the first information I1 and the simulated second information I2' generated in the above steps. Thereby, integrated information with higher accuracy than the single first information I1 can be obtained.

[0141] In step S45, the information integration unit 9 transmits the integrated information to an external system (an in-vehicle automatic driving system or a control system).

[0142] According to the present embodiment described above, even in a sensing system not provided with the sensor 300, vehicle control equivalent to that of Example 2 can be realized.

Example

[0143] Next, Example 4 of the present invention will be described with reference to FIG. 16. Note that duplicate descriptions of the common points with the above-described examples will be omitted. Although the sensor 300 is illustrated in FIG. 16, a system configuration in which the sensor 300 is omitted as in Example 3 may also be used.

[0144] In the sensing systems of Examples 1 to 3, the sensing device 100 was arranged in an environment where the camera 200 and the sensor 300 were installed. In contrast, in the present example, an edge device 400 is arranged in an environment where the camera 200 and the sensor 300 are installed, and a server connected via a communication network such as the Internet to this edge device 400 is used as the sensing device 100.

[0145] According to the system configuration of the present example, in the installation environment of the camera 200 or the like, it is only necessary to install a simple edge device 400 having only a function of transmitting data collected by the camera 200 or the like to the communication network, and the sensing device 100 that executes advanced arithmetic processing may be prepared as a server. Therefore, a sensing system for collectively managing a plurality of environments can be configured at low cost.

Explanation of Signs

[0146] 100 Sensing device 101 Processing unit 102 Storage unit 103 Input unit 104 Display unit 105 Communication unit 1 Image data reception unit 2 Point cloud data reception unit 3 First information generation unit 31 Camera parameter acquisition unit 32 Object detection unit 32a Inference unit 33 Position estimation unit 4 Second Information Generation Unit 5 Conversion Parameter Generation Unit 51 Projection Unit 52 Comparison Unit 53 Learning Unit 54 Map Generation Unit 6 Conversion Unit 7 Judgment Unit 8 Update Instruction Unit 9 Information Integration Unit 200 Camera 300 Sensor 400 Edge Device P0 Calibration Parameter P1 Conversion Parameter P2 Machine Learning Parameter

Claims

1. An image data reception unit that acquires image data from a camera, A point cloud data reception unit that acquires point cloud data from a sensor, A first information generation unit that generates first information related to an object included in the image data using the image data, A second information generation unit that generates second information related to an object included in the point cloud data using the point cloud data, A conversion parameter generation unit that generates a conversion parameter for converting the first information into pseudo-second information related to the object included in the image data using the first information and the second information, Comprising, The sensing device, wherein the pseudo-second information includes information of the same type as the second information.

2. In the sensing device according to claim 1, further A conversion unit that converts the first information into the pseudo-second information using the conversion parameter, A determination unit that determines whether the difference between the second information generated by the second information generation unit and the information of the same type of the pseudo-second information converted by the conversion unit is within a threshold value, An update instruction unit that instructs an update of the conversion parameter, Comprising, When the difference exceeds the threshold value, the update instruction unit Instructs the image data reception unit, the point cloud data reception unit, the first information generation unit, and the second information generation unit to add data, and Instructs the conversion parameter generation unit to update the conversion parameter using the added data. The sensing device is characterized by this.

3. In the sensing device according to claim 2, The threshold value is small for an object close to the camera and large for an object far from the camera. The sensing device is characterized by this.

4. In the sensing device according to claim 2, The second information includes a reliability regarding the information included in the second information, The conversion parameter generation unit updates the conversion parameter using the second information for which the reliability is higher than a predetermined threshold value and the corresponding first information. The sensing device is characterized by this.

5. In the sensing device according to claim 2, The determination unit determines whether the difference is within the threshold value for each small region obtained by dividing the space into a plurality of regions. The sensing device is characterized by this.

6. In the sensing device according to claim 1, The conversion parameter generation unit inputs the first information and performs machine learning with the second information as a teacher to generate the conversion parameter. A sensing device characterized by this.

7. In the sensing device according to claim 1, The conversion parameter generation unit generates, as the conversion parameter, a map that corrects the position of an object based on the first information to the position of the object based on the second information. A sensing device characterized by this.

8. In the sensing device according to claim 7, The conversion parameter generation unit has a comparison unit that compares the information of the object included in the first information with the information of the object included in the second information to link the same object, and generates, as the conversion parameter, a map that corrects the position of the object based on the first information to the position of the object based on the second information only for the object determined to be the same by the comparison unit. A sensing device characterized by this.

9. A sensing device that uses the conversion parameter generated by the sensing device according to any one of claims 1 to 7, an image data reception unit that acquires image data from a camera, a point cloud data reception unit that acquires point cloud data from a sensor, a first information generation unit that generates first information related to an object included in the image data using the image data, a second information generation unit that generates second information related to an object included in the point cloud data using the point cloud data, a conversion unit that converts the first information into the pseudo-second information using the conversion parameter, and an information integration unit that integrates the first information, the second information, and the pseudo-second information. A sensing device characterized by including these.

10. A sensing device that uses the conversion parameter generated by the sensing device according to any one of claims 1 to 7, an image data reception unit that acquires image data from a camera, a first information generation unit that generates first information related to an object included in the image data using the image data, a conversion unit that converts the first information into the pseudo-second information using the conversion parameter, and an information integration unit that integrates the first information and the pseudo-second information. A sensing device characterized by including these.

11. an image data reception step of acquiring image data from a camera, a point cloud data reception step of acquiring point cloud data from a sensor, A first information generation step of generating first information related to an object included in the image data using the image data; A second information generation step of generating second information related to an object included in the point cloud data using the point cloud data; A conversion parameter generation step of generating a conversion parameter for converting the first information into pseudo-second information related to an object included in the image data using the first information and the second information; A conversion unit step of converting the first information into the pseudo-second information using the conversion parameter; An information integration step of integrating the first information, the second information, and the pseudo-second information; comprising; The pseudo-second information includes information of the same type as the second information, and a sensing method characterized by this.

12. In the sensing method according to claim 11, further comprising a determination step of determining whether the difference between the second information and the information of the same type of the pseudo-second information is within a threshold value, when the difference exceeds the threshold value, adding data in the image data reception step, the point cloud data reception step, the first information generation step, and the second information generation step, and updating the conversion parameter using the added data in the conversion parameter generation step, and a sensing method characterized by this.

13. A camera installed in the environment, A sensor installed in the environment, A sensing device for detecting an object in the environment based on the image data from the camera and the point cloud data from the sensor, A sensing system comprising: The sensing device is an image data reception unit that acquires image data from the camera, a point cloud data reception unit that acquires point cloud data from the sensor, a first information generation unit that generates first information related to an object included in the image data using the image data, a second information generation unit that generates second information related to an object included in the point cloud data using the point cloud data, a conversion parameter generation unit that generates a conversion parameter for converting the first information into pseudo-second information related to an object included in the image data using the first information and the second information, comprising; The pseudo-second information includes information of the same type as the second information, and a sensing system characterized by this.

14. In the sensing system according to claim 13, the sensing device further has a conversion unit that converts the first information into the pseudo-second information using the conversion parameter, A determination unit that determines whether the difference between the second information generated by the second information generation unit and the information of the same type as the pseudo-second information converted by the conversion unit is within a threshold value; An update instruction unit that instructs an update of the conversion parameter; Comprising; When the difference exceeds the threshold value, the update instruction unit Instructs the image data reception unit, the point cloud data reception unit, the first information generation unit, and the second information generation unit to add data, and Instructs the conversion parameter generation unit to update the conversion parameter using the added data. A sensing system characterized by this.

15. The sensing system according to claim 13 or claim 14, wherein The camera, the sensor, and the sensing device are connected via an edge device and a communication network. A sensing system characterized by this.

Citation Information

Patent Citations

  • Object position estimation device

    JP2015215299A