Sensing device, sensing system, and sensing method

JP7866493B2Active Publication Date: 2026-05-27HITACHI LTD
View PDF 2 Cites 0 Cited by

Patent Information

Authority / Receiving Office
JP · JP
Patent Type
Patents
Current Assignee / Owner
HITACHI LTD
Filing Date
2022-12-08
Publication Date
2026-05-27

AI Technical Summary

Technical Problem

Existing methods for generating training data for object detection in infrastructure sensors result in high capacity and cost, with potential scene bias due to fixed viewpoints, object types, acquisition times, and weather conditions, leading to decreased performance in varied environments.

Method used

A sensing device and method that constructs training data by combining camera images with point clouds, extracts and corrects misinferred objects, and adjusts data subsets to reduce bias, using a system with image and point cloud recognition units, generation and adjustment units to enhance object detection accuracy.

Benefits of technology

Enables the creation of high-precision training data that reduces scene bias, improving object detection performance across diverse environments and conditions, ensuring accurate detection of objects even in challenging scenarios.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure 0007866493000009
    Figure 0007866493000009
  • Figure 0007866493000010
    Figure 0007866493000010
  • Figure 0007866493000011
    Figure 0007866493000011
Patent Text Reader

Abstract

To provide a sensing apparatus capable of appropriately constructing learning data of a detector which detects an object from an image acquired from a camera.SOLUTION: A sensing apparatus 100 includes: an image reception unit 1 which receives an image acquired by a camera that images an object; a point cloud reception unit 2 which receives a point cloud acquired by a sensor that measures the object; an image recognition unit 3 which infers the object included in the image using machine learning; a point cloud recognition unit 4 which recognizes the object included in the point cloud; an extraction unit 5 which extracts an object which has not been inferred or falsely inferred by the image recognition unit, using a result recognized by the point cloud recognition unit 4; a generation unit 6 which generates object information related to the object not inferred or falsely inferred extracted by the extraction unit 5; a storage unit 7 which stores the image in association with the object information generated by the generation unit 6; and an adjustment unit 8 which extracts a set of subsets so as to prevent the contents of the images from being biased, for pairs of the images and multiple pieces of object information stored by the storage unit 7.SELECTED DRAWING: Figure 1
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention relates to a sensing device, a sensing system, and a sensing method for constructing learning data of a detector that detects an object from an image acquired by a camera.

Background Art

[0002] In recent years, there has been an increasing need for sensing technologies that detect, track, and predict the path of a target by analyzing data acquired by infrastructure sensors such as cameras and LiDAR (Laser Imaging Detection And Ranging) installed in the environment. In particular, in the field of autonomous driving, by installing an infrastructure sensor that measures areas that are blind spots from the vehicle, it is expected to avoid accidents by estimating the presence or absence of objects on the vehicle path and the predicted trajectory and notifying the vehicle, thereby realizing safe and secure movement.

[0003] In object recognition using a camera, in recent years, a method of estimating the type and area of an object using a machine learning-based detector such as a neural network has shown high performance. On the other hand, for machine learning, it is necessary to create learning data in advance for parameter adjustment of the detector, and it takes time and effort when manually performing the annotation work of assigning true values to the true values of the object type and area. In particular, in order to achieve a high recognition rate in machine learning for infrastructure sensors, it is effective to additionally learn local data acquired in the installation environment in addition to data acquired in various scenes.

[0004] In Patent Document 1, data is sequentially output from two types of sensors such as a camera and LiDAR installed in a vehicle, and when the reliability of the recognition result based on the first sensor data is above a predetermined degree, the second sensor output is used as an input in machine learning, and the recognition result of the first sensor is associated as teacher data, thereby generating and storing teacher data for parameter adjustment in the machine learning of the first sensor, reducing the labor of annotation.

Prior Art Documents

Patent Documents

[0005] [Patent Document 1] Patent No. 6682833 [Overview of the Initiative] [Problems that the invention aims to solve]

[0006] However, when training data is generated from sequentially acquired sensor data using the technology described in Patent Document 1, the amount of training data becomes enormous, resulting in high capacity and training costs. In particular, with infrastructure sensors, the viewpoint is fixed, so there is a possibility that a large amount of the data will include scenes that can already be recognized through past training. In addition, there is a possibility of scene bias due to the position and type of objects, the time of acquisition, and the weather, and if additional training to update the parameters of the machine learning detector is performed with biased data, the performance may be high for that scene, but the performance may decrease for other scenes.

[0007] The present invention is for solving the aforementioned problems and provides a sensing device, a sensing system, and a sensing method that can appropriately construct training data for a detector that detects objects using images acquired from a camera. [Means for solving the problem]

[0008] To solve the aforementioned problems, the present invention provides a sensing device that constructs learning data for a detector that detects objects using images acquired from a camera, comprising: an image receiving unit that receives images acquired by a camera that images the object; a point cloud receiving unit that receives point clouds acquired by a sensor that measures the object; an image recognition unit that infers objects included in the image using machine learning; a point cloud recognition unit that recognizes objects included in the point cloud; an extraction unit that extracts objects not inferred or misinferred by the image recognition unit using the recognition results from the point cloud recognition unit; a generation unit that generates object information related to the uninferred or misinferred objects extracted by the extraction unit; a storage unit that associates and stores the object information and images generated by the generation unit; and an adjustment unit that extracts a subset of the multiple object information and image sets stored by the storage unit such that the bias in image content is reduced. Furthermore, the bias in the image content refers to a bias in at least one of the following: the type of object, the position of the object, the date and time of image capture, or the weather conditions at the time of image capture. The adjustment unit includes a single-map adjustment unit that reduces bias within a map showing the frequency of positions in a single scene, and a multi-map adjustment unit that reduces bias between multiple maps. This invention is characterized by the following embodiments. Other aspects of the present invention will be described in the embodiments described below. [Effects of the Invention]

[0009] Images acquired from a camera can be used to appropriately construct training data for object detection detectors. [Brief explanation of the drawing]

[0010] [Figure 1] This figure shows the functional block of the sensing device according to the embodiment. [Figure 2] This figure shows an example of detection by the image recognition unit according to the embodiment. [Figure 3A] This is a flowchart showing the detection process (part 1) by the point cloud recognition unit according to the embodiment. [Figure 3B] This is a flowchart showing the detection process (part 2) by the point cloud recognition unit according to the embodiment. [Figure 3C] This is a flowchart showing the detection process (part 3) by the point cloud recognition unit according to the embodiment. [Figure 4] This diagram shows the functional block of the extraction unit according to the embodiment. [Figure 5]It is a diagram showing an example of projection of point group recognition information by a projection unit according to an embodiment. [Figure 6] It is a diagram showing an example of a scene calculated to have a low reliability by a reliability determination unit according to an embodiment. [Figure 7] It is a diagram showing a functional block of a generation unit according to an embodiment. [Figure 8] It is a diagram showing an example of generating object information by a point group-based generation unit according to an embodiment. [Figure 9] It is a diagram showing a functional block of an adjustment unit according to an embodiment. [Figure 10] It is a diagram showing an example of associating an image with a map by a map generation unit according to an embodiment. [Figure 11] It is a diagram showing an example of reducing the bias in a map by a single map adjustment unit according to an embodiment. [Figure 12] It is a diagram showing an example of reducing the bias between maps by a plurality of map adjustment units according to an embodiment. [Figure 13A] It is a diagram showing the relationship (Part 1) between a camera and a sensor and a sensing device. [Figure 13B] It is a diagram showing the relationship (Part 2) between a camera and a sensor and a sensing device.

Embodiments for Carrying Out the Invention

[0011] Hereinafter, specific embodiments of the present invention will be described with reference to the drawings. In the example of this embodiment, the camera is installed so as to view the objects around the crosswalk obliquely, and the sensor is installed so as to measure the objects around the crosswalk, but it is not limited to this example.

[0012] (Overall Configuration) FIG. 1 is a diagram showing a functional block of a sensing device 100 according to an embodiment. FIG. 13A is a diagram showing the relationship (Part 1) between the camera 110 and the sensor 120 and the sensing device 100. FIG. 13B is a diagram showing the relationship (Part 2) between the camera 110 and the sensor 120 and the sensing device 100.

[0013] In this embodiment, an embodiment in which a monocular camera is used as the camera 110 and LiDAR is used as the sensor 120 will be described, but it is not limited thereto. The camera 110 may be an infrared camera or a stereo camera, and other sensors such as a distance sensor may be used as the sensor 120 instead.

[0014] As shown in FIG. 13A, the sensing device 100 is capable of receiving data from the camera 110 and the sensor 120. The sensing device 100 includes a processing unit 20, a storage unit 21, an input unit 22, a display unit 23, a communication unit 24, and the like. The processing unit 20 is a central processing unit (CPU) that executes various programs stored in a RAM, an HDD, or the like. The storage unit 21 is an HDD or the like that stores various data for the sensing device 100 to execute processing. The input unit 22 is a device for inputting instructions to a computer such as a keyboard or a mouse, and inputs instructions such as program startup. The display unit 23 is a display or the like that displays the execution status and execution results of processing by the sensing device 100. The communication unit 24 is a device that exchanges various data and commands with other devices via a network.

[0015] Also, as shown in FIG. 13B, the sensing device 100 may receive the data of the camera 110 and the sensor 120 via the edge device 130. Note that an edge refers to a point (a terminal of a network) that sends out data collected in a terminal and a terminal-side network to a line.

[0016] The sensing device 100 shown in Figure 1 is a device that automatically generates training data for additional training of a machine learning-based detector that recognizes moving objects such as people and cars from camera images, using images captured by the camera 110 and point clouds acquired from the sensor 120, and performs additional training. Each part (each function) in Figure 1 is realized in the processing unit 20. The memory unit 21 stores machine learning parameters 201 (see Figure 2), image recognition information 202 (see Figure 4), point cloud recognition information 203 (see Figure 4), position and orientation parameters 204 (see Figure 4), camera parameters 205 (see Figure 4), projection information 206 (see Figure 5), advanced inference parameters 208 (see Figure 7), etc., which will be described later.

[0017] The overview of each function shown in Figure 1 is explained below. The image receiving unit 1 has the function of acquiring images captured by the camera 110, and the point cloud receiving unit 2 has the function of acquiring point clouds acquired from the sensor 120 in synchronization with the image captured by the image receiving unit 1.

[0018] The image recognition unit (image inference unit) 3 has the function of inferring image recognition information 202 (see Figure 4), which is information about the position (region) and type of objects in the image, by inputting the image into a machine learning-based detector, and the point cloud recognition unit 4 has the function of estimating point cloud recognition information 203 (see Figure 4), which is information about the position (region), size, and type of objects included in the point cloud, by analyzing the point cloud.

[0019] The extraction unit 5 has the function of extracting objects that are not detected or are falsely detected by the detector from the image recognition information 202 and point cloud recognition information 203; the generation unit 6 has the function of assigning object information, which is information about the location (region) and type of the object, to the undetected or falsely detected objects; and the storage unit 7 has the function of storing the object information and the image together. Specifically, the object information includes rectangular information that encloses the area surrounding the object on the image, and type information that indicates the type of object.

[0020] The adjustment unit 8 has the function of extracting subsets of multiple stored object information and image sets that reduce the bias in image content; the learning unit 9 has the function of updating the parameters of a machine learning-based detector by additional learning using the subsets; and the correction unit 10 has the function of presenting object information and images to the user and allowing the user to correct the generated object information.

[0021] Image content bias refers to a bias in at least one of the following: the type of object, the position of the object, the date and time the image was taken, or the weather conditions at the time the image was taken. Furthermore, a subset means a part, a subset, or a subgroup, and is a small group formed by taking some of the elements that make up the whole group.

[0022] The following provides a detailed explanation of each function. (Explanation of Image Recognition Unit 3) Figure 2 shows an example of detection by the image recognition unit 3 according to the embodiment. Figure 2 shows an example of how the image recognition unit 3 infers image recognition information 202. The image recognition unit 3 receives an image from the image receiving unit 1 and uses machine learning parameters 201, which are parameters such as the weights of a machine learning-based detector, to infer image recognition information 202 in the inference unit 31.

[0023] In this embodiment, the inference unit 31 is a detector that performs a detection task of inferring image recognition information 202 consisting of the types of multiple moving objects in the image and the position and size of the rectangle surrounding the moving objects. However, the example is not limited to this, and for example, objects other than moving objects may be targeted, or the task may be segmentation, such as a task to predict the type of object for all pixels in the image, a task to predict labels for all objects in the image and assign a unique ID, or a task to predict class labels for all pixels in the image and estimate a unique ID.

[0024] In Figure 2, if the machine learning parameter 201 has not acquired the ability to detect objects in the input image during inference, non-inference or misinference may occur. Non-inference means that even though the target object is actually present, no object information related to that object can be obtained from the image. Misinference means that even though the target object is not actually present, object information is incorrectly inferred to be present.

[0025] In Figure 2, although there is actually a person in the center of the image, object information related to that person has not been inferred, and in this case, it is considered uninferred. However, the image recognition unit 3 has not received any information that it has not been able to infer object information related to that person.

[0026] Furthermore, the sensor mistakenly identified the structure in the upper left as a person and generated object information, which is a misinference. For infrastructure sensors, it is desirable to update the detector's machine learning parameter 201 by additionally learning scenes that the sensor struggles with in the installation environment, thereby improving its object detection capability for images acquired in the installation environment.

[0027] (Explanation of Point Cloud Recognition Unit 4) Figure 3A is a flowchart showing the detection process (1) by the point cloud recognition unit 4 according to the embodiment. Figure 3B is a flowchart showing the detection process (2) by the point cloud recognition unit 4 according to the embodiment. Figure 3C is a flowchart showing the detection process (3) by the point cloud recognition unit 4 according to the embodiment.

[0028] The point cloud recognition unit 4 estimates point cloud recognition information 203 by inputting a point cloud. In this embodiment, an example is shown in which the object position (region, centroid), size, and type are estimated in the point cloud recognition information, but the reliability of recognition may also be estimated from the reflection intensity of the point cloud of the object part, for example, and the embodiment is not limited to this example.

[0029] The flow of the point cloud recognition unit 4 is explained using Figure 3A. If there are multiple sensors, the point cloud recognition unit 4 integrates the point clouds from these sensors (step S1), then removes the background point cloud by taking the difference with a pre-generated background point cloud, and extracts the point cloud of the moving object (step S2). The point cloud recognition unit 4 then divides the 3D region of the point cloud of the moving object into multiple grids called Pillars, extracts features such as the number and variance of points in each Pillar, generates a 2D map that provides an overview of this grid (step S3), determines the presence or absence of a moving object in each grid of the map (step S4), groups the moving object grids (step S5), and estimates the position, size, and type of the object (step S6). This method (method example M1) is a point cloud recognition method called Pillar-based, but there are also voxel-based (method example M2) and machine learning-based (method example M3) methods.

[0030] The flow of the point cloud recognition unit 4 is explained using Figure 3B. In the voxel-based method (method example M2), the point cloud recognition unit 4 performs steps S1 and S2, converts the point cloud in space into voxels composed of cubes (step S7), determines the presence or absence of an object for each voxel (step S8), groups the object voxels (step S9), and estimates the object's center of gravity, size, and object type (step S10).

[0031] The flow of the point cloud recognition unit 4 is explained using Figure 3C. In a machine learning-based method (method example M3), the point cloud recognition unit 4 performs steps S1 and S2, and then inputs the point cloud into a detector such as a neural network to directly estimate the position, size, and type of the object (step S11).

[0032] In this embodiment, the point cloud of the moving object is extracted by removing the background point cloud in step S2, thereby improving the accuracy of subsequent processing. However, this is not limited to this example; for example, stationary objects may be recognized without removing background differences. Furthermore, any method capable of estimating point cloud recognition information 203 can be used similarly, not limited to method examples M1, M2, and M3. The point cloud recognition unit 4 described above can estimate information regarding the type, location (region), and size of objects included in the point cloud.

[0033] (Explanation of extraction unit 5) Figure 4 shows the functional block of the extraction unit 5 according to the embodiment. The extraction unit 5 extracts objects that are not detected or are falsely detected by the detector from the image recognition information 202 and the point cloud recognition information 203. The extraction unit 5 includes a projection unit 51, a consistency determination unit 52, and a reliability determination unit 53. The position and orientation parameters 204 are parameters that indicate the relative position and orientation relationship between the camera 110 and the sensor 120, and the camera parameters 205 are camera parameters that include information on the camera's focal length and optical axis.

[0034] The projection unit 51 has the function of projecting image recognition information 202 and point cloud recognition information 203 onto the same coordinate system using position and orientation parameters 204 and camera parameters 205.

[0035] The consistency determination unit 52 has the function of determining an unpredicted or mispredicted object by determining the consistency of the two projected recognition pieces of information, and the reliability determination unit 53 has the function of determining the final unpredicted or mispredicted object by calculating the reliability of the information from the consistency determination unit 52.

[0036] Specifically, in the sensing device 100 of this embodiment, the projection unit 51 superimposes the point cloud corresponding to the object region identified by the point cloud recognition unit 4, or the position information calculated from the point cloud corresponding to the object region, by associating it with the pixels of the image, and the consistency determination unit 52 determines the presence or absence of uninferred or misinferred objects based on the proportion of position information contained within a predetermined range from the rectangle containing the object information inferred by the image recognition unit 3.

[0037] The projection unit 51, consistency determination unit 52, and reliability determination unit 53 will be described below. Figure 5 shows an example of projection of point cloud recognition information 203 by the projection unit 51 according to the embodiment. Figure 5 shows an example of superimposing point cloud recognition information 203 onto an image. In this embodiment, an example is shown in which the object point cloud and position obtained from the object region included in the point cloud recognition information 203 are superimposed onto the image. However, it is also possible to add image information to the point cloud or to define a separate common coordinate system and superimpose it.

[0038] To map the point cloud sP of the sensor coordinate system to the pixels of the image, first, in equation (1), the point cloud sP of the sensor coordinate system is converted to the point cloud cP of the camera coordinate system using the position and orientation parameter cTs (position and orientation parameter 204).

[0039] Note that the point cloud cP in the camera coordinate system and the point cloud sP in the sensor coordinate system shown in equation (2) are usually represented as three-dimensional points (x, y, z), but for dimensional alignment in calculations, they are written in a form where 1 is added in the dimensional direction, such as (x, y, z, 1). The position and orientation parameter cTs shown in equation (3) is a position and orientation parameter that represents the transformation from the camera coordinate system to the sensor coordinate system, and is a 4x4 matrix composed of a 3x3 rotation matrix cRs representing the rotation shown in equation (4) and a translation matrix cTs representing the translation shown in equation (5).

[0040]

number

number

number

number

number

[0041] Next, using equation (6), the point cloud cP in the camera coordinate system is transformed into pixels (u, v) on the image using the camera parameters 205. In this embodiment, various processing is performed on an image that has been distorted, but this is not limited to this.

[0042] In equation (6), fx and fy are the camera's focal length parameters, and cx and cy are the principal points of the image center, i.e., the optical center parameters of the lens. These camera parameters 205 can be determined in advance using a method called camera calibration. Note that the values ​​of Xc, Yc, and Zc in equation (6) are those obtained in equation (1). By solving equation (6), the point cloud cP in the camera coordinate system can be transformed into pixels (u, v) on the image using the camera parameters 205, as shown in equations (7) and (8).

[0043]

number

number

number

[0044] The projection information 206 in Figure 5 is an example of visualizing the superimposition of point cloud recognition information 203 on image recognition information 202, representing the point cloud of the human region, its centroid position, and size (region) information using the pixel positions in the image. In other words, in projection information 206, both the point cloud recognition information 203 and the image recognition information 202 are correlated in the same dimension.

[0045] The consistency determination unit 52 has the function of determining unpredicted or mispredicted objects by determining the consistency of the two types of recognition information contained in the projection information 206. To explain using the example of projection information 206 in Figure 5, the point cloud recognition information 203 determines that there is a person as a moving object in the person region of the image, but the image recognition information 202 does not detect a moving object in the person region of the image. In other words, it is determined that the object is unpredicted (not detected).

[0046] Similarly, if the image recognition unit 3 detects an object that the point cloud recognition unit 4 has not detected, it is determined to be a false inference (false detection). This determination method includes, for example, determining whether the centroid position determined by the point cloud recognition information 203 is included within the rectangle determined by the image recognition information 202, or determining whether a certain percentage or more of the object region point cloud determined by the point cloud recognition information 203 is included within the rectangle. Furthermore, considering the timing difference between the camera 110 and the sensor 120, and errors in the position and orientation parameters 204 and camera parameters 205, the vicinity of the rectangle determined by the image may be searched and the presence or absence of a matching object detected by the point cloud recognition information 203 may be determined, and the method is not limited to this example.

[0047] The reliability determination unit 53 calculates the reliability of the determination for objects that the consistency determination unit 52 has determined to be misinferred or not inferred, and then decides whether or not to actually consider the data to be misinferred or not inferred. In the camera 110, due to the effect of brightness, an object may not actually be obtained as an image even if it is within the field of view. Also, sensors such as LiDAR cannot always assign the true correct value, and there are scenes and objects that are difficult to measure and recognize. In such scenes, errors may occur in the consistency determination unit 52's determination of misinferred or not inferred.

[0048] Figure 6 shows examples of scenes in which the reliability determination unit 53 according to the embodiment calculates a low reliability. The reliability determination unit 53 will be explained using Figure 6. Scene 6a in Figure 6 is a case where a lot of light is incident from the left side of the image, and the object in the image is not clearly visible due to overexposure. Scene 6b in Figure 6 is a case where the lighting is not on and it is dark, such as at night, and the image of the object in the image is not clearly visible. In such cases, even if the consistency determination unit 52 determines that it is not inferred, it is not desirable to assign training data to this object and perform learning.

[0049] Therefore, if the brightness around a moving object detected by the point cloud recognition unit 4 but not detected by the image recognition unit 3 is higher or lower than a predetermined value, the confidence level can be calculated lower. This prevents undesirable data from being judged as uninferred due to image brightness. Furthermore, while a method using a threshold for the brightness value of the image around an object was shown for brightness determination, this is not limited to this example. For example, the confidence level could be calculated lower if the difference between past data and the brightness value of a stationary object region is greater than or equal to a threshold, or the confidence level could be calculated by detecting overexposure using machine learning or by directly estimating the brightness of the scene. Similarly, a detection unit that detects abnormalities in the camera or sensors could be provided, and the confidence level could be calculated lower if an abnormality is detected.

[0050] In the sensing device 100 of this embodiment, if the extraction unit 5 determines that an image of an object cannot be obtained in a predetermined area of ​​the image due to the influence of the brightness of the image, it can set the object in the predetermined area as not being unpredicted.

[0051] Scene 6c in Figure 6 is a case where the LiDAR sensor's point cloud has large errors or omissions due to the color or material of the object, such as the object being black or made of an absorbent material. Scene 6d in Figure 6 is a scene that the sensor struggles with, such as when the person is positioned sideways in a direction that is far from the LiDAR sensor and / or where the LiDAR sensor's resolution is low. In such cases, even if the consistency determination unit 52 determines that it is a misinference, it is undesirable to add information that there is no object in this object region to the training data and perform learning.

[0052] Therefore, for example, when detecting an object region on a point cloud image, if the reflectivity of the point cloud corresponding to the region is low, or if the color of the image of the object region is determined to be black, or if the number of point clouds of the object portion estimated from the LiDAR resolution is small, calculating a lower confidence level can prevent undesirable data such as the object's color, material, distance, and shape from being misinferred.

[0053] In the sensing device 100 of this embodiment, the extraction unit 5 can be set to not be an erroneous inference if the confidence level of the point cloud recognition unit 4 is below a predetermined level.

[0054] (Explanation of generation unit 6) The generation unit 6 has the function of assigning object information, which is information about the location (region) and type of the object, to undetected or falsely detected objects. More specifically, for unpredicted objects, it assigns the type and location (centroid, region) as object information on the image, and for falsely predicted objects, it generates training data by deleting the object type and location information contained in the corresponding image recognition information 202.

[0055] Figure 7 shows the functional blocks of the generation unit 6 according to the embodiment. The generation unit 6 includes a point cloud-based generation unit 61, an image-based generation unit 62, and a false detection information deletion unit 63. The point cloud-based generation unit 61 has the function of generating object information using point cloud recognition information 203, the image-based generation unit 62 has the function of generating object information using inference results using a different detector than the image recognition unit, and the false detection information deletion unit 63 has the function of deleting object information related to falsely inferred objects. The point cloud-based generation unit 61 and the image-based generation unit 62 will be described in detail below.

[0056] Figure 8 shows an example of object information being generated by the point cloud-based generation unit 61. Refer to Figure 7 as appropriate. The point cloud-based generation unit 61 includes a bounding rectangle generation unit 611 that surrounds the bounding rectangle of the point cloud of the object region, as shown in generation example 8a of Figure 8, and an expansion unit 612 that expands the bounding rectangle, as shown in generation example 8b of Figure 8. In the bounding rectangle of the point cloud generated by the bounding rectangle generation unit 611, a rectangle smaller than the actual object may be generated, as shown in generation example 8a of Figure 8, due to reasons such as the low resolution of the sensor, such as LiDAR. In particular, the vertical direction, where the resolution of LiDAR is low, tends to be estimated to be smaller. Therefore, the expansion unit 612 expands the rectangle to generate highly accurate object information. In the case of generation example 8b, the rectangle is expanded in the vertical direction compared to generation example 8a.

[0057] As a means of expanding the rectangle, for example, the magnification factor can be determined from the LiDAR resolution and distance to the object information, or the background scene can be acquired as an image in advance, and the object region can be easily estimated by taking the difference between the current image and the background image, and then the rectangle can be expanded to match that object region, or the rectangle can be expanded to match the estimated object region by inputting it into a machine learning-based classifier that performs image segmentation. Furthermore, in addition to simple magnification, assuming that the point cloud superimposed on the image will be shifted due to errors in the position and orientation parameters 204 and camera parameters 205, the position can be corrected by matching the object region estimated by the above method with the bounding rectangle generated by the bounding rectangle generation unit 611 through a nearest neighbor search.

[0058] The image-based generation unit 62 is a function that generates object information using inference results from a detector different from that of the image recognition unit 3. The image-based generation unit 62 has an advanced inference unit 613 that infers object information using advanced inference parameters 208, which are parameters of a machine learning-based detector different from those of the image recognition unit 3. Generally, in object detection processing using infrastructure sensors for autonomous driving, real-time inference is required, so the detectors included in the image recognition unit 3 are often lightweight. On the other hand, although real-time performance is low, if a high-precision detector with a large number of parameters and calculations is used, it may be possible to generate object information for objects that have not been inferred by the image recognition unit 3.

[0059] In other words, the image-based generation unit 62 attempts to generate object information using a machine learning-based detector different from the image recognition unit 3. If the position and size of the unpredicted object obtained by superimposing the point cloud recognition information 203 onto the image match, it generates the object information.

[0060] By combining the image-based generation unit 62 and the point cloud-based generation unit 61 as follows, it is conceivable to generate object information more accurately. For example, the image-based generation unit 62 is first applied to an image containing unpredicted objects. In this case, if object information can be generated for objects determined to be unpredicted, this object information is saved in the subsequent storage unit 7. On the other hand, if the image-based generation unit 62 is unable to generate reasonable object information for the unpredicted object region, the point cloud-based generation unit 61 attempts to generate object information. Generally, since images have higher resolution than point clouds, the above order is used because estimating the scene, including the type of object, is superior to methods using point clouds.

[0061] Furthermore, since the object information generated by the point cloud-based generation unit 61 using the above procedure may contain more errors in object type compared to the image-based generation unit 62, it is preferable to display the object information on the display unit 23 (see Figure 13A) in a format like the generation example 8b in Figure 8, so that the user can input and correct whether the generated object type is correct or not via the input unit 22 (see Figure 13A). In other words, the sensing device 100 has a correction unit 10 that presents object information and images to the user and allows the user to correct the generated object information, as described above.

[0062] Furthermore, the same processing may be performed not only on the point cloud-based generation unit 61 but also on the object information generated by the image-based generation unit 62. In addition to the object type, the object information may also be modified to include area information such as rectangular information surrounding the area of ​​the object on the image. Moreover, the modification unit 10 may be performed on data generated by the storage unit 7 or the adjustment unit 8, rather than on data generated by the generation unit 6.

[0063] (Explanation of adjustment unit 8) The adjustment unit 8 is a function that extracts subsets of stored object information and image sets that reduce the bias in image content. When performing additional training to update the parameters of a machine learning-based detector, it is common to collect a certain number of data on-site and combine this data with other pre-prepared data for training. When data is collected sequentially from fixed cameras or sensors, such as infrastructure sensors, the scene may be biased in terms of the type and location of objects in the image, the time of acquisition, and the weather. By reducing the bias in a certain number of on-site data used for additional training, it is possible to prevent the detector from specializing only in specific scenes.

[0064] Figure 9 shows the functional blocks of the adjustment unit 8 according to the embodiment. The adjustment unit 8 includes a map generation unit 81, a single map adjustment unit 82, and a multiple map adjustment unit 83. The map generation unit 81 has the function of generating a map that links the frequency of occurrence of the positions of moving objects included in a data set in a particular scene to data, the single map adjustment unit 82 has the function of selecting data to reduce the frequency of occurrence of positions in the map, and the multiple map adjustment unit 83 has the function of adjusting the ratio of the number of data points between multiple maps. The map generation unit 81, the single map adjustment unit 82, and the multiple map adjustment unit 83 will be described in detail below with reference to the figures.

[0065] Figure 10 shows an example of mapping an image to a map using the map generation unit 81 according to this embodiment. In this embodiment, the map is shaped by dividing the XY plane, which provides an overview of the sensing space, into multiple grids, and each grid is associated with the frequency of appearance of a specific object and its data.

[0066] In this embodiment, maps are generated for each scene, and also for each type of object. For example, maps are created for each acquisition time, uninferred or misinferred object type, and weather, such as "20xx / 04 Daytime Uninferred Object Car Sunny" or "20xx / 04 Nighttime Uninferred Person Rainy". The acquisition time is determined by the processing unit 20 (see Figure 13A), and the weather can be input by a human, for example, the weather at the time of data acquisition, or a detector that automatically estimates the weather from the image can be provided, and the results can be used. Note that the method of dividing the scenes when generating the maps is not limited to this example.

[0067] In the example in Figure 10, a map is generated for unspecified individuals during the daytime on April 20xx. Based on the location of people detected by sensors such as LiDAR, data including these unspecified individuals is mapped to the grid of the map. The grid color indicates the number of data points assigned, and the distribution of color intensity visualizes the frequency of object appearance.

[0068] The sensing device 100 may also have a storage unit 21 (see Figure 13A) that stores information regarding the frequency of unpredicted or mispredicted values, and a notification unit that compares the latest frequency information with past information and notifies if it is determined that the frequency has increased.

[0069] Figure 11 shows an example of reducing bias within a map using the single map adjustment unit 82 according to the embodiment. As shown in Figure 11, data is extracted and generated so that the frequency of data occurrence is not biased. Methods for generating a corrected map using a map include, for example, randomly deleting data from grids with high frequency in the map and updating the map that does not contain that data, but are not limited to this example. In this case, the maximum number of data points to be included in each grid of the corrected map may be determined by setting a threshold in advance, or the most frequent range of the number of objects in each grid may be used. Note that it is not necessarily required to strictly match the number of data points included in all grids.

[0070] Figure 12 shows an example of reducing bias between maps using the multiple map adjustment unit 83 according to the embodiment. Specifically, as shown in Figure 12, the multiple map adjustment unit 83 adjusts the ratio of data between multiple correction maps generated by the single map adjustment unit 82. The ratio may be made equal, or the ratio between scenes may be predetermined, and the adjustment method may be random sampling or sampling in the time direction, and is not limited to this example. Furthermore, the adjustment may be performed after adding a map consisting of data used in past additional training.

[0071] (Explanation of Learning Section 9) The learning unit 9 has the function of updating the parameters of the machine learning-based detector by additional learning using a subset generated by the adjustment unit 8. The learning unit 9 additionally learns the parameters using a set of scene data with little bias saved by the storage unit 7 based on the map generated by the generation unit 6.

[0072] The data used for additional training may consist only of the subset mentioned above, or it may also include data previously generated by the generation unit 6 and used for additional training, data acquired in other scenes, or data that has been manually annotated.

[0073] As a method for additional learning, if a neural network is used as a machine learning-based detector, optimization methods such as gradient descent can be used. In this case, the entire set of parameters of the neural network may be adjusted, or some parameters (for example, the layer that performs feature extraction close to the input layer of the neural network) may be fixed while other parameters (for example, the layer that performs tasks close to the output layer) may be adjusted. Alternatively, the adjustment weights may be changed for each parameter (for example, the learning rate may be higher for layers closer to the output layer). These methods may be appropriately selected depending on the number and type of data used for additional learning. In this way, the machine learning parameters 201 in the image recognition unit 3 may be updated, and the advanced inference parameters 208 may also be updated in the same manner.

[0074] (Explanation of the entire loop) For example, during system implementation, after the learning unit 9 updates the machine learning parameters 201 from the image receiving unit 1 to achieve sufficient accuracy for practical use, the system enters an operational phase where infrastructure sensors detect moving objects and notify vehicles. In this operational phase, if environmental changes occur that were not considered during system implementation (for example, if the system is built in the summer and then snow accumulates in the winter), the machine learning parameters 201 may not be able to cope with that scene.

[0075] Therefore, the sensing device 100 may be operated even during operation, and the extraction unit 5 may collect unpredicted or mispredicted data. If it is detected that the number of unpredicted or mispredicted objects has increased beyond a predetermined level compared to previous additional learning, the notification unit may notify the system of the performance degradation, and the learning unit 9 may be executed again based on that notification. In this way, accuracy can be ensured by detecting the performance degradation in response to environmental changes and performing additional learning on the detector again.

[0076] (Effects of the invention) The sensing device 100 described above is a sensing device that constructs training data for a detector that detects objects. It can automatically extract scenes that are difficult to recognize, annotate them, generate a dataset with less scene bias, and then perform additional training.

[0077] Furthermore, the extraction unit 5 of this embodiment projects the image recognition information 202 and the point cloud recognition information 203 onto the same coordinate system, and then determines the consistency of the two recognition pieces of information to extract candidates for unpredicted or mispredicted objects. The reliability determination unit 53 then calculates the reliability of the unpredicted or mispredicted candidates based on the characteristics of the camera and sensors, thereby enabling high-precision extraction of unpredicted or mispredicted objects.

[0078] Furthermore, the generation unit 6 of this embodiment can assign type and position (centroid, region) as object information on the image to unpredicted objects, and delete the object type and position information included in the corresponding image recognition information 202 for mispredicted objects.

[0079] Furthermore, the adjustment unit 8 of this embodiment can extract subsets of stored object information and image sets that reduce the bias in image content. In particular, the single-map adjustment unit 82 can reduce positional bias by selecting data to reduce the frequency of occurrence of a location within a map in a single scene, and the multiple-map adjustment unit 83 can adjust the ratio of the number of data points between multiple maps.

[0080] Furthermore, the learning unit 9 of this embodiment can generate parameters that can handle various scenarios by updating the parameters of the machine learning-based detector using the generated subset.

[0081] Furthermore, even during operation, the extraction unit 5 collects data on unpredicted or mispredicted objects, and saves these numbers as logs. If the number of unpredicted or mispredicted objects increases beyond a predetermined level compared to logs after previous additional training, the notification unit notifies the system of performance degradation. Based on this notification, the learning unit 9 is made to run again, thereby detecting performance degradation in response to environmental changes and ensuring accuracy by performing additional training on the detector.

[0082] With the above configuration, even if there is a bias in the number of scenes in the collected data, such as the location and type of objects, the time of acquisition, or the weather, the system can automatically extract and annotate scenes that are difficult to recognize, generate a dataset with less scene bias, and efficiently perform additional training to improve performance early on.

[0083] Furthermore, with the sensing device 100 described above, if there is a moving object such as a person or vehicle in a blind spot for an autonomously driving vehicle, the sensing device can detect the moving object using a detector trained on the learning data it generates, and notify the vehicle of the object information, enabling control such as slowing down or stopping the vehicle.

[0084] Furthermore, by installing visible and infrared cameras and sensors in areas with extreme differences in light and shadow, such as tunnel exits, where object detection by cameras installed in autonomous vehicles is likely to fail, it is possible to further improve safety by notifying the vehicle of objects that it has failed to detect.

[0085] The sensing device 100 has been described above, but the sensing system and sensing method are as follows. The sensing system is a sensing system that constructs training data for a detector that detects objects using images acquired from a camera, and comprises: an image receiving unit 1 that receives images acquired by a camera that images objects; a point cloud receiving unit 2 that receives point clouds acquired by a sensor that measures objects; an image recognition unit 3 that infers objects contained in the image using machine learning; a point cloud recognition unit 4 that recognizes objects contained in the point cloud; an extraction unit 5 that extracts objects that were not inferred or misinferred by the image recognition unit 3 using the recognition results from the point cloud recognition unit 4; a generation unit 6 that generates object information related to the objects that were not inferred or misinferred by the extraction unit 5; a storage unit 7 that associates and stores the object information and images generated by the generation unit 6; and an adjustment unit 8 that extracts a subset of the multiple object information and image sets stored by the storage unit 7 such that the bias in image content is reduced. This allows for the appropriate construction of training data for a detector that detects objects using images acquired from a camera.

[0086] The sensing method is a sensing method for constructing training data for an object detection detector using images acquired from a camera, and comprises: an image reception step that receives images acquired by a camera that images objects; a point cloud reception step that receives point clouds acquired by a sensor that measures objects; an image recognition step that infers objects contained in the image using machine learning; a point cloud recognition step that recognizes objects contained in the point cloud; an extraction step that extracts objects that were not inferred or misinferred in the image recognition step using the recognition results from the point cloud recognition step; a generation step that generates object information related to the objects that were not inferred or misinferred extracted in the extraction step; a storage step that associates and saves the object information and images generated in the generation step; and an adjustment step that extracts a subset of the multiple object information and image sets saved in the storage step such that the bias in image content is reduced. This makes it possible to appropriately construct training data for an object detection detector using images acquired from a camera. [Explanation of Symbols]

[0087] 1. Image Reception Section 2. Point cloud reception section 3. Image Recognition Unit (Image Inference Unit) 4 Point cloud recognition section 5 Extraction part 6 Generation part 7 Storage section 8 Adjustment section 9. Learning Department 10 Correction section 20 Processing Units 21 Memory section 22 Input section 23 Display section 24 Communications Department 31 Reasoning part 51 Projection section 52 Consistency judgment section 53. Confidence Determination Unit 61 Point cloud base generation unit 62 Image-based generation unit 81 Map Generation Unit 82 Single Map Adjustment Unit 83 Multiple Map Adjustment Section 100 Sensing devices 110 Camera 120 sensors 130 Edge Device 201 Machine Learning Parameters 202 Image Recognition Information 203 Point cloud recognition information 204 Position and Attitude Parameters 205 Camera Parameters 206 Projection Information 208 Advanced Inference Parameters 611 Circumscribed rectangle generator 612 Expansion section 613 Advanced Reasoning Department

Claims

1. A sensing device that constructs training data for a detector that detects objects from images acquired from a camera, An image receiving unit that receives an image acquired by a camera that images the aforementioned object, A point cloud receiving unit that receives the point cloud acquired by the sensor that measures the aforementioned object, An image recognition unit that uses machine learning to infer objects contained in the image, A point cloud recognition unit that recognizes objects included in the point cloud, An extraction unit that uses the recognition results from the point cloud recognition unit to extract objects that the image recognition unit has not inferred or has misinferred, A generation unit generates object information related to the unpredicted or mispredicted objects extracted by the extraction unit, A storage unit that associates and stores the object information and image generated by the generation unit, The storage unit has an adjustment unit that extracts subsets of object information and image sets stored by the storage unit that reduce the bias in the image content. The aforementioned bias in image content means that at least one of the following is biased: type of object, position of object, date and time of image capture, or weather conditions at the time of image capture. The adjustment unit includes a single-map adjustment unit that reduces bias within a map showing the frequency of positions in a single scene, and a multi-map adjustment unit that reduces bias between multiple maps. A sensing device characterized by the following features.

2. A sensing device that constructs training data for a detector that detects objects from images acquired from a camera, An image receiving unit that receives an image acquired by a camera that images the aforementioned object, A point cloud receiving unit that receives the point cloud acquired by the sensor that measures the aforementioned object, An image recognition unit that uses machine learning to infer objects contained in the image, A point cloud recognition unit that recognizes objects included in the point cloud, An extraction unit that uses the recognition results from the point cloud recognition unit to extract objects that the image recognition unit has not inferred or has misinferred, A generation unit generates object information related to the unpredicted or mispredicted objects extracted by the extraction unit, A storage unit that associates and stores the object information and image generated by the generation unit, An adjustment unit extracts subsets of multiple object information and image sets stored by the storage unit that reduce the bias in image content, A storage unit that stores information regarding the frequency of unpredicted or mispredicted results, It includes a notification unit that compares the latest frequency information with past information and notifies if it is determined that the frequency has increased. A sensing device characterized by the following features.

3. A sensing device that constructs training data for a detector that detects objects from images acquired from a camera, An image receiving unit that receives an image acquired by a camera that images the aforementioned object, A point cloud receiving unit that receives the point cloud acquired by the sensor that measures the aforementioned object, An image recognition unit that uses machine learning to infer objects contained in the image, A point cloud recognition unit that recognizes objects included in the point cloud, An extraction unit that uses the recognition results from the point cloud recognition unit to extract objects that the image recognition unit has not inferred or has misinferred, A generation unit generates object information related to the unpredicted or mispredicted objects extracted by the extraction unit, A storage unit that associates and stores the object information and image generated by the generation unit, The storage unit has an adjustment unit that extracts subsets of object information and image sets stored by the storage unit that reduce the bias in the image content. The aforementioned sensor is a LiDAR, If the extraction unit determines that an image of an object cannot be obtained in a predetermined area of ​​the image due to the influence of the image's brightness, it sets the object in the predetermined area as not being unpredicted. A sensing device characterized by the following features.

4. A sensing device that constructs training data for a detector that detects objects from images acquired from a camera, An image receiving unit that receives an image acquired by a camera that images the aforementioned object, A point cloud receiving unit that receives the point cloud acquired by the sensor that measures the aforementioned object, An image recognition unit that uses machine learning to infer objects contained in the image, A point cloud recognition unit that recognizes objects included in the point cloud, An extraction unit that uses the recognition results from the point cloud recognition unit to extract objects that the image recognition unit has not inferred or has misinferred, A generation unit generates object information related to the unpredicted or mispredicted objects extracted by the extraction unit, A storage unit that associates and stores the object information and image generated by the generation unit, The storage unit has an adjustment unit that extracts subsets of object information and image sets stored by the storage unit that reduce the bias in the image content. The aforementioned sensor is a LiDAR, The sensing device is characterized in that the extraction unit sets the confidence level in the point cloud recognition unit to be below a predetermined level, in which case it is not an erroneous inference.

5. In the sensing device according to claim 4, The sensing device is characterized in that the reliability is calculated based on one or more of the reflectance of the point cloud of the object region in the LiDAR, the number of points in the point cloud, or the distribution of the point cloud.

6. A sensing system that constructs training data for a detector that detects objects from images acquired from a camera, An image receiving unit that receives an image acquired by a camera that images the aforementioned object, A point cloud receiving unit that receives the point cloud acquired by the sensor that measures the aforementioned object, An image recognition unit that uses machine learning to infer objects contained in the image, A point cloud recognition unit that recognizes objects included in the point cloud, An extraction unit that uses the recognition results from the point cloud recognition unit to extract objects that the image recognition unit has not inferred or has misinferred, A generation unit generates object information related to the unpredicted or mispredicted objects extracted by the extraction unit, A storage unit that associates and stores object information and images generated by the generation unit, The storage unit has an adjustment unit that extracts subsets of object information and image sets stored by the storage unit that reduce the bias in the image content. The aforementioned bias in image content means that at least one of the following is biased: type of object, position of object, date and time of image capture, or weather conditions at the time of image capture. The adjustment unit includes a single-map adjustment unit that reduces bias within a map showing the frequency of positions in a single scene, and a multi-map adjustment unit that reduces bias between multiple maps. A sensing system characterized by the following:

7. A sensing system that constructs training data for a detector that detects objects from images acquired from a camera, An image receiving unit that receives an image acquired by a camera that images the aforementioned object, A point cloud receiving unit that receives the point cloud acquired by the sensor that measures the aforementioned object, An image recognition unit that uses machine learning to infer objects contained in the image, A point cloud recognition unit that recognizes objects included in the point cloud, An extraction unit that uses the recognition results from the point cloud recognition unit to extract objects that the image recognition unit has not inferred or has misinferred, A generation unit generates object information related to the unpredicted or mispredicted objects extracted by the extraction unit, A storage unit that associates and stores object information and images generated by the generation unit, An adjustment unit extracts subsets of multiple object information and image sets stored by the storage unit that reduce the bias in image content, A storage unit that stores information regarding the frequency of unpredicted or mispredicted results, It includes a notification unit that compares the latest frequency information with past information and notifies if it is determined that the frequency has increased. A sensing system characterized by the following:

8. A sensing system that constructs training data for a detector that detects objects from images acquired from a camera, An image receiving unit that receives an image acquired by a camera that images the aforementioned object, A point cloud receiving unit that receives the point cloud acquired by the sensor that measures the aforementioned object, An image recognition unit that uses machine learning to infer objects contained in the image, A point cloud recognition unit that recognizes objects included in the point cloud, An extraction unit that uses the recognition results from the point cloud recognition unit to extract objects that the image recognition unit has not inferred or has misinferred, A generation unit generates object information related to the unpredicted or mispredicted objects extracted by the extraction unit, A storage unit that associates and stores object information and images generated by the generation unit, The storage unit has an adjustment unit that extracts subsets of object information and image sets stored by the storage unit that reduce the bias in the image content. The aforementioned sensor is a LiDAR, If the extraction unit determines that an image of an object cannot be obtained in a predetermined area of ​​the image due to the influence of the image's brightness, it sets the object in the predetermined area as not being unpredicted. A sensing system characterized by the following:

9. A sensing system that constructs training data for a detector that detects objects from images acquired from a camera, An image receiving unit that receives an image acquired by a camera that images the aforementioned object, A point cloud receiving unit that receives the point cloud acquired by the sensor that measures the aforementioned object, An image recognition unit that uses machine learning to infer objects contained in the image, A point cloud recognition unit that recognizes objects included in the point cloud, An extraction unit that uses the recognition results from the point cloud recognition unit to extract objects that the image recognition unit has not inferred or has misinferred, A generation unit generates object information related to the unpredicted or mispredicted objects extracted by the extraction unit, A storage unit that associates and stores object information and images generated by the generation unit, The storage unit has an adjustment unit that extracts subsets of object information and image sets stored by the storage unit that reduce the bias in the image content. The aforementioned sensor is a LiDAR, The extraction unit sets the point cloud recognition unit to not be an error if the confidence level is below a predetermined level. A sensing system characterized by the following:

10. In the sensing system according to claim 9, The sensing system is characterized in that the reliability is calculated based on one or more of the reflectance of the point cloud of the object region in the LiDAR, the number of points in the point cloud, or the distribution of the point cloud.

11. A sensing method for constructing training data for an object detection detector from images acquired from a camera, An image receiving step in which an image acquired by a camera that images the aforementioned object is received, A point cloud reception step that receives the point cloud acquired by the sensor that measures the aforementioned object, An image recognition step of inferring objects contained in the image using machine learning, A point cloud recognition step of recognizing an object included in the point cloud, An extraction step is performed to extract objects that were not inferred or misinferred in the image recognition step using the recognition results from the point cloud recognition step, The extraction step generates object information related to the uninferred or misinferred objects extracted, A saving step which involves linking and saving the object information and image generated in the generation step, The system includes an adjustment step of extracting subsets of object information and image sets from the multiple object information and image sets saved in the saving step, such that the bias in image content is reduced. The aforementioned bias in image content means that at least one of the following is biased: type of object, position of object, date and time of image capture, or weather conditions at the time of image capture. The adjustment step includes a single-map adjustment step that reduces bias within a map showing the frequency of positions in a single scene, and a multi-map adjustment step that reduces bias between multiple maps. A sensing method characterized by the following features.

12. A sensing method for constructing training data for an object detection detector from images acquired from a camera, An image receiving step in which an image acquired by a camera that images the aforementioned object is received, A point cloud reception step that receives the point cloud acquired by the sensor that measures the aforementioned object, An image recognition step of inferring objects contained in the image using machine learning, A point cloud recognition step of recognizing an object included in the point cloud, An extraction step is performed to extract objects that were not inferred or misinferred in the image recognition step using the recognition results from the point cloud recognition step, The extraction step generates object information related to the uninferred or misinferred objects extracted, A saving step which involves linking and saving the object information and image generated in the generation step, The adjustment step involves extracting a subset of the multiple object information and image sets saved in the saving step such that the bias in the image content is reduced. A memory step of storing information about the frequency of being determined to be unpredicted or mispredicted, The system includes a notification step that compares the latest frequency information with past information and notifies if it is determined that the frequency has increased. A sensing method characterized by the following features.

13. A sensing method for constructing training data for an object detection detector from images acquired from a camera, An image receiving step in which an image acquired by a camera that images the aforementioned object is received, A point cloud reception step that receives the point cloud acquired by the sensor that measures the aforementioned object, An image recognition step of inferring objects contained in the image using machine learning, A point cloud recognition step of recognizing an object included in the point cloud, An extraction step is performed to extract objects that were not inferred or misinferred in the image recognition step using the recognition results from the point cloud recognition step, The extraction step generates object information related to the uninferred or misinferred objects extracted, A saving step which involves linking and saving the object information and image generated in the generation step, The system includes an adjustment step of extracting subsets of object information and image sets from the multiple object information and image sets saved in the saving step, such that the bias in image content is reduced. The aforementioned sensor is a LiDAR, The extraction step determines that an image of an object cannot be obtained in a predetermined region of the image due to the influence of image brightness, and sets the object in the predetermined region as not being uninferred. A sensing method characterized by the following features.

14. A sensing method for constructing training data for an object detection detector from images acquired from a camera, An image receiving step in which an image acquired by a camera that images the aforementioned object is received, A point cloud reception step that receives the point cloud acquired by the sensor that measures the aforementioned object, An image recognition step of inferring objects contained in the image using machine learning, A point cloud recognition step of recognizing an object included in the point cloud, An extraction step is performed to extract objects that were not inferred or misinferred in the image recognition step using the recognition results from the point cloud recognition step, The extraction step generates object information related to the uninferred or misinferred objects extracted, A saving step which involves linking and saving the object information and image generated in the generation step, The system includes an adjustment step of extracting subsets of object information and image sets from the multiple object information and image sets saved in the saving step, such that the bias in image content is reduced. The aforementioned sensor is a LiDAR, The extraction step is defined as not being an error if the confidence level in the point cloud recognition step is below a predetermined level. A sensing method characterized by the following features.

15. In the sensing method according to claim 14, The sensing method is characterized in that the reliability is calculated based on one or more of the reflectance of the point cloud of the object region in the LiDAR, the number of points in the point cloud, or the distribution of the point cloud.