Information processing device, information processing method, and information processing system
A multi-task learning model for human presence detection using wide-angle lenses addresses accuracy issues due to varying camera positions by incorporating tasks like human distance and keypoint detection, ensuring consistent performance across different installations and reducing training complexity.
Patent Information
- Application Number
- PCT/JP2025/023603
- Authority / Receiving Office
- WO · WO
- Patent Type
- Applications
- Current Assignee / Owner
- Priority Date
- 2024-07-18
- Filing Date
- 2025-07-01
- Publication Date
- 2026-01-22
Smart Images

Figure JP2025023603_22012026_PF_FP_ABST
Abstract
Description
Information processing device, information processing method, and information processing system
[0001] The present technology relates to an information processing device, an information processing method, and an information processing system, and more particularly to an information processing device, an information processing method, and an information processing system that are capable of realizing subject detection that is robust regardless of the mounting position of a camera.
[0002] There is a technology to detect people in captured images. Such technology is used, for example, in automobile parking monitoring systems. By detecting people within a certain distance from the camera's installation position, it is possible to detect suspicious individuals approaching a parked car.
[0003] Japanese Patent Application Laid-Open No. 2020-107070
[0004] Cameras equipped with wide-angle lenses such as fisheye lenses are often used to capture images for human presence detection. Multiple cameras equipped with wide-angle lenses are attached to various positions on a single vehicle.
[0005] Images taken with a wide-angle lens tend to show people and other subjects with significant distortion. Therefore, if a human presence detection function is designed with only images taken with a camera mounted in a specific position in mind, the accuracy of human presence detection for images taken with a camera mounted in another position will be low.
[0006] The present technology has been made in view of such circumstances, and makes it possible to realize subject detection that is robust regardless of the mounting position of the camera.
[0007] An information processing device according to one aspect of the present technology includes a memory unit that stores an inference model generated by multi-task learning based on captured images for learning, to which a distance label from the shooting position to the position of the subject has been added, and the first task is presence / absence detection, which is a task of detecting a subject that is within a threshold distance range, and the second task is detection of other information related to the subject, and an inference unit that inputs the captured image into the inference model and outputs the result of the presence / absence detection based on the output of the inference model.
[0008] In one aspect of the present technology, an inference model is generated by multi-task learning based on captured images for training, to which distance labels from a shooting position to a position of the subject are added, and the first task is presence / absence detection, which is a task of detecting a subject within a distance range that is a threshold, and the second task is detection of other information related to the subject. The captured image is input to the inference model, and a result of the presence / absence detection is output based on an output of the inference model.
[0009] 1 is a diagram showing an example of an installation position of an in-vehicle camera. FIG. 2 is a diagram showing an example of human presence detection. FIG. 3 is a diagram showing an example of a wide-angle image with distance labels set. FIG. 4 is a diagram showing examples of images captured by cameras with different depression angles. FIG. 5 is a diagram showing an example of learning of an inference model in the present technology. FIG. 6 is a diagram showing an example of human distance detection. FIG. 7 is a diagram showing an example of a model configuration used for multi-task learning. FIG. 8 is a diagram showing an example of a model configuration of a human presence detection DNN. FIG. 9 is a block diagram showing an example of the configuration of an in-vehicle camera. FIG. 10 is a diagram showing an example of the chip structure of an image sensor. FIG. 11 is a diagram showing a schematic example of a configuration of an information processing system. FIG. 12 is a block diagram showing an example of the configuration of a learning device. FIG. 13 is a flowchart explaining learning processing of a human presence detection DNN. FIG. 14 is a flowchart explaining human presence detection processing. FIG. 15 is a diagram showing another example of multi-task learning. FIG. 16 is a diagram showing another example of inference using a human presence detection DNN. FIG. 17 is a block diagram showing another example of the configuration of an in-vehicle camera. FIG. 18 is a block diagram showing an example of a schematic configuration of a vehicle control system. FIG. 19 is an explanatory diagram showing an example of installation positions of an outside-vehicle information detection unit and an imaging unit.
[0010] Hereinafter, embodiments of the present technology will be described. The description will be made in the following order: 1. Human detection in an in-vehicle camera 2. Model learning in the present technology 3. Configuration of each device 4. Operation of each device 5. Modification 6. Example of application to a moving body
[0011] <Human detection in an in-vehicle camera> Figure 1 is a diagram showing an example of the installation position of an in-vehicle camera that constitutes an SVM (Surround View Monitoring) system. An example of the installation position in a plan view is shown in A of Figure 1. For example, a parking monitoring function used for crime prevention measures when parking can be realized using the SVM system.
[0012] As shown by small circular dots in A of Fig. 1, an on-board camera is mounted at position P1 on the front of the vehicle, position P2 on the rear, position P3 on the left side, and position P4 on the right side. The on-board cameras mounted at positions P1 to P4 all have the same configuration, for example. The on-board camera at position P1 is mounted so that its shooting direction is forward, and the on-board camera at position P2 is mounted so that its shooting direction is backward. The on-board camera at position P3 is mounted so that its shooting direction is forward on the left side, and the on-board camera at position P4 is mounted so that its shooting direction is forward on the right side.
[0013] The semicircular area #1 shown in color in A of Fig. 1 indicates the shooting range (angle of view) of the onboard camera at position P4. The onboard cameras installed at positions P1 to P4 are cameras equipped with wide-angle lenses such as fisheye lenses. When looking at the right side of the vehicle, the shooting range of the onboard camera at position P4 is a circular area as shown in B of Fig. 1.
[0014] During parking surveillance, each in-vehicle camera detects the presence or absence of a person based on wide-angle images captured using a wide-angle lens. The image sensor installed in the in-vehicle camera is equipped with a DNN (Deep Neural Network) that takes wide-angle images as input and outputs information indicating the probability that a person is captured in the input image.
[0015] FIG. 2 is a diagram showing an example of human presence detection using an image sensor mounted on an in-vehicle camera.
[0016] As shown on the left side of Figure 2, low-resolution wide-angle images, such as 120 x 96 pixels, captured during parking surveillance are used as input images for the human presence detection DNN, which is a DNN for inference. The human presence detection DNN outputs a value indicating the probability that a person is in the input image, and a threshold process is used to determine whether a person is present or not.
[0017] The human presence detection DNN is an inference model generated by machine learning using wide-angle images with distance labels as training data, as shown in Figure 3. The distance label indicates the distance from the camera to a person. For example, the human presence detection DNN is trained to determine that a person is present if the distance is within 3 m, and that a person is not present if the distance is beyond 3 m or if no person is present at all. The wide-angle images shown in Figure 3 each contain people at different distances. The upward arc-shaped stripes with different shades of gray indicate areas on the ground that are equidistant from the camera position.
[0018] The mounting positions of the in-vehicle cameras that make up the SVM system vary. Different mounting positions also mean different mounting heights and angles (depression angles). As the mounted lenses are wide-angle lenses, differences in the depression angles in particular have a significant impact on the size and appearance of people in the captured images.
[0019] FIG. 4 shows examples of images captured by cameras with different depression angles.
[0020] The captured images shown in Figure 4 were taken from the same shooting position but with different depression angles. From the left, the images were taken with depression angles of 0 degrees, 30 degrees, and 45 degrees. The range indicated by arrow A1 on each image is a range at a distance of 1 meter from the shooting position, and the range indicated by arrow A2 is a range at a distance of 2 meters from the shooting position. All three images show a person H standing at the same distance, between 1 meter and 2 meters.
[0021] As shown in Figure 4, the degree of distortion of the subject person and the ground changes depending on the depression angle, and the size and appearance of person H in the same position in the captured image differ significantly. This poses the following problems when performing human presence detection using the same human presence detection DNN on in-vehicle cameras installed in different positions.
[0022] (1) When training is performed using various images taken with cameras attached in different positions as training data, the distance labels attached to each image become ambiguous, leading to a deterioration in inference accuracy. In other words, when the camera's attachment position changes, the boundary position within the image between the distance of a person that should be determined as "person present" and the distance of a person that should be determined as "not present" changes, making it unclear how far the "person present" should be determined, leading to a deterioration in the accuracy of the DNN.
[0023] (2) Even for images with the same distance label, the size and appearance of people in the image change significantly depending on the mounting position. Therefore, if the training data contains many images with a certain depression angle, the inference accuracy for images with a different depression angle will deteriorate. For example, if the training data contains many images with a depression angle of 45 degrees, a model will be trained that determines "people present" only if the size of the person in the image is large, and the number of incorrect determinations will increase when images with other depression angles, such as 0 degrees, are input.
[0024] In this technology, a DNN that is robust to the mounting position is trained and used to detect the presence or absence of people in each vehicle-mounted camera, which is mounted in different positions.
[0025] <Model Learning in the Present Technology> FIG. 5 is a diagram showing an example of inference model learning in the present technology.
[0026] As shown in Figure 5, multi-task learning is used in the training of the inference model in this technology. Multi-task learning is a machine learning technique for a single model that simultaneously infers multiple tasks. By training the model so that it can accurately infer multiple related tasks, it is possible to acquire common factors between the tasks and improve the inference accuracy of the main task.
[0027] As shown in Fig. 5, human presence detection, which is a task for determining whether a person is present or absent, is the primary task (first task). In the example of Fig. 5, human distance detection, segmentation detection, and keypoint detection are each auxiliary tasks (second tasks). These auxiliary tasks are tasks related to the primary task, as they detect other information about a person appearing as a subject in a captured image.
[0028] Human distance detection is the task of detecting the distance from the shooting position to the position of a person in the input image. Segmentation detection is the task of segmenting the input image. Segmentation detects the area of each object that is the subject, including the area where the subject person is captured. Keypoint detection is the task of detecting specific parts, such as the skeleton points, of a person in the input image.
[0029] Other tasks related to human presence detection may be used as auxiliary tasks. Rather than using the three types of tasks, human distance detection, segmentation detection, and keypoint detection, as auxiliary tasks, it is possible to use at least one of these tasks as an auxiliary task.
[0030] A large number of images taken by cameras mounted in different positions are prepared as training data and input into the DNN in sequence. For example, each image is labeled with the correct answer data for each task. Correct answer data for "people present" or "people absent," a distance label that is the correct answer data for the position of the person from the shooting position, correct answer data for the area in which the person appears, and correct answer data for the position of the person's skeleton points in the input image are added to each image.
[0031] In feature extraction DNNs, features of the input image are extracted by performing operations such as convolution, and the extracted features are used for the tasks of human presence detection, human distance detection, segmentation detection, and keypoint detection.
[0032] The feature values output by the feature extraction DNN are used as input, and the human presence detection results are output from the human presence detection layer, which includes layers such as a fully connected layer.
[0033] Furthermore, the human distance detection layer outputs human distance detection results. The segmentation detection layer outputs segmentation detection results, and the keypoint detection layer outputs keypoint detection results. Each layer, including the feature extraction DNN, is trained by updating parameters based on the error between the human presence detection result and the correct data, the error between the human distance detection result and the correct data, the error between the segmentation detection result and the correct data, and the error between the keypoint detection result and the correct data.
[0034] In this way, in this technology, multi-task learning is performed in which human presence detection is the primary task and human distance detection, segmentation detection, and keypoint detection are each auxiliary tasks.
[0035] Fig. 6 is a diagram showing an example of human distance detection. In human distance detection, as shown in Fig. 6, the distance of the subject is treated as a class, and the probability for each class is output, and the class with the highest probability is processed as the human distance detection result. By using human distance detection as an auxiliary task, it is possible to avoid ambiguity in distance labels and perform learning even when captured images with different distance labels corresponding to various mounting positions are input.
[0036] Fig. 7 shows an example of a model configuration used for multi-task learning, which includes a feature extraction DNN, a layer for detecting human presence as the main task, and a layer for detecting segmentation as an auxiliary task.
[0037] The feature extraction DNN consists of layers L1 to L6. Layers L1 and L6 perform processes such as convolution and batch normalization, while layers L2 to L5 are inverted residual block layers. The input image is input to layer L1, and processing is carried out in order in each layer. The image size and number of channels input to each layer are shown in the upper left of each layer.
[0038] The output of layer L6, the final layer of the feature extraction DNN, is input to the human presence detection layer. As indicated by the dashed arrows, the outputs of layers L1 to L4 and layer L6 are input to the segmentation detection layer as features.
[0039] The human presence detection layer consists of layers L11 and L12. Layer L11 performs processing such as average pooling, and layer L12 is an output layer using a sigmoid function. Layer L11 processes the features output from the feature extraction DNN, and layer L12 outputs the human presence detection results.
[0040] The segmentation detection layer is composed of layers L21 to L26. Layers L21 to L24 perform processes such as upconversion and concat. The image size and number of channels input to layers L21 to L24 are indicated in the lower right corner of each layer. In layer L21, processing is performed based on the output of layer L6 of the feature extraction DNN and the output of layer L4. In layer L22, processing is performed based on the output of layer L21 and the output of layer L3 of the feature extraction DNN. In layer L23, processing is performed based on the output of layer L22 and the output of layer L2 of the feature extraction DNN. In layer L24, processing is performed based on the output of layer L23 and the output of layer L1 of the feature extraction DNN. The output of layer L24 is input to layer L25.
[0041] Binary classification is performed by processing layers L25 and L26, and a mask image showing areas where people appear is output from layer L26 as the segmentation detection result, as shown in the lower right of Fig. 7. The white areas in the mask image shown in the lower right of Fig. 7 correspond to areas in the input image where people appear.
[0042] A layer for human distance detection and a layer for keypoint detection are provided in the same manner as the segmentation detection layer in FIG.
[0043] Multi-task learning with human distance detection as an auxiliary task makes it possible to learn changes in the distance to people in wide-angle images caused by differences in depression angle, etc. Furthermore, multi-task learning with segmentation detection and keypoint detection as auxiliary tasks makes it possible to learn changes in human shape, such as distortions, caused by differences in depression angle, etc.
[0044] The human presence detection DNN generated by such multi-task learning is provided to the image sensor of each vehicle-mounted camera.
[0045] FIG. 8 is a diagram illustrating an example of a model configuration of a human presence detection DNN, which is a DNN for inference.
[0046] As shown in Figure 8, a model including the feature extraction DNN used for the main task and the human presence detection layer is generated as the human presence detection DNN. This enables highly accurate human presence detection for each in-vehicle camera even if the installation positions are different.
[0047] <Configuration of Each Device> Configuration of the Vehicle-Mounted Camera FIG. 9 is a block diagram showing an example of the configuration of the vehicle-mounted camera 1. As shown in FIG.
[0048] The vehicle-mounted camera 1 has a lens 11 , an image sensor 12 , a processor 13 , and a memory 14 .
[0049] The lens 11 is configured as a wide-angle lens such as a fisheye lens, a focus lens, or the like. While the projection method of a normal lens is a central projection method, the projection method of a fisheye lens is an equidistant projection method, an equisolid angle projection method, an orthogonal projection method, or the like. A projection method different from the central projection method is used for the projection method of a fisheye lens. Light collected by the lens 11 is guided to the imaging unit 21 of the image sensor 12. The lens 11 may be configured as a zoom lens.
[0050] The image sensor 12 has an imaging unit 21, a signal processing unit 22, a DSP 23, a memory 24, and a selector 25. For example, the image sensor 12 is a one-chip CMOS (Complementary Metal Oxide Semiconductor) image sensor having a stacked structure of a pixel chip 12A and a logic chip 12B, as shown in Fig. 10. For example, the imaging unit 21 is provided on the pixel chip 12A, and the signal processing unit 22, the DSP 23, the memory 24, and the selector 25 are provided on the logic chip 12B.
[0051] The imaging unit 21 has a pixel array unit in which pixels, each including a light-receiving element such as a photodiode, are arranged in a matrix. Each pixel in the pixel array unit performs photoelectric conversion of light incident on the light-receiving element and accumulates an electric charge according to the amount of light received. The imaging unit 21 also has an A / D conversion unit. The analog pixel signal of each pixel is converted to a digital value in the A / D conversion unit, and the digital image data is supplied to the signal processing unit 22 and the DSP 23.
[0052] The signal processing unit 22 performs various signal processing such as noise removal and white balance adjustment on the digital image data read out from the imaging unit 21. The image data after the signal processing by the signal processing unit 22 is supplied to a selector 25. The image data after the signal processing by the signal processing unit 22 may be supplied to the DSP 23 and used for processing in the DSP 23.
[0053] The DSP 23 reads and acquires the human presence detection DNN stored in the memory 24 and performs human presence detection using the human presence detection DNN. The memory 24 stores the human presence detection DNN generated by the multitask learning described above. The DSP 23 acquires the human presence detection result by inputting the wide-angle image provided from the image capture unit 21 to the human presence detection DNN and outputs it to the selector 25. The DSP 23 functions as an inference unit that inputs the wide-angle image to the human presence detection DNN, which is an inference model stored in the memory 24, and outputs the human presence detection result based on the output of the human presence detection DNN. The image sensor 12 is an information processing device having the memory 24 as a storage unit that stores the human presence detection DNN and the DSP 23 as an inference unit.
[0054] For example, during normal driving, the selector 25 outputs image data supplied from the signal processing unit 22 to the processor 13. Wide-angle images captured during normal driving are used in an ECU (Electronic Control Unit) to recognize the situation around the vehicle. Furthermore, during parking monitoring, the selector 25 outputs the occupancy detection result supplied from the DSP 23 to the processor 13.
[0055] The processor 13 executes programs stored in the memory 14 and controls the overall operation of the vehicle-mounted camera 1. For example, the processor 13 outputs image data supplied from the selector 25 to the ECU. Furthermore, when the occupant presence detection result supplied from the DSP 23 indicates "occupant present," the processor 13 outputs an interrupt signal to the ECU.
[0056] 11 is a diagram showing a schematic configuration example of an information processing system using the vehicle-mounted camera 1 having the above configuration. A parking monitoring function is realized by the information processing system shown in FIG.
[0057] On-board cameras 1-1 to 1-4 are mounted at positions P1 to P4, respectively. On-board cameras 1-1 to 1-4 are connected to ECU 2 via wired or wireless communication. Each of on-board cameras 1-1 to 1-4 detects the presence or absence of a person using a human presence detection DNN, and if a "person is present," an interrupt signal is output to ECU 2.
[0058] The ECU 2 is a control device that controls the overall operation of the vehicle. For example, the ECU 2 sets the vehicle's operation mode to a parking monitoring mode. When the parking monitoring mode is set, each vehicle-mounted camera 1 performs occupancy detection.
[0059] The ECU 2 is activated in response to an interrupt signal from one of the vehicle-mounted cameras 1-1 to 1-4, and executes predetermined processing such as issuing an alarm or sending a notification to the vehicle owner. Because each vehicle-mounted camera 1 is equipped with a human presence detection function, the ECU 2 can be set to a power-saving operating state during parking monitoring. Typically, the power consumption of the ECU 2 is greater than that of the vehicle-mounted camera 1. This makes it possible to reduce power consumption compared to when human presence detection is performed by the ECU 2 based on images captured by each vehicle-mounted camera 1.
[0060] 12 is a block diagram showing an example of the configuration of a learning device 101 for a human presence detection DNN. The learning device 101 is configured by a computer such as a PC.
[0061] A CPU (Central Processing Unit) 111, a ROM (Read Only Memory) 112, and a RAM (Random Access Memory) 113 are interconnected by a bus 114. In the CPU 111, a learning data acquisition unit 111A and a learning processing unit 111B are realized by executing a predetermined program.
[0062] The training data acquisition unit 111A acquires a large number of wide-angle images that constitute training data. Each wide-angle image is labeled with correct answer data for the main task and the auxiliary task. The training processing unit 111B performs multi-task learning as described with reference to FIG. 5 based on the training data acquired by the training data acquisition unit 111A, and generates a human presence / absence detection DNN. The training device 101 is an information processing device that includes the training data acquisition unit 111A as an acquisition unit that acquires captured images for training, and the training processing unit 111B as a generation unit that performs multi-task learning based on the captured images for training and generates a human presence / absence detection DNN.
[0063] An input / output interface 115 is connected to the bus 114. An input unit 116 including a keyboard, a mouse, etc., and an output unit 117 including a display, a speaker, etc. are connected to the input / output interface 115. Also connected to the input / output interface 115 are a storage unit 118 including a hard disk, a nonvolatile memory, etc., a communication unit 119 including a network interface, etc., and a drive 120 that drives removable media 121.
[0064] <Operation of Each Device> Here, the operation of each device having the above-described configuration will be described.
[0065] Operation of the Learning Device First, the learning process of the human presence detection DNN will be described with reference to the flowchart in Fig. 13. The process in Fig. 13 starts when learning data transmitted from an external device is received by, for example, the communication unit 119 and imported into the learning device 101.
[0066] In step S1, the learning data acquisition unit 111A acquires learning data.
[0067] In step S2, the learning processing unit 111B inputs the wide-angle image that constitutes the learning data into the feature extraction DNN and extracts features of the input image by performing various processes such as convolution.
[0068] In step S3, the learning processing unit 111B performs human presence detection as a primary task based on the feature amount of the input image.
[0069] In step S4, the learning processing unit 111B performs human distance detection as an auxiliary task based on the feature amount of the input image.
[0070] In step S5, the learning processing unit 111B performs segmentation detection as an auxiliary task based on the feature amount of the input image.
[0071] In step S6, the learning processing unit 111B performs keypoint detection as an auxiliary task based on the feature amount of the input image.
[0072] In step S7, the learning processing unit 111B updates parameters based on the error between the correct data and each of the detection results for human presence detection, human distance detection, segmentation detection, and keypoint detection. The above process is repeated to generate a human presence detection DNN. An image sensor 12 having the human presence detection DNN generated by the learning device 101 is manufactured, and multiple vehicle-mounted cameras 1 equipped with the image sensor 12 are attached to the vehicle.
[0073] Operation of the In-Vehicle Camera The human presence detection process of the in-vehicle camera 1 will be described with reference to the flowchart of Fig. 14. The process of Fig. 14 is performed in each in-vehicle camera 1 when, for example, the parking monitoring mode is set.
[0074] In step S11, the photographing unit 21 photographs the surroundings of the vehicle and acquires a wide-angle image.
[0075] In step S12, the DSP 23 inputs the captured wide-angle image to the human presence detection DNN and extracts features. The feature extraction is performed in the feature extraction layer that constitutes the human presence detection DNN.
[0076] In step S13, the DSP 23 performs human presence detection based on the feature amounts of the input image by inputting the feature amounts to a human presence detection layer that constitutes the human presence detection DNN.
[0077] In step S14, the processor 13 determines whether the human presence detection result is "human presence" based on the output from the DSP 23.
[0078] If it is determined in step S14 that there is an occupant present, then in step S15 the processor 13 outputs an interrupt signal to the ECU 2. After the interrupt signal is output, or if it is determined in step S14 that there is no occupant present, the process returns to step S11, and the above processing is repeated.
[0079] The above processing realizes a robust occupancy detection DNN system using wide-angle images for various installation positions. That is, when performing occupancy detection using wide-angle images with in-vehicle cameras 1 installed in different positions, it is possible to perform occupancy detection with the same accuracy using the same DNN regardless of the installation position.
[0080] Furthermore, compared to preparing a human presence detection DNN for each vehicle-mounted camera with a different installation location, this method reduces training costs. For example, preparing a human presence detection DNN for each installation location requires preparing wide-angle images with a variety of camera parameters as training data, but this method eliminates the need for such work. Furthermore, because it can learn the shape and size of people in the training images, training can be performed without preprocessing such as distortion correction, regardless of the lens used for the image capture.
[0081] Although human presence detection, which is the task of determining whether a person is present or absent, is considered to be the primary task, it is possible to use objects other than humans as the subject of presence detection by the primary task. Any object, such as an animal, a moving object, or a building, can be used as the subject of detection.
[0082] Modified Example of Learning FIG. 15 is a diagram showing another example of multitask learning.
[0083] 15 , a camera parameter inference DNN is provided after the layer that performs keypoint detection. The camera parameter inference DNN is a DNN that receives keypoint detection results as input and outputs camera parameters such as the depression angle and degree of distortion. The camera parameter inference DNN generated by learning based on the keypoint detection results and camera parameters prepared as learning data is provided in the learning device 101.
[0084] During training of the occupancy detection DNN, the keypoint detection results are input to the camera parameter inference DNN, and the inference results of the camera parameters are obtained. Furthermore, training is performed by updating the parameters based on the error between the inference results of the camera parameters and the camera parameters prepared as ground truth data. By training based on the error from the actual camera parameters, it becomes possible to train a recognizer that is robust to the mounting angle and distortion of the in-vehicle camera 1.
[0085] Modified Example of Inference FIG. 16 is a diagram showing another example of inference using a human presence detection DNN.
[0086] Although human distance detection, segmentation detection, and keypoint detection using the output of the feature extraction DNN are performed only during learning and not during inference, human distance detection, segmentation detection, and keypoint detection may also be performed during inference, with the respective inference results being used to determine the results of human presence / absence detection as the main task. For example, it is possible to change the threshold used to determine whether a "human is present" or "no human is present" depending on the reliability of each auxiliary task of human distance detection, segmentation detection, and keypoint detection.
[0087] Instead of performing all auxiliary tasks, at least one of human distance detection, segmentation detection, and keypoint detection may be performed during inference. Also, a camera parameter inference DNN may be provided, and camera parameters may be inferred based on keypoint detection results during inference. The inferred camera parameters are used for processing such as correcting images captured during parking monitoring and adjusting the angle of view when the lens 11 is a zoom lens.
[0088] Modification of the Configuration of the Vehicle-Mounted Camera 1 FIG. 17 is a block diagram showing another example of the configuration of the vehicle-mounted camera 1. In FIG.
[0089] Although inference using the human presence detection DNN is assumed to be performed in the image sensor 12, it is also possible to have it performed in the processor 13, which is an external configuration of the image sensor 12. In this case, the human presence detection DNN is prepared in the memory 14 as shown in FIG.
[0090] The processor 13 reads and acquires the human presence detection DNN stored in the memory 14, and inputs the wide-angle image captured by the image sensor 12 to the human presence detection DNN to detect the presence or absence of a person. In this way, the processor 13 functions as an inference unit that inputs the wide-angle image to the human presence detection DNN, which is an inference model stored in the memory 14, and outputs the result of human presence detection based on the output of the human presence detection DNN. In the example of Figure 17, the vehicle-mounted camera 1 is an information processing device having the memory 14 as a storage unit that stores the human presence detection DNN, and the processor 13 as an inference unit.
[0091] Inference using the human presence detection DNN is performed in the image sensor 12 or the in-vehicle camera 1, but it may also be performed in various information processing devices with wide-angle lenses, such as smartphones, tablet terminals, video cameras, HMDs, PCs, etc. Images captured using a normal lens with a central projection projection method, rather than a fisheye lens, may also be used for training the human presence detection DNN and for inference by the human presence detection DNN.
[0092] - Example of a program The above-described series of processes can be executed by hardware or software. When the series of processes is executed by software, the program that constitutes the software is installed in a computer that is built into dedicated hardware, or in a general-purpose personal computer.
[0093] The program to be installed is provided by being recorded on removable media such as an optical disk (CD-ROM (Compact Disc-Read Only Memory), DVD (Digital Versatile Disc), etc.) or semiconductor memory. It may also be provided via wired or wireless transmission media such as a local area network, the Internet, or digital broadcasting. The program can be pre-installed in a ROM or memory unit.
[0094] The program executed by the computer may be a program that processes in chronological order according to the order described in this specification, or may be a program that processes in parallel or at the required timing, such as when called.
[0095] In this specification, a system refers to a collection of multiple components (devices, modules (components), etc.), regardless of whether all of the components are housed in the same housing. Therefore, multiple devices housed in separate housings and connected via a network, and a single device housed in a single housing with multiple modules, are both systems.
[0096] The effects described in this specification are merely examples and are not limiting, and other effects may also be present.
[0097] The embodiments of the present technology are not limited to the above-described embodiments, and various modifications are possible without departing from the spirit of the present technology.
[0098] Each step described in the above flowchart can be executed by one device, or can be shared and executed by a plurality of devices.
[0099] When one step includes multiple processes, the multiple processes included in that one step can be executed by one device or can be shared and executed by multiple devices.
[0100] <Application to a Mobile Body> The technology according to the present disclosure (the present technology) can be applied to various products. For example, the technology according to the present disclosure may be realized as a device mounted on any type of mobile body, such as an automobile, an electric vehicle, a hybrid electric vehicle, a motorcycle, a bicycle, personal mobility, an airplane, a drone, a ship, or a robot.
[0101] FIG. 18 is a block diagram showing a schematic configuration example of a vehicle control system, which is an example of a mobile object control system to which the technology according to the present disclosure can be applied.
[0102] The vehicle control system 12000 includes a plurality of electronic control units connected via a communication network 12001. In the example shown in Fig. 18, the vehicle control system 12000 includes a drive system control unit 12010, a body system control unit 12020, an outside-vehicle information detection unit 12030, an inside-vehicle information detection unit 12040, and an integrated control unit 12050. Also shown as functional components of the integrated control unit 12050 are a microcomputer 12051, an audio / video output unit 12052, and an in-vehicle network I / F (Interface) 12053.
[0103] The drivetrain control unit 12010 controls the operation of devices related to the drivetrain of the vehicle in accordance with various programs. For example, the drivetrain control unit 12010 functions as a control device for a drive force generating device for generating a drive force of the vehicle, such as an internal combustion engine or a drive motor, a drive force transmission mechanism for transmitting the drive force to the wheels, a steering mechanism for adjusting the steering angle of the vehicle, and a braking device for generating a braking force of the vehicle.
[0104] The body system control unit 12020 controls the operation of various devices equipped in the vehicle body according to various programs. For example, the body system control unit 12020 functions as a control device for a keyless entry system, a smart key system, a power window device, or various lamps such as headlamps, backup lamps, brake lamps, turn signals, and fog lamps. In this case, radio waves transmitted from a portable device that serves as a key or signals from various switches can be input to the body system control unit 12020. The body system control unit 12020 receives these radio waves or signals and controls the vehicle's door lock device, power window device, lamps, etc.
[0105] The outside-vehicle information detection unit 12030 detects information outside the vehicle equipped with the vehicle control system 12000. For example, an imaging unit 12031 is connected to the outside-vehicle information detection unit 12030. The outside-vehicle information detection unit 12030 causes the imaging unit 12031 to capture images outside the vehicle and receives the captured images. The outside-vehicle information detection unit 12030 may perform object detection processing or distance detection processing for people, cars, obstacles, signs, characters on the road surface, etc. based on the received images.
[0106] The imaging unit 12031 is an optical sensor that receives light and outputs an electrical signal corresponding to the amount of light received. The imaging unit 12031 can output the electrical signal as an image or as distance measurement information. The light received by the imaging unit 12031 may be visible light or invisible light such as infrared light.
[0107] The in-vehicle information detection unit 12040 detects information inside the vehicle. For example, a driver state detection unit 12041 that detects the state of the driver is connected to the in-vehicle information detection unit 12040. The driver state detection unit 12041 includes, for example, a camera that captures an image of the driver, and the in-vehicle information detection unit 12040 may calculate the degree of fatigue or concentration of the driver based on the detection information input from the driver state detection unit 12041, or may determine whether the driver is dozing off.
[0108] The microcomputer 12051 can calculate control target values for the driving force generating device, steering mechanism, or braking device based on the information inside and outside the vehicle acquired by the outside-vehicle information detection unit 12030 or the inside-vehicle information detection unit 12040, and output control commands to the drive system control unit 12010. For example, the microcomputer 12051 can perform cooperative control aimed at realizing the functions of an ADAS (Advanced Driver Assistance System), including vehicle collision avoidance or impact mitigation, following driving based on the distance between vehicles, maintaining vehicle speed, vehicle collision warning, vehicle lane departure warning, etc.
[0109] In addition, the microcomputer 12051 can perform cooperative control for the purpose of autonomous driving, which allows the vehicle to travel autonomously without relying on driver operation, by controlling the driving force generating device, steering mechanism, braking device, etc. based on information about the surroundings of the vehicle obtained by the outside vehicle information detection unit 12030 or the inside vehicle information detection unit 12040.
[0110] Furthermore, the microcomputer 12051 can output a control command to the body system control unit 12020 based on the information outside the vehicle acquired by the outside information detection unit 12030. For example, the microcomputer 12051 can control the headlamps according to the position of a preceding vehicle or an oncoming vehicle detected by the outside information detection unit 12030, and perform cooperative control aimed at preventing glare, such as switching from high beams to low beams.
[0111] The audio / video output unit 12052 transmits at least one of audio and video output signals to an output device capable of visually or audibly notifying the passengers of the vehicle or the outside of the vehicle of information. In the example of Fig. 18, the output devices are exemplified by an audio speaker 12061, a display unit 12062, and an instrument panel 12063. The display unit 12062 may include, for example, at least one of an on-board display and a head-up display.
[0112] FIG. 19 is a diagram showing an example of the installation position of the imaging unit 12031.
[0113] In FIG. 19, the imaging unit 12031 includes imaging units 12101, 12102, 12103, 12104, and 12105.
[0114] The imaging units 12101, 12102, 12103, 12104, and 12105 are provided, for example, at positions such as the front nose, side mirrors, rear bumper, back door, and the top of the windshield inside the vehicle cabin of the vehicle 12100. The imaging unit 12101 provided on the front nose and the imaging unit 12105 provided on the top of the windshield inside the vehicle cabin mainly acquire images of the front of the vehicle 12100. The imaging units 12102 and 12103 provided on the side mirrors mainly acquire images of the sides of the vehicle 12100. The imaging unit 12104 provided on the rear bumper or back door mainly acquires images of the rear of the vehicle 12100. The imaging unit 12105 provided on the top of the windshield inside the vehicle cabin is mainly used to detect preceding vehicles, pedestrians, obstacles, traffic lights, traffic signs, lanes, etc.
[0115] 19 shows an example of the imaging ranges of the imaging units 12101 to 12104. Imaging range 12111 indicates the imaging range of the imaging unit 12101 provided on the front nose, imaging ranges 12112 and 12113 indicate the imaging ranges of the imaging units 12102 and 12103 provided on the side mirrors, respectively, and imaging range 12114 indicates the imaging range of the imaging unit 12104 provided on the rear bumper or back door. For example, by overlaying the image data captured by the imaging units 12101 to 12104, an overhead image of the vehicle 12100 viewed from above can be obtained.
[0116] At least one of the image capturing units 12101 to 12104 may have a function of acquiring distance information. For example, at least one of the image capturing units 12101 to 12104 may be a stereo camera made up of multiple image capturing elements, or may be an image capturing element having pixels for phase difference detection.
[0117] For example, based on the distance information obtained from the imaging units 12101 to 12104, the microcomputer 12051 can calculate the distance to each three-dimensional object within the imaging ranges 12111 to 12114 and the change in this distance over time (relative speed with respect to the vehicle 12100), thereby extracting as a preceding vehicle, in particular, the three-dimensional object that is the closest three-dimensional object on the path of the vehicle 12100 and traveling in approximately the same direction as the vehicle 12100 at a predetermined speed (e.g., 0 km / h or higher). Furthermore, the microcomputer 12051 can set a vehicle-to-vehicle distance to be maintained in advance in front of the preceding vehicle, and perform automatic braking control (including follow-up stop control), automatic acceleration control (including follow-up start control), etc. In this way, cooperative control can be performed for the purpose of autonomous driving, which allows the vehicle to travel autonomously without relying on driver operation.
[0118] For example, the microcomputer 12051 classifies and extracts three-dimensional object data regarding three-dimensional objects into two-wheeled vehicles, ordinary vehicles, large vehicles, pedestrians, utility poles, and other three-dimensional objects based on distance information obtained from the imaging units 12101 to 12104, and can use the data for automatic obstacle avoidance. For example, the microcomputer 12051 distinguishes obstacles around the vehicle 12100 into obstacles that are visible to the driver of the vehicle 12100 and obstacles that are difficult to see. The microcomputer 12051 then determines a collision risk that indicates the risk of collision with each obstacle, and when the collision risk is equal to or greater than a set value and a collision is possible, the microcomputer 12051 can provide driving assistance for collision avoidance by outputting an alarm to the driver via the audio speaker 12061 or the display unit 12062, or by performing forced deceleration or avoidance steering via the drive system control unit 12010.
[0119] At least one of the image capturing units 12101 to 12104 may be an infrared camera that detects infrared rays. For example, the microcomputer 12051 can recognize a pedestrian by determining whether a pedestrian is present in the images captured by the image capturing units 12101 to 12104. Such pedestrian recognition is performed, for example, by extracting feature points from the images captured by the image capturing units 12101 to 12104 as infrared cameras and performing pattern matching on a series of feature points that indicate the outline of an object to determine whether the object is a pedestrian. When the microcomputer 12051 determines that a pedestrian is present in the images captured by the image capturing units 12101 to 12104 and recognizes the pedestrian, the audio / image output unit 12052 controls the display unit 12062 to superimpose a rectangular outline on the recognized pedestrian for emphasis. The audio / image output unit 12052 may also control the display unit 12062 to display an icon or the like indicating the pedestrian at a desired position.
[0120] The above describes an example of a vehicle control system to which the technology according to the present disclosure can be applied. The technology according to the present disclosure can be applied to the imaging unit 12031 and the like among the above-described configurations. Specifically, the technology can be applied when the imaging unit 12031 detects the presence or absence of a person during parking monitoring.
[0121] <Examples of Combinations of Configurations> The present technology can also have the following configurations.
[0122] (1) An information processing device comprising: a memory unit that stores an inference model generated by multi-task learning based on captured images for learning to which a distance label from a shooting position to a position of the subject has been added, the first task being presence / absence detection, which is a task of detecting an object captured within a distance range that is a threshold, and a second task being detection of other information related to the object; and an inference unit that inputs the captured images to the inference model and outputs a result of the presence / absence detection based on an output of the inference model. (2) The information processing device described in (1), wherein the second task is at least one of detecting the distance from the shooting position to the object, detecting an area in which the object is captured, and detecting a specific part of the object. (3) The information processing device described in (1) or (2), wherein the information processing device is an image sensor further comprising a capturing unit that captures images that are input to the inference model. (4) The information processing device described in (1) or (2), wherein the information processing device is an image sensor further comprising an image sensor that captures images that are input to the inference model, and wherein the memory unit and the inference unit are external to the image sensor. (5) The subject to be detected for presence or absence is a person. (6) An information processing method, comprising: an information processing device having a memory unit that stores an inference model generated by multi-task learning based on training images to which a distance label from a shooting position to a position of the subject is added, the first task being presence or absence detection, which is a task of detecting a subject that appears within a distance range that is a threshold, and a second task being detection of other information related to the subject, the information processing device inputs the captured images to the inference model and outputs a result of the presence or absence detection based on an output of the inference model. (7) An information processing device comprising: an acquisition unit that acquires training images to which a distance label from a shooting position to a position of the subject is added, and a learning unit that performs multi-task learning, based on the training images, to generate an inference model, the first task being presence or absence detection, which is a task of detecting a subject that appears within a distance range that is a threshold, and a second task being detection of other information related to the subject.(8) The information processing device according to (7), wherein the second task is at least one of detecting a distance from a shooting position to the subject, detecting an area in which the subject is photographed, and detecting a specific part of the subject. (9) The information processing device according to (7) or (8), wherein the subject for which the presence / absence detection is performed is a person. (10) An information processing method, wherein an information processing device acquires captured images for learning to which a distance label from the shooting position to a position of the subject is added, and performs multi-task learning based on the captured images for learning, with presence / absence detection, which is a task of detecting the subject photographed within a distance range that is a threshold, as a first task, and detecting other information related to the subject as a second task, to generate an inference model. (11) An information processing system having: a storage unit that stores an inference model generated by multi-task learning based on captured images for learning to which a distance label from a capture position to a position of the subject has been added, the first task being presence detection, which is a task of detecting an object captured within a distance range that is a threshold, and a second task being detection of other information related to the object; an image capture unit that captures images to be input to the inference model; and an inference unit that inputs the captured images to the inference model and outputs a result of the presence detection based on an output of the inference model; and a control device that has a control unit that executes predetermined processing based on information supplied from one of the multiple image capture devices that has detected the object within a distance range that is a threshold. (12) The multiple image capture devices are attached to different positions on a vehicle, and the presence detection in each of the image capture devices is performed based on an image captured during parking surveillance.
[0123] REFERENCE SIGNS LIST 1 on-board camera, 11 lens, 12 image sensor, 13 processor, 14 memory, 101 learning device, 111A learning data acquisition unit, 111B learning processing unit
Claims
1. An information processing device comprising: a memory unit that stores an inference model generated by multi-task learning based on learning images to which distance labels from the shooting position to the position of the subject are added, where the first task is presence / absence detection, which is the task of detecting an object within a threshold distance range, and the second task is detection of other information related to the object; and an inference unit that inputs the captured image into the inference model and outputs the result of the presence / absence detection based on the output of the inference model.
2. The information processing device according to claim 1, wherein the second task is at least one of: detecting the distance from the shooting position to the subject; detecting the area in which the subject is captured; and detecting a specific part of the subject.
3. The information processing device described in claim 1, further comprising an image sensor that captures images that serve as input to the inference model.
4. An information processing device as described in claim 1, further comprising an image sensor that captures images that serve as input for the inference model, and wherein the memory unit and the inference unit are located outside the image sensor.
5. The information processing device according to claim 1, wherein the subject for which the presence or absence detection is performed is a person.
6. An information processing method in which an information processing device has a memory unit that stores an inference model generated by multi-task learning based on captured images for learning, to which distance labels from the shooting position to the position of the subject are added, where the first task is presence / absence detection, which is a task of detecting an object within a threshold distance range, and the second task is detection of other information related to the object, inputs the captured image into the inference model and outputs the result of the presence / absence detection based on the output of the inference model.
7. An information processing device comprising: an acquisition unit that acquires captured images for learning to which distance labels from the shooting position to the position of the subject have been added; and a learning unit that performs multi-task learning based on the captured images for learning, with a first task being presence / absence detection, which is a task of detecting the subject that is captured within a threshold distance range, and a second task being detection of other information related to the subject, and generates an inference model.
8. The information processing device according to claim 7, wherein the second task is at least one of: detecting the distance from the shooting position to the subject; detecting the area in which the subject is captured; and detecting a specific part of the subject.
9. The information processing device according to claim 7, wherein the subject for which the presence or absence detection is performed is a person.
10. An information processing method in which an information processing device acquires learning images to which distance labels from the shooting position to the position of the subject have been added, and performs multi-task learning based on the learning images, with a first task being presence / absence detection, which is a task of detecting the subject that is captured within a threshold distance range, and a second task being detection of other information related to the subject, to generate an inference model.
11. An information processing system having: a memory unit that stores an inference model generated by multi-task learning based on learning images to which distance labels from the shooting position to the position of the subject have been added, where the first task is presence detection, which is a task of detecting an object within a threshold distance range, and the second task is detection of other information related to the object; an imaging device that includes: a shooting unit that captures images that serve as input to the inference model; and an inference unit that inputs the captured images to the inference model and outputs the results of the presence / absence detection based on the output of the inference model; and a control device that includes a control unit that executes predetermined processing based on information supplied from one of the imaging devices that detects the object within a threshold distance range, among multiple imaging devices.
12. The information processing system according to claim 11, wherein the plurality of image capturing devices are attached to different positions on the vehicle, and the presence / absence detection in each of the image capturing devices is performed based on an image captured during parking monitoring.
Citation Information
Patent Citations
System and method to detect fraudulent / shoplifting activities during self-checkout operations
EP4401025A1
Detection device, learning device, detection method, learning method and program
JP2015005237A
Image processing program and image processing device
JP2021051530A
Image processing device and image processing method
WO2021090943A1