Information processing device, information processing method, and information processing program
By integrating an IPU for high-resolution image processing and a MoPU for high-frame-rate point information generation, the system addresses frame rate challenges in advanced autonomous driving, achieving precise object recognition and movement analysis for improved autonomous driving accuracy.
Patent Information
- Authority / Receiving Office
- JP · JP
- Patent Type
- Patents
- Current Assignee / Owner
- SOFTBANK GROUP CORP
- Filing Date
- 2023-03-09
- Publication Date
- 2026-04-27
AI Technical Summary
Existing automatic driving systems face challenges in setting optimal frame rates for multiple cameras, which affects the accuracy of object recognition and movement analysis, particularly in advanced autonomous driving levels like Level 6.
The system employs two processors: an IPU for high-resolution image processing and a MoPU for high-frame-rate point information generation, with the MoPU capturing object movement at 100 frames per second and the IPU identifying objects, allowing for efficient data reduction and accurate object recognition.
This configuration enables precise object recognition and movement analysis with significantly reduced data output, enhancing the accuracy of autonomous driving by combining high-resolution image information with high-frame-rate point information, thereby improving obstacle avoidance and overall driving control.
Smart Images

Figure 0007851876000003 
Figure 0007851876000004 
Figure 0007851876000005
Abstract
Description
Technical Field
[0001] This disclosure relates to an information processing apparatus, an information processing method, and an information processing program.
Background Art
[0002] Patent Document 1 describes a vehicle having an automatic driving function.
Prior Art Documents
Patent Documents
[0003]
Patent Document 1
Summary of the Invention
Problems to be Solved by the Invention
[0004] By the way, when automatically driving a vehicle as in Patent Document 1, automatic driving control is performed using a plurality of images obtained by photographing the surroundings of the vehicle with a camera. Here, when controlling automatic driving using a plurality of images obtained by a plurality of cameras, there is room for improvement in how to set the frame rate for each camera.
[0005] Therefore, an object of the present disclosure is to provide an information processing apparatus, an information processing method, and an information processing program that can cause an object to be photographed at a frame rate suitable for each camera when the object is photographed by a plurality of cameras.
Means for Solving the Problems
[0006] The information processing device of this disclosure includes: a first processor that outputs point information from an image of an object captured by a first camera, capturing the captured object as a point; a second processor that outputs identification information from an image of the object captured by a second camera facing a direction corresponding to the first camera, identifying the captured object; and a third processor that associates the point information output from the first processor with the identification information output from the second processor, wherein the frame rate of the first camera is greater than the frame rate of the second camera.
[0007] Furthermore, in the information processing device of this disclosure, the frame rate of the first camera is 10 times or more the frame rate of the second camera.
[0008] Furthermore, the information processing device of the present disclosure has a frame rate of 100 frames / second or more for the first camera and a frame rate of 10 frames / second for the second camera.
[0009] The information processing method of this disclosure involves a computer performing the following steps: outputting point information from an image of an object captured by a first camera, identifying the captured object as a point from an image of the object captured by a second camera having a lower frame rate than the first camera and facing a direction corresponding to the first camera, and associating the point information with the identification information.
[0010] The information processing program of this disclosure causes a computer to output point information from an image of an object captured by a first camera, which captures the captured object as a point; output identification information from an image of the object captured by a second camera, which has a lower frame rate than the first camera and is oriented in a direction corresponding to the first camera, which identifies the captured object; and perform a process to associate the point information and the identification information.
[0011] It should be noted that the above summary of the invention does not list all the necessary features of the present invention. Furthermore, subcombinations of these features may also constitute an invention. [Brief explanation of the drawing]
[0012] [Figure 1] This is a schematic diagram showing an example of a vehicle equipped with Central Brain. [Figure 2] This is the first block diagram showing an example of the configuration of an information processing device. [Figure 3] This is a second block diagram showing an example of the configuration of an information processing device. [Figure 4] This is an explanatory diagram showing an example of point information output by the MoPU. [Figure 5] This is a third block diagram showing an example of the configuration of an information processing device. [Figure 6] This is a fourth block diagram showing an example of the configuration of an information processing device. [Figure 7] This is an explanatory diagram illustrating an example of the correspondence between point information and label information. [Figure 8] This is an explanatory diagram showing the general configuration of the vehicle. [Figure 9] This is a block diagram showing an example of the functional configuration of a cooling execution device. [Figure 10] This is the fifth block diagram showing an example of the configuration of an information processing device. [Figure 11] This is the sixth block diagram showing an example of the configuration of an information processing device. [Figure 12] This diagram schematically illustrates the detection of coordinates in a time series of an object. [Figure 13] This is the seventh block diagram showing an example of the configuration of an information processing device. [Figure 14] This is an explanatory diagram for describing images of objects captured by an event camera. [Figure 15] This is a schematic diagram illustrating an example of a computer hardware configuration that functions as an information processing device or a cooling execution device. [Modes for carrying out the invention]
[0013] Hereinafter, embodiments of the present disclosure will be described. However, the following embodiments do not limit the invention according to the claims. Also, not all combinations of features described in the embodiments are essential for the solution of the invention.
[0014] (First Embodiment) First, the first embodiment according to this embodiment will be described. The information processing apparatus according to the present disclosure is, for example, at least partially mounted on the vehicle 100 and performs automatic driving control of the vehicle 100. Further, the information processing apparatus can be realized in real time based on data obtained by AI / multivariate analysis / goal seeking / strategy formulation / optimal probability solution / optimal speed solution / optimal course management / multiple sensor inputs at the edge in Autonomous Driving at Level 6, and can provide a driving system adjusted based on the delta optimal solution. The vehicle 100 is an example of an "object".
[0015] Here, "Level 6" is a level representing autonomous driving and corresponds to a level higher than Level 5 representing fully autonomous driving. Although Level 5 represents fully autonomous driving, it is at a level equivalent to human driving, and there is still a probability of accidents and the like occurring. Level 6 represents a level higher than Level 5 and corresponds to a level with a lower probability of accidents than Level 5.
[0016] The computing power at Level 6 is about 1000 times that of Level 5. Therefore, high-performance driving control that could not be achieved at Level 5 can be realized.
[0017] Figure 1 is a schematic diagram showing an example of a vehicle 100 equipped with Central Brain 15. Multiple Gateways are connected to Central Brain 15 in a communicative manner. Central Brain 15 is connected to an external cloud server via Gateways. Central Brain 15 is configured to be able to access the external cloud server via Gateways. On the other hand, the presence of Gateways prevents direct access to Central Brain 15 from the outside.
[0018] The Central Brain 15 outputs a request signal to the cloud server at predetermined intervals. Specifically, the Central Brain 15 outputs a request signal representing an inquiry to the cloud server every one billionth of a second. As an example, the Central Brain 15 controls the autonomous driving of Level 6 based on multiple pieces of information acquired via the Gateway.
[0019] Figure 2 is a first block diagram showing an example of the configuration of the information processing device 10. The information processing device 10 comprises an IPU (Image Processing Unit) 11, a MoPU (Motion Processing Unit) 12, a Central Brain 15, and memory 16. The Central Brain 15 is composed of a GNPU (Graphics Neural Network Processing Unit) 13 and a CPU (Central Processing Unit) 14.
[0020] The IPU 11 is built into an ultra-high resolution camera (not shown) installed on the vehicle 100. The IPU 11 performs predetermined image processing, such as Bayer transform, demosaicing, noise reduction, and sharpening, on images of objects surrounding the vehicle 100 captured by the ultra-high resolution camera, and outputs the processed images of the objects at a frame rate of, for example, 10 frames per second and a resolution of 12 million pixels. The IPU 11 also outputs identification information that identifies the captured objects from the images of the objects captured by the ultra-high resolution camera. Identification information is necessary to identify what the captured object is (for example, whether it is a person or an obstacle). In this embodiment, the IPU 11 outputs label information indicating the type of the captured object (for example, information indicating whether the captured object is a dog, a cat, or a bear) as identification information. Furthermore, the IPU 11 outputs position information indicating the position of the captured object in the camera coordinate system of the ultra-high resolution camera. The images, label information, and position information output from the IPU 11 are supplied to the Central Brain 15 and the memory 16. The IPU11 is an example of a "second processor," and the ultra-high-resolution camera is an example of a "second camera."
[0021] The MoPU 12 is built into a separate camera (not shown) from the ultra-high resolution camera installed on the vehicle 100. The MoPU 12 outputs point information, for example, at a frame rate of 100 frames per second or more, from images of objects captured by the separate camera facing the same direction as the ultra-high resolution camera at a frame rate of 100 frames per second or more. The point information output from the MoPU 12 is supplied to the Central Brain 15 and memory 16. Thus, the image used by the MoPU 12 to output point information and the image used by the IPU 11 to output identification information are images captured by the separate camera and the ultra-high resolution camera facing the same direction. Here, "corresponding direction" refers to the direction in which the shooting range of the separate camera and the shooting range of the ultra-high resolution camera overlap. In the above case, the separate camera captures the object facing the direction that overlaps with the shooting range of the ultra-high resolution camera. Note that capturing an object with the ultra-high resolution camera and the separate camera facing the same direction can be achieved, for example, by pre-determining the correspondence between the camera coordinate systems of the ultra-high resolution camera and the separate camera.
[0022] For example, MoPU12 outputs coordinate values for at least two coordinate axes in a three-dimensional Cartesian coordinate system of a point indicating the location of an object, as point information. These coordinate values, as an example, represent the center point (or centroid) of the object. MoPU12 also outputs coordinate values for two coordinate axes: the coordinate value of the axis along the width direction (x-axis) in the three-dimensional Cartesian coordinate system (hereinafter referred to as "x-coordinate value") and the coordinate value of the axis along the height direction (y-axis) (hereinafter referred to as "y-coordinate value"). The x-axis is the axis along the width direction of the vehicle 100, and the y-axis is the axis along the height direction of the vehicle 100.
[0023] With the above configuration, the point information output by MoPU12 for one second includes x and y coordinate values for more than 100 frames. Based on this point information, it is possible to understand the movement (direction of movement and speed of movement) of an object on the x and y axes in the three-dimensional Cartesian coordinate system. In other words, the point information output by MoPU12 includes position information indicating the position of the object in the three-dimensional Cartesian coordinate system and motion information indicating the movement of the object.
[0024] As described above, the point information output from MoPU12 does not contain any information necessary to identify what the captured object is (for example, whether it is a person or an obstacle), but only information indicating the movement (direction of movement and speed of movement) of the object's center point (or center of gravity) along the x and y axes. Furthermore, since the point information output from MoPU12 does not contain image information, the amount of data output to the Central Brain15 and memory16 can be drastically reduced. MoPU12 is an example of a "first processor," and the other camera is an example of a "first camera."
[0025] As described above, in this embodiment, the frame rate of the separate camera with the built-in MoPU12 is greater than the frame rate of the ultra-high resolution camera with the built-in IPU11. Specifically, the frame rate of the separate camera is 100 frames / second or more, while the frame rate of the ultra-high resolution camera is 10 frames / second. In other words, the frame rate of the separate camera is more than 10 times that of the ultra-high resolution camera.
[0026] The Central Brain 15 associates the point information output from the MoPU 12 with the label information output from the IPU 11. For example, due to the difference in frame rates between the other camera and the ultra-high resolution camera, the Central Brain 15 may acquire point information about an object but not label information. In this state, the Central Brain 15 recognizes the x and y coordinates of the object based on the point information, but does not recognize what the object is.
[0027] Subsequently, if label information for the object is obtained, Central Brain 15 derives the type of label information (e.g., PERSON). Then, Central Brain 15 associates this label information with the point information obtained above. As a result, Central Brain 15 recognizes the x and y coordinates of the object based on the point information, and also recognizes what the object is. Central Brain 15 is an example of a "third processor".
[0028] Here, if there are multiple objects captured by the ultra-high resolution camera and another camera, for example, object A and object B, Central Brain 15 associates point information and label information for each object as follows. Due to the frame rate difference between the other camera and the ultra-high resolution camera, there are situations where Central Brain 15 acquires point information for object A and object B (hereinafter referred to as "point information A" and "point information B") but does not acquire label information. In this state, Central Brain 15 recognizes the x and y coordinate values of object A based on point information A, and recognizes the x and y coordinate values of object B based on point information B, but does not recognize what those objects are.
[0029] Subsequently, when a label is obtained, the Central Brain 15 derives the type of that label (e.g., PERSON). Then, based on the location information output from the IPU 11 along with the label, and the location information contained in the acquired point information A and point information B, the Central Brain 15 identifies the point information to associate with the label. For example, the Central Brain 15 identifies the point information containing location information that is closest to the location of the object indicated by the location information output from the IPU 11, and associates that point information with the label. If the point information identified above is point information A, the Central Brain 15 associates the label with point information A, recognizes the x and y coordinate values of object A based on point information A, and recognizes what object A is.
[0030] As explained above, when there are multiple objects captured by the ultra-high resolution camera and other cameras, Central Brain 15 associates point information and label information based on the position information output from IPU 11 and the position information included in the point information output from MoPU 12.
[0031] Furthermore, the Central Brain 15 recognizes objects (people, animals, roads, traffic lights, signs, crosswalks, obstacles, buildings, etc.) present around the vehicle 100 based on the image and label information output from the IPU 11. The Central Brain 15 also recognizes the position and movement of the recognized objects present around the vehicle 100 based on the point information output from the MoPU 12. Based on the recognized information, the Central Brain 15 controls the autonomous driving of the vehicle 100, for example, by controlling the motors that drive the wheels (speed control), brake control, and steering control. For example, the Central Brain 15 controls the autonomous driving of the vehicle 100 to avoid collisions with objects based on the position and movement information contained in the point information output from the MoPU 12. In the Central Brain 15, the GNPU 13 may be responsible for image recognition processing, and the CPU 14 may be responsible for vehicle control processing.
[0032] In general, ultra-high resolution cameras are used for image recognition in autonomous driving. While it's possible to recognize what an object is from an image captured by an ultra-high resolution camera, this alone is insufficient for Level 6 autonomous driving. Level 6 requires more accurate recognition of object movement. By using MoPU12 to recognize object movement with greater precision, for example, a vehicle 100 operating autonomously can perform obstacle avoidance maneuvers with greater accuracy. However, ultra-high resolution cameras can only acquire about 10 frames per second, resulting in lower accuracy in analyzing object movement compared to cameras equipped with MoPU12. On the other hand, cameras equipped with MoPU12 can output at high frame rates, such as 100 frames per second.
[0033] Therefore, the information processing device 10 according to the first embodiment includes two independent processors, an IPU 11 and a MoPU 12. The information processing device 10 assigns the IPU 11, which is built into the ultra-high resolution camera, the role of acquiring information necessary to identify what the captured object is, and the MoPU 12, which is built into a separate camera, the role of detecting the position and movement of the object. The MoPU 12 captures the captured object as a point and analyzes in which direction and at what speed the coordinates of that point move on at least the x and y axes in the three-dimensional Cartesian coordinate system. Since the overall contour of the object and the detection of what the object is can be done from the image from the ultra-high resolution camera, the MoPU 12 can, for example, determine how the center point of the object moves, and thus determine how the entire object behaves.
[0034] By analyzing only the movement and velocity of the object's central point, it is possible to significantly reduce the amount of data output to Central Brain15 and the computational load on Central Brain15 compared to determining how the entire image of the object moves. For example, when outputting a 1000x1000 pixel image to Central Brain15 at a frame rate of 1000 frames / second, including color information, 4 billion bits / second of data would be output to Central Brain15. By having MoPU12 output only point information indicating the movement of the object's central point, the amount of data output to Central Brain15 can be compressed to 20,000 bits / second. In other words, the amount of data output to Central Brain15 is compressed to 1 / 200,000th.
[0035] In this way, by combining the low-frame-rate, high-resolution image and label information output from IPU11 with the high-frame-rate, lightweight point information output from MoPU12, it becomes possible to perform object recognition, including object motion, with a small amount of data.
[0036] Furthermore, the information processing device 10 can understand information about what kind of object is moving and how, by associating the point information output from the MoPU 12 with the label information output from the IPU 11.
[0037] (Second embodiment) Next, a second embodiment according to this embodiment will be described, omitting or simplifying any parts that overlap with the above embodiment.
[0038] Figure 3 is a second block diagram showing an example of the configuration of the information processing device 10. As shown in Figure 3, the information processing device 10 mounted on the vehicle 100 includes a MoPU 12L corresponding to the left eye, a MoPU 12R corresponding to the right eye, an IPU 11, and a Central Brain 15.
[0039] The MoPU12L is equipped with a camera 30L, a radar 32L, an infrared camera 34L, and a core 17L. The MoPU12R is equipped with a camera 30R, a radar 32R, an infrared camera 34R, and a core 17R. In the following, if MoPU12L and MoPU12R are not distinguished, they will be referred to as "MoPU12", if camera 30L and camera 30R are not distinguished, they will be referred to as "camera 30", if radar 32L and radar 32R are not distinguished, they will be referred to as "radar 32", if infrared camera 34L and infrared camera 34R are not distinguished, they will be referred to as "infrared camera 34", and if core 17L and core 17R are not distinguished, they will be referred to as "core 17".
[0040] The camera 30 on MoPU12 captures objects at a higher frame rate (120, 240, 480, 960, or 1920 frames / second) than the ultra-high resolution camera on IPU11 (e.g., 10 frames / second). The frame rate of camera 30 is variable. Camera 30 is an example of a "first camera".
[0041] The radar 32 equipped in MoPU12 acquires radar signals, which are signals based on the reflected waves from electromagnetic waves irradiated onto an object. The infrared camera 34 equipped in MoPU12 is a camera that takes infrared images.
[0042] The core 17 of the MoPU12 (for example, composed of one or more CPUs) extracts feature points for each frame of image captured by the camera 30 and outputs the x and y coordinate values of the object in the three-dimensional Cartesian coordinate system as point information. For example, the core 17 uses the center point (centroid) of the object extracted from the image as a feature point. The point information output by the core 17 includes position information and motion information, as in the above embodiment.
[0043] IPU11 is equipped with an ultra-high resolution camera (not shown) and outputs an image of an object captured by the ultra-high resolution camera, label information indicating the type of object, and position information indicating the position of the object in the camera coordinate system of the ultra-high resolution camera.
[0044] The Central Brain 15 acquires point information output from the MoPU 12, and images, label information, and position information output from the IPU 11. The Central Brain 15 then associates the label information of an object located at the position corresponding to the position information contained in the point information output from the MoPU 12 and the position information output from the IPU 11 with the point information. This makes it possible for the information processing device 10 to associate information about what the object indicated by the label information is with the position and movement of the object indicated by the point information.
[0045] Here, the MoPU 12 changes the frame rate of the camera 30 according to predetermined factors. In this embodiment, the MoPU 12 changes the frame rate of the camera 30 according to a score related to the external environment, as an example of predetermined factors. In this case, the MoPU 12 calculates a score related to the external environment for the vehicle 100 and changes the frame rate of the camera 30 according to the calculated score. The MoPU 12 then outputs a control signal to the camera 30 to capture an image at the changed frame rate. As a result, the camera 30 captures an image at the frame rate indicated by the control signal. With this configuration, the information processing device 10 can capture an image of an object at a frame rate suitable for the external environment.
[0046] The information processing device 10 mounted on the vehicle 100 is equipped with multiple types of sensors (not shown). Based on sensor information taken in from multiple types of sensors (for example, weight shift, road material detection, outside temperature detection, outside humidity detection, up / down / side / diagonal inclination angle detection, road freezing condition, moisture content detection, tire material, wear condition, air pressure detection, road width, whether overtaking is prohibited, oncoming vehicles, vehicle type information of vehicles in front and behind, cruising status of those vehicles, or surrounding conditions (birds, animals, soccer balls, accident vehicles, earthquakes, fires, wind, typhoons, heavy rain, light rain, blizzards, fog, etc.)) and point information, the MoPU 12 calculates the degree of risk related to the movement of the vehicle 100 as a score related to the external environment for the vehicle 100. The degree of risk indicates the degree to which the vehicle 100 will travel through dangerous areas in the future. In this case, the MoPU 12 changes the frame rate of the camera 30 according to the calculated degree of risk. The vehicle 100 is an example of a "moving object". With this configuration, the information processing device 10 can change the frame rate of the camera 30 according to the degree of risk associated with the movement of the vehicle 100. The sensor is an example of a "detection unit," and the sensor information is an example of "detection information."
[0047] For example, MoPU12 increases the frame rate of camera 30 as the calculated risk level increases. If the calculated risk level is less than the first threshold, MoPU12 changes the frame rate of camera 30 to 120 frames / second. If the calculated risk level is greater than or equal to the first threshold but less than the second threshold, MoPU12 changes the frame rate of camera 30 to 240, 480, or 960 frames / second. If the calculated risk level is greater than or equal to the second threshold, MoPU12 changes the frame rate of camera 30 to 1920 frames / second. In addition, if the risk level falls within any of the above ranges, MoPU12 may output control signals to radar 32 and infrared camera 34 to cause camera 30 to capture images at the selected frame rate, and to acquire radar signals and capture infrared images at values corresponding to that frame rate.
[0048] For example, MoPU12 lowers the frame rate of camera 30 the lower the calculated risk level. If MoPU12 has set the frame rate of camera 30 to 1920 frames / second and the calculated risk level is above the first threshold but below the second threshold, it changes the frame rate of camera 30 to 240, 480, or 960 frames / second. Also, if MoPU12 has set the frame rate of camera 30 to 1920 frames / second and the calculated risk level is below the first threshold, it changes the frame rate of camera 30 to 120 frames / second. Furthermore, if MoPU12 has set the frame rate of camera 30 to 240, 480, or 960 frames / second and the calculated risk level is below the first threshold, it changes the frame rate of camera 30 to 120 frames / second. In this case as well, control signals may be output to the radar 32 and infrared camera 34 to acquire radar signals and capture infrared images at a value corresponding to the frame rate of the modified camera 30, similar to the above.
[0049] Furthermore, MoPU12 may calculate the degree of risk by using big data related to driving that is known before vehicle 100 is in motion, such as long-tail incident AI (Artificial Intelligence) data (for example, trip data of vehicles with Level 5 autonomous driving control systems implemented) or map information, as information to predict the degree of risk.
[0050] In the above, the risk level was calculated as a score related to the external environment, but the indicators used to score the external environment are not limited to the risk level. For example, MoPU12 may calculate a score related to the external environment other than the risk level based on the direction of movement or speed of objects captured by camera 30, and change the frame rate of camera 30 according to that score. The following describes a case where MoPU12 calculates a speed score, which is a score related to the speed of objects captured by camera 30, and changes the frame rate of camera 30 according to the speed score. As an example, the speed score is set so that it is higher the faster the object's speed and lower the slower the object's speed. MoPU12 then increases the frame rate of camera 30 the higher the calculated speed score and decreases the frame rate of camera 30 the lower the calculated speed score. For this reason, if the calculated speed score exceeds a threshold due to the object's speed being high, MoPU12 changes the frame rate of camera 30 to 1920 frames / second. Furthermore, if the calculated speed score falls below a threshold due to the slow speed of the object, MoPU12 changes the frame rate of camera 30 to 120 frames / second. In this case as well, control signals may be output to radar 32 and infrared camera 34 to acquire radar signals and capture infrared images at a value corresponding to the changed frame rate of camera 30, as described above.
[0051] Next, we will explain how MoPU12 calculates a direction score, which is a score related to the direction of movement of an object captured by camera 30, and how it changes the frame rate of camera 30 according to the direction score. As an example, the direction score is set to be higher when the object is moving towards the road and lower when it is moving away from the road. MoPU12 then increases the frame rate of camera 30 as the calculated direction score increases, and decreases it as the calculated direction score decreases. Specifically, MoPU12 identifies the direction of movement of an object by using AI, etc., and calculates a direction score based on the identified direction of movement. Then, if the calculated direction score is above a threshold because the object's direction of movement was towards the road, MoPU12 changes the frame rate of camera 30 to 1920 frames / second. Also, if the calculated direction score is below a threshold because the object's direction of movement was away from the road, MoPU12 changes the frame rate of camera 30 to 120 frames / second. In this case as well, control signals may be output to the radar 32 and infrared camera 34 to acquire radar signals and capture infrared images at a value corresponding to the frame rate of the modified camera 30, similar to the above.
[0052] Furthermore, MoPU12 may output point information only for objects whose calculated external environment score is above a predetermined threshold. In this case, for example, MoPU12 may determine whether or not to output point information for an object depending on the direction of movement of the object captured by camera 30. For example, MoPU12 does not need to output point information for objects that have little impact on the driving of vehicle 100. Specifically, MoPU12 calculates the direction of movement of objects captured by camera 30 and does not output point information for objects such as pedestrians moving away from the road. On the other hand, MoPU12 outputs point information for objects approaching the road (for example, objects such as pedestrians that may jump into the road). With this configuration, the information processing device 10 does not need to output point information for objects that have little impact on the driving of vehicle 100.
[0053] Furthermore, while the above explanation illustrates a case where MoPU12 calculates the risk level, the disclosed technology is not limited to this embodiment. For example, Central Brain15 may calculate the risk level instead of MoPU12. In this case, Central Brain15 calculates the risk level related to the movement of vehicle 100 as a score related to the external environment for vehicle 100, based on sensor information taken in from multiple types of sensors and point information output from MoPU12. Then, Central Brain15 outputs an instruction to MoPU12 to change the frame rate of camera 30 according to the calculated risk level.
[0054] Furthermore, while the above description illustrates a case where MoPU12 outputs point information based on an image captured by camera 30, the disclosed technology is not limited to this embodiment. For example, MoPU12 may output point information based on radar signals and infrared images instead of images captured by camera 30. MoPU12 can derive the x and y coordinate values of an object from an infrared image of an object captured by infrared camera 34, similar to the image captured by camera 30. Radar 32 can acquire three-dimensional point cloud data of an object based on radar signals. In other words, radar 32 can detect the coordinate of the z axis in the above three-dimensional Cartesian coordinate system. Here, the z axis is the axis along the depth direction of the object and the direction of travel of the vehicle 100, and the coordinate value of the z axis will be referred to as the "z coordinate value" below. In this case, MoPU12 utilizes the principle of a stereo camera to combine the x and y coordinate values of the object captured by the infrared camera 34 at the same time that the radar 32 acquires the object's 3D point cloud data, with the z coordinate value of the object indicated by the 3D point cloud data, to derive the coordinate values of the object's three coordinate axes (x, y, and z axes) as point information. Then, MoPU12 outputs the derived point information to the Central Brain 15.
[0055] Furthermore, while the above description illustrates the case where MoPU12 derives point information, the disclosed technology is not limited to this embodiment. For example, Central Brain15 may derive point information instead of MoPU12. Central Brain15 deriving point information can be achieved, for example, by combining information detected by camera 30L, camera 30R, radar 32, and infrared camera 34. Specifically, Central Brain15 derives the coordinate values of the object's three coordinate axes (x-axis, y-axis, and z-axis) as point information by performing tripoint surveying based on the x- and y-coordinate values of the object captured by camera 30L and the x- and y-coordinate values of the object captured by camera 30R.
[0056] Furthermore, while the above description illustrates a case where the Central Brain 15 controls the autonomous driving of the vehicle 100 based on image and label information output from the IPU 11 and point information output from the MoPU 12, the disclosed technology is not limited to this embodiment. For example, the Central Brain 15 may control the movements of a robot based on the above information output from the IPU 11 and MoPU 12. The robot may be a humanoid smart robot that performs tasks in place of a human. In this case, the Central Brain 15 controls the movements of the robot's arms, palms, fingers, and feet based on the above information output from the IPU 11 and MoPU 12 to perform actions such as grasping, grabbing, holding, carrying, moving, transporting, throwing, kicking, and avoiding objects. When the Central Brain 15 controls the movements of a robot, the IPU 11 and MoPU 12 may be mounted at the positions of the robot's right eye and left eye. In other words, the right eye may be equipped with an IPU11 and MoPU12 designed for the right eye, and the left eye may be equipped with an IPU11 and MoPU12 designed for the left eye.
[0057] (Third embodiment) Next, a third embodiment according to this embodiment will be described, omitting or simplifying any parts that overlap with the above embodiments. As an example, the information processing device 10 according to the third embodiment has the same configuration as shown in Figure 2 as in the first embodiment.
[0058] In the third embodiment, the MoPU12 outputs, as point information, the coordinate values of at least two diagonal points at the vertices of a polygon surrounding the contour of an object recognized from an image captured by another camera. These coordinate values are the x and y coordinate values of the object in the three-dimensional Cartesian coordinate system, as in the first embodiment.
[0059] Figure 4 is an explanatory diagram illustrating an example of point information output by MoPU12. In Figure 4, MoPU12 shows bounding boxes 21, 22, 23, and 24 that enclose the outlines of four objects included in an image captured by another camera. Figure 4 illustrates how MoPU12 outputs the coordinate values of two points that are diagonally opposite each other at the vertices of the bounding boxes 21, 22, 23, and 24 that enclose the outlines of the objects as point information. In this way, MoPU12 may treat objects not as points, but as objects with a certain size.
[0060] Furthermore, when an object is perceived as having a certain size, MoPU12 may output the coordinate values of multiple vertices of the polygon surrounding the object's outline as point information, rather than the coordinate values of two diagonally opposite vertices of the polygon surrounding the object's outline as recognized from an image captured by another camera. For example, using Figure 4 as an example, MoPU12 may output the coordinate values of all four vertices of the bounding box 21, 22, 23, and 24 that encloses the object's outline in a rectangle as point information.
[0061] (Fourth embodiment) Next, a fourth embodiment according to this embodiment will be described, omitting or simplifying any parts that overlap with the above embodiments. As an example, the information processing device 10 according to the fourth embodiment has the same configuration as shown in Figure 2 as in the first embodiment.
[0062] A vehicle 100 equipped with the information processing device 10 according to the fourth embodiment includes sensors comprising at least one of the following: radar, LiDAR, high-resolution, telephoto, ultra-wide-angle, 360-degree, high-performance camera, vision sensor, sound sensor, ultrasonic sensor, vibration sensor, infrared sensor, ultraviolet sensor, radio wave sensor, temperature sensor, and humidity sensor. Sensor information taken in by the information processing device 10 from the sensors includes: shift in the center of gravity of the vehicle's weight, detection of the material of the road, detection of the outside temperature, detection of the outside humidity, detection of the up-down, side-to-side, and diagonal inclination angles of slopes, detection of the degree of road freezing and moisture content, detection of the material, wear status, and air pressure of each tire, road width, whether or not overtaking is prohibited, oncoming vehicles, vehicle type information of vehicles in front and behind, cruising status of those vehicles, and surrounding conditions (birds, animals, soccer balls, accident vehicles, earthquakes, fires, wind, typhoons, heavy rain, light rain, blizzards, fog, etc.). A sensor is an example of a "detection unit," and sensor information is an example of "detection information."
[0063] In the fourth embodiment, the Central Brain 15 calculates control variables for controlling the autonomous driving of the vehicle 100 based on sensor information detected by the sensors. The Central Brain 15 acquires sensor information every one billionth of a second. Specifically, the Central Brain 15 calculates control variables for controlling the wheel speed, tilt, and suspension supporting each of the four wheels of the vehicle 100. Note that the wheel tilt includes both the tilt of the wheel with respect to an axis horizontal to the road and the tilt of the wheel with respect to an axis perpendicular to the road. In this case, the Central Brain 15 calculates a total of 16 control variables for controlling the wheel speed of each of the four wheels, the tilt of each of the four wheels with respect to an axis horizontal to the road, the tilt of each of the four wheels with respect to an axis perpendicular to the road, and the suspension supporting each of the four wheels.
[0064] The Central Brain 15 then controls the autonomous driving of the vehicle 100 based on the control variables calculated above, the point information output from the MoPU 12, and the label information output from the IPU 11. Specifically, the Central Brain 15 controls the in-wheel motors mounted on each of the four wheels based on the control variables described in 16 above, thereby controlling the wheel speed, tilt, and suspension supporting each of the four wheels of the vehicle 100 to perform autonomous driving. In addition, the Central Brain 15 recognizes the position and movement of recognized objects present around the vehicle 100 based on the point information and label information, and based on this recognized information, controls the autonomous driving of the vehicle 100 to avoid collisions with objects, for example. By controlling the autonomous driving of the vehicle 100 in this way, the Central Brain 15 can, for example, perform optimal steering to suit the mountain road when the vehicle 100 is driving on a mountain road, and drive at the optimal angle to suit the parking lot when the vehicle 100 is parking in a parking lot.
[0065] Here, Central Brain15 may be capable of inferring control variables from the sensor information and information obtainable via the network from servers (not shown) using machine learning, more specifically, deep learning. In other words, Central Brain15 can be composed of AI.
[0066] Central Brain15 can determine control variables by performing multivariate analysis using the integral method shown in equation (1) below (see, for example, equation (2)), using the computational power used to achieve Level 6 (hereinafter also referred to as "Level 6 computational power"), which is the computational power of the above sensor information and long-tail incident AI data every billionth of a second. More specifically, by calculating the integral values of the delta values of various Ultra High Resolution sensors with Level 6 computational power, each control variable can be determined at the edge level and in real time, and the results (i.e., each control variable) that occur in the next billionth of a second can be obtained with the highest probability value. To achieve this, for example, the integral values obtained by integrating the delta values (e.g., the change in a small amount of time) of a function that can identify each variable (e.g., the above sensor information and information obtainable via the network) such as air resistance, road resistance, road elements (e.g., garbage), and slip coefficient over time are input into Central Brain15's deep learning model (e.g., a trained model obtained by performing deep learning on a neural network). The Central Brain15 deep learning model outputs control variables corresponding to the input integral value (for example, the control variable with the highest confidence level (i.e., the evaluation value)). The output of the control variables occurs in units of one billionth of a second.
[0067]
number
[0068]
number
[0069] For example, in equation (1), “f(A)” is a simplified representation of a function that shows the behavior of each variable, such as air resistance, road resistance, road elements (e.g., debris), and slip coefficient. Also, for example, equation (1) is an equation that shows the time integral v of “f(A)” from time a to time b. In equation (2), DL stands for deep learning (for example, a deep learning model optimized by performing deep learning on a neural network), and dA n / dt represents the delta value of f(A,B,C,D,...,N), where A,B,C,D,...,N represent variables such as air resistance, road resistance, road elements (e.g., debris), and slip coefficient, and f(A,B,C,D,...,N) represents a function that shows the behavior of A,B,C,D,...,N, V n This represents the values (control variables) output from a deep learning model that has been optimized by applying deep learning to a neural network.
[0070] Here, we have given an example of inputting the integral value obtained by integrating the delta value of a function over time into the Central Brain15 deep learning model, but this is merely one example. For example, the Central Brain15 deep learning model may infer the integral value obtained by integrating the delta value of a function that represents the behavior of each variable such as air resistance, road resistance, road elements, and slip coefficient over time (for example, the result that will occur in the next 1 / 1 billionth of a second), and as an inference result, the integral value with the highest confidence level (i.e., evaluation value) may be obtained by Central Brain15 every 1 / 1 billionth of a second.
[0071] Furthermore, while examples of inputting integral values into a deep learning model and outputting integral values from a deep learning model are given here, these are merely examples, and the technology of this disclosure can be implemented without using integral values. For example, deep learning may be performed on a neural network using training data where values corresponding to A, B, C, D, ..., N are used as example data and values corresponding to at least one control variable (for example, the result that occurs in the next 1 / 100-billionth of a second) are used as ground truth data, so that at least one control variable is inferred by an optimized deep learning model.
[0072] The control variables obtained in Central Brain15 can be further refined by increasing the number of Deep Learning iterations. For example, more accurate control variables can be calculated using vast amounts of data such as tire and motor rotation, steering angle, road material, weather, the effects of debris and quadratic deceleration, slippage, and steering and speed control methods for balance loss and recovery, as well as long-tail incident AI data.
[0073] (Fifth embodiment) Next, a fifth embodiment according to this embodiment will be described, omitting or simplifying any parts that overlap with the above embodiments. Figure 5 is a third block diagram showing an example of the configuration of the information processing device 10. Note that Figure 5 shows only a portion of the configuration of the information processing device 10.
[0074] As shown in Figure 5, in the MoPU 12, visible light images and infrared images of objects captured by the camera 30 are input to the core 17 at a frame rate of 100 frames / second or more. The camera 30 consists of a visible light camera 30A capable of capturing visible light images of objects and an infrared camera 30B capable of capturing infrared images of objects. The core 17 then outputs point information to the Central Brain 15 based on at least one of the input visible light images and infrared images.
[0075] Here, if the core 17 can identify an object from the visible light image of the object captured by the visible light camera 30A, it outputs point information based on the visible light image. On the other hand, if the core 17 cannot capture the object from the visible light image due to predetermined factors, it outputs point information based on the infrared image of the object captured by the infrared camera 30B. For example, one predetermined factor is the effect of darkness, which may prevent the core 17 from capturing the object from the visible light image. In this case, the core 17 uses the infrared camera 30B to detect the heat of the object and outputs point information of the object based on the infrared image obtained from that detection. However, the core 17 may also output point information based on both the visible light image and the infrared image.
[0076] Furthermore, MoPU12 synchronizes the timing of visible light image capture by the visible light camera 30A and the timing of infrared image capture by the infrared camera 30B. Specifically, MoPU12 outputs control signals to the camera 30 so that visible light and infrared images are captured at the same time. As a result, the number of images captured per second by the visible light camera 30A and the number of images captured per second by the infrared camera 30B are synchronized (for example, 1920 frames / second).
[0077] (Sixth embodiment) Next, a sixth embodiment according to this embodiment will be described, omitting or simplifying any parts that overlap with the above embodiments. Figure 6 is a fourth block diagram showing an example of the configuration of the information processing device 10. Note that Figure 6 shows only a portion of the configuration of the information processing device 10.
[0078] As shown in Figure 6, in MoPU12, images of objects captured by camera 30 and radar signals based on reflected electromagnetic waves from objects irradiated by radar 32 are input to core 17 at a frame rate of 100 frames / second or more. Core 17 then outputs point information to Central Brain 15 based on the input images of objects and radar signals. Core 17 can derive the x and y coordinate values of an object from the input images of objects. As described above, radar 32 can acquire 3D point cloud data of an object based on radar signals and detect the coordinate of the z axis in the 3D Cartesian coordinate system. In this case, core 17 uses the principle of a stereo camera to derive the coordinate values of the object's three coordinate axes (x axis, y axis, and z axis) as point information by combining the x and y coordinate values of the object captured by camera 30 at the same time that radar 32 acquires the 3D point cloud data of the object, with the z coordinate value of the object indicated by the 3D point cloud data. Furthermore, the image of the object input to core 17 as described above may include at least one of a visible light image and an infrared image.
[0079] Furthermore, MoPU12 synchronizes the timing of image capture by camera 30 with the timing of radar 32 acquiring 3D point cloud data of objects based on radar signals. Specifically, MoPU12 outputs control signals to camera 30 and radar 32 to capture images and acquire 3D point cloud data of objects at the same time. As a result, the number of images captured per second by camera 30 and the number of 3D point cloud data acquired per second by radar 32 are synchronized (for example, 1920 frames / second). Thus, the number of images captured per second by camera 30 and the number of 3D point cloud data acquired per second by radar 32 are greater than the frame rate of the ultra-high resolution camera equipped with IPU11, i.e., the number of images captured per second by the ultra-high resolution camera.
[0080] (Seventh Embodiment) Next, a seventh embodiment according to this embodiment will be described, omitting or simplifying any parts that overlap with the above embodiments. As an example, the information processing device 10 according to the seventh embodiment has the same configuration as shown in Figure 2 as in the first embodiment.
[0081] In the seventh embodiment, the Central Brain 15 associates point information output from the MoPU 12 with the label information at the same time that the IPU 11 outputs the label information. Furthermore, if new point information is output from the MoPU 12 after the Central Brain 15 has associated the point information and the label information, it also associates the new point information with the label information. The new point information is point information of the same object as the point information associated with the label information, and is one or more pieces of point information between the time the association is made and the next label information is output. In the seventh embodiment, as in the above embodiment, the frame rate of the separate camera with the built-in MoPU 12 is 100 frames / second or more (for example, 1920 frames / second), and the frame rate of the ultra-high resolution camera with the built-in IPU 11 is 10 frames / second.
[0082] Figure 7 is an explanatory diagram illustrating an example of the correspondence between point information and label information. In the following explanation, the number of point information points output per second from MoPU12 will be referred to as the "point information output rate," and the number of label information points output per second from IPU11 will be referred to as the "label information output rate."
[0083] Figure 7 shows the time series of the output rate of point information P4 for object B14. The output rate of point information P4 for object B14 is 1920 frames / second. Also, point information P4 moves from right to left in the figure. The output rate of label information for object B14 is 10 frames / second, which is lower than the output rate of point information P4.
[0084] First, at time t0, no label information for object B14 has been output from IPU11. Therefore, at time t0, Central Brain15 recognizes the coordinate values (position information) of object B14 based on point information P4, but does not recognize what object B14 is.
[0085] Next, at time t1, label information for object B14 is output from IPU11. Therefore, Central Brain15 derives the label information "PERSON" for object B14 based on this label information. Then, Central Brain15 associates the label information "PERSON" derived at time t1 with the coordinate values (position information) of point information P4 output from MoPU12 at time t1. As a result, at time t1, Central Brain15 recognizes the coordinate values (position information) of object B14 based on point information P4, and also recognizes what object B14 is.
[0086] In Figure 7, time t2 is defined as the timing when the next label information for object B14 is output from IPU11. Therefore, at time t2, Central Brain15 derives the label information "PERSON" for object B14 based on the label information output from IPU11. Then, Central Brain15 associates the label information "PERSON" derived at time t2 with the coordinate values (position information) of point information P4 output from MoPU12 at time t2.
[0087] Here, due to the frame rate difference between the separate camera with the built-in MoPU12 and the ultra-high resolution camera with the built-in IPU11, during the period from time t1 to time t2, the Central Brain15 acquires point information P4 for object B14, but not label information. In this case, the Central Brain15 associates the point information P4 acquired during the period from time t1 to time t2 with the label information "PERSON" that was associated with the preceding time t1. Here, the point information P4 acquired by the Central Brain15 during the period from time t1 to time t2 is an example of "new point information". In the example shown in Figure 7, since multiple point information P4s were output from the MoPU12 during the period from time t1 to time t2, the Central Brain15 acquired multiple point information P4s. Therefore, in the example shown in Figure 7, Central Brain 15 associates each of the multiple point information P4 acquired during the period from time t1 to time t2 with the label information "PERSON" associated with the previous time t1. However, unlike the example shown in Figure 7, if one point information P4 is output from MoPU 12 during the period from time t1 to time t2, Central Brain 15 associates that one point information P4 with the label information "PERSON" associated with the previous time t1.
[0088] Here, even if there is a period when the type of object being tracked is uncertain, Central Brain15 continuously outputs the object's point information at a high frame rate, so the risk of losing the object's coordinate values (position information) is low. For this reason, once Central Brain15 has associated point information with label information, it can estimate the label information of the most recent point information acquired until the next label information is acquired.
[0089] (Eighth embodiment) Next, an eighth embodiment according to this embodiment will be described, omitting or simplifying any parts that overlap with the above embodiments. When the information processing device 10 that controls the autonomous driving of the vehicle 100 performs advanced computational processing, heat generation becomes a problem. Therefore, the eighth embodiment provides a vehicle 100 that has a cooling function for the information processing device 10.
[0090] Figure 8 is an explanatory diagram showing the schematic configuration of vehicle 100. As shown in Figure 8, vehicle 100 is equipped with an information processing device 10, a cooling execution device 110, and a cooling unit 120.
[0091] The information processing device 10 according to the eighth embodiment is a device for controlling the automatic driving of the vehicle 100, and as an example, has the same configuration as shown in Figure 2 as in the first embodiment. The cooling execution device 110 acquires the object detection result from the information processing device 10 and, based on the detection result, causes the cooling unit 120 to perform cooling of the information processing device 10. The cooling unit 120 cools the information processing device 10 using at least one cooling means such as air cooling means, water cooling means, and liquid nitrogen cooling means. In the following description, the object to be cooled in the information processing device 10 will be described as the Central Brain 15 (more specifically, the CPU 14 that constitutes the Central Brain 15) that controls the automatic driving of the vehicle 100, but is not limited to this.
[0092] The information processing device 10 and the cooling execution device 110 are connected to each other via a network (not shown). This network may be a vehicle network, the Internet, a LAN (Local Area Network), or a mobile communication network. The mobile communication network may conform to any of the following communication methods: 5G (5th Generation), LTE (Long Term Evolution), 3G (3rd Generation), or 6G (6th Generation) or later.
[0093] Figure 9 is a block diagram showing an example of the functional configuration of the cooling execution device 110. As shown in Figure 9, the cooling execution device 110 has an acquisition unit 112, an execution unit 114, and a prediction unit 116 as its functional configuration.
[0094] The acquisition unit 112 acquires the object detection result from the information processing device 10. For example, as a result of the detection, the acquisition unit 112 acquires the point information of the object output from the MoPU 12.
[0095] The execution unit 114 initiates cooling of the Central Brain 15 based on the object detection results acquired by the acquisition unit 112. For example, if the execution unit 114 recognizes that an object is moving based on the point information of the object output from the MoPU 12, it initiates cooling of the Central Brain 15 by the cooling unit 120.
[0096] Furthermore, the execution unit 114 is not limited to performing cooling on the Central Brain 15 based on the object detection result, but may also perform cooling on the Central Brain 15 based on the prediction result of the operating status of the information processing device 10.
[0097] Here, the prediction unit 116 predicts the operating status of the information processing device 10, specifically the Central Brain 15, based on the object detection results acquired by the acquisition unit 112. For example, the prediction unit 116 acquires a learning model stored in a predetermined memory area. Then, the prediction unit 116 predicts the operating status of the Central Brain 15 by inputting the object point information output from the MoPU 12 acquired by the acquisition unit 112 into the learning model. Here, the learning model outputs the computing power status and change amount of the Central Brain 15 as the operating status. The prediction unit 116 may also predict and output the temperature change of the information processing device 10, specifically the Central Brain 15, along with the operating status. For example, the prediction unit 116 predicts the temperature change of the Central Brain 15 based on the number of object point information output from the MoPU 12 acquired by the acquisition unit 112. In this case, the prediction unit 116 predicts that the more point information there is, the larger the temperature change will be, and that the fewer point information there is, the smaller the temperature change will be.
[0098] In the above case, the execution unit 114 initiates cooling of the Central Brain 15 by the cooling unit 120 based on the prediction result of the prediction unit 116 regarding the operating status of the Central Brain 15. For example, the execution unit 114 initiates cooling by the cooling unit 120 if the predicted computing power status and change amount of the Central Brain 15 exceeds a predetermined threshold. Also, the execution unit 114 initiates cooling by the cooling unit 120 if the temperature based on the temperature change of the Central Brain 15, as predicted for the operating status, exceeds a predetermined threshold.
[0099] Furthermore, the execution unit 114 may perform cooling on the Central Brain 15 using cooling means corresponding to the temperature change prediction result of the prediction unit 116. For example, the higher the predicted temperature of the Central Brain 15, the more cooling means the cooling unit 120 may use to perform cooling. Specifically, if the execution unit 114 predicts that the temperature of the Central Brain 15 will exceed a first threshold, it may perform cooling on the cooling unit 120 using one cooling means. On the other hand, if the execution unit 114 predicts that the temperature of the Central Brain 15 will exceed a second threshold that is higher than the first threshold, it may perform cooling on the cooling unit 120 using multiple cooling means.
[0100] Furthermore, the execution unit 114 may use a more powerful cooling means to cool the Central Brain 15, especially if the predicted temperature of the Central Brain 15 is higher. For example, if the execution unit 114 predicts that the temperature of the Central Brain 15 will exceed a first threshold, it may instruct the cooling unit 120 to perform cooling using an air cooling means. If the execution unit 114 predicts that the temperature of the Central Brain 15 will exceed a second threshold, which is higher than the first threshold, it may instruct the cooling unit 120 to perform cooling using a water cooling means. Furthermore, if the execution unit 114 predicts that the temperature of the Central Brain 15 will exceed a third threshold, which is higher than the second threshold, it may instruct the cooling unit 120 to perform cooling using a liquid nitrogen cooling means.
[0101] Furthermore, the execution unit 114 may determine the cooling means to be used for cooling based on the number of object point information output from the MoPU 12 acquired by the acquisition unit 112. In this case, the execution unit 114 may perform cooling of the Central Brain 15 using a more powerful cooling means the more point information there is. For example, if the number of point information exceeds a first threshold, the execution unit 114 will have the cooling unit 120 perform cooling using air cooling means. Also, if the number of point information exceeds a second threshold higher than the first threshold, the execution unit 114 will have the cooling unit 120 perform cooling using water cooling means. Furthermore, if the number of point information exceeds a third threshold higher than the second threshold, the execution unit 114 will have the cooling unit 120 perform cooling using liquid nitrogen cooling means.
[0102] Incidentally, one trigger for the activation of the Central Brain 15 is the detection of a moving object on the roadway. For example, if a moving object is detected on the roadway while the vehicle 100 is performing autonomous driving, the Central Brain 15 may perform calculations to control the vehicle 100 in relation to that object. As mentioned above, the heat generated when the Central Brain 15 performs advanced calculations to control the autonomous driving of the vehicle 100 is a problem. Therefore, the cooling execution device 110 according to the eighth embodiment predicts the heat dissipation of the Central Brain 15 based on the object detection result by the information processing device 10, and performs cooling on the Central Brain 15 before or simultaneously with the start of heat dissipation. As a result, the Central Brain 15 is prevented from becoming hot during the autonomous driving of the vehicle 100, enabling advanced calculations during autonomous driving.
[0103] (Ninth embodiment) Next, a ninth embodiment according to this embodiment will be described, omitting or simplifying any parts that overlap with the above embodiments. The MoPU 12 in the information processing device 10 according to the ninth embodiment derives the z-coordinate value of an object as point information from an image of the object captured by the camera 30. Hereinafter, each aspect of the information processing device 10 according to the ninth embodiment will be described in order.
[0104] The information processing device 10 according to the first embodiment has the same configuration as shown in Figure 3 as in the second embodiment.
[0105] In the first embodiment described above, the MoPU 12 derives the z-coordinate value of an object as point information from images of the object captured by multiple cameras 30, specifically cameras 30L and 30R. As described above, when using one MoPU 12, it is possible to derive the x-coordinate and y-coordinate values of an object as point information. Here, when using two MoPU 12s, it is possible to derive the z-coordinate value of an object as point information based on images of the object captured by two cameras 30, utilizing the principle of a stereo camera. Therefore, in this first embodiment, the z-coordinate value of an object is derived as point information based on images of the object captured by camera 30L of MoPU 12L and camera 30R of MoPU 12R, respectively, utilizing the principle of a stereo camera.
[0106] The information processing device 10 according to the second embodiment has the same configuration as shown in Figure 3 as in the second embodiment.
[0107] In the second embodiment described above, the MoPU 12 derives the x, y, and z coordinate values of an object as point information from the image of the object captured by the camera 30 and the radar signal based on the reflected waves from the object of electromagnetic waves irradiated onto the object by the radar 32. As described above, the radar 32 is capable of acquiring 3D point cloud data of an object based on the radar signal. In other words, the radar 32 is capable of detecting the coordinate of the z axis in the 3D Cartesian coordinate system. In this case, the MoPU 12 uses the principle of a stereo camera to combine the x and y coordinate values of the object captured by the camera 30 at the same time that the radar 32 acquires the 3D point cloud data of the object, with the z coordinate value of the object indicated by the 3D point cloud data, to derive the coordinate values of the object's three coordinate axes as point information.
[0108] The information processing device 10 according to the third embodiment has the configuration shown in Figure 10. Figure 10 is a fifth block diagram showing an example of the configuration of the information processing device 10. Note that Figure 10 shows only a part of the configuration of the information processing device 10.
[0109] In the third embodiment described above, the MoPU 12 derives the z-coordinate value of the object as point information from the image of the object captured by the camera 30 and the structured light captured by the illumination device 130.
[0110] As shown in Figure 10, in MoPU12, images of objects captured by camera 30 and distortion information showing the distortion of the structured light pattern, which is the result of camera 140 capturing structured light illuminating the object by illumination device 130, are input to core 17 at a frame rate of 100 frames / second or more. Core 17 then outputs point information to Central Brain 15 based on the input object images and distortion information.
[0111] Here, one method for identifying the three-dimensional position or shape of an object is the structured light method. The structured light method involves shining a dot-patterned structured light onto an object and obtaining depth information from the distortion of the pattern. The structured light method is disclosed, for example, in the reference (http: / / ex-press.jp / wp-content / uploads / 2018 / 10 / 018_teledyne_3rd.pdf).
[0112] The illumination device 130 shown in Figure 10 illuminates an object with structured light. The camera 140 captures the structured light illuminated on the object by the illumination device 130. The camera 140 then outputs distortion information based on the distortion of the captured structured light pattern to the core 17.
[0113] Here, MoPU12 synchronizes the timing of image capture by camera 30 with the timing of image capture of structured light by camera 140. Specifically, MoPU12 outputs control signals to camera 30 and camera 140 so that images are captured at the same time. As a result, the number of images captured per second by camera 30 and the number of images captured per second by camera 140 are synchronized (for example, 1920 frames / second). Thus, the number of images captured per second by camera 30 and the number of images captured per second by camera 140 are greater than the frame rate of the ultra-high resolution camera equipped with IPU11, i.e., the number of images captured per second by the ultra-high resolution camera.
[0114] Then, Core 17 combines the x and y coordinate values of the object, which were captured by Camera 30 at the same time that the structured light was captured by Camera 140, with distortion information based on the distortion of the structured light pattern, to derive the z coordinate value of the object as point information.
[0115] The information processing device 10 according to the fourth embodiment has the configuration shown in Figure 11. Figure 11 is a sixth block diagram showing an example of the configuration of the information processing device 10. Note that Figure 11 shows only a part of the configuration of the information processing device 10.
[0116] The block diagram shown in Figure 11 is the same as the block diagram shown in Figure 2, but with the addition of a Lidar sensor 18. The Lidar sensor 18 is a sensor that acquires point cloud data including objects in three-dimensional space and the road surface on which the vehicle 100 is traveling. The information processing device 10 can use the point cloud data acquired by the Lidar sensor 18 to derive positional information in the depth direction of an object, that is, the z-coordinate value of the object. It is assumed that the point cloud data acquired by the Lidar sensor 18 is acquired at longer intervals than the x-coordinate and y-coordinate values of the object output from the MoPU 12. The MoPU 12 is also equipped with a camera 30, similar to the above embodiment of the ninth embodiment.
[0117] In the fourth aspect, the MoPU 12 utilizes the principle of a stereo camera to combine the x and y coordinate values of the object captured by the camera 30 at the same time that the Lidar sensor 18 acquires point cloud data of the object, with the z coordinate value of the object indicated by the point cloud data, to derive the coordinate values of the object's three coordinate axes as point information.
[0118] In this fourth embodiment, MoPU12 derives the z-coordinate of the object at time t+1 as point information from the x, y, and z-coordinates of the object at time t and the x and y-coordinates of the object at the next time point (e.g., time t+1). Time t is an example of a "first time point," and time t+1 is an example of a "second time point." In the fourth embodiment, the z-coordinate of the object at time t+1 is derived using shape information, i.e., geometry. The details of this will be explained below.
[0119] Figure 12 schematically illustrates the time-series coordinate detection of an object. In Figure 12, J represents the position of the object, which is represented by a rectangle, and the object's position moves from J1 to J2 in a time series. In Figure 12, the coordinates of the object at time t when it is located at J1 are (x1, y1, z1), and the coordinates of the object at time t+1 when it is located at J2 are (x2, y2, z2).
[0120] First, let's explain the time t. The MoPU12 derives the x and y coordinate values of an object from the image of the object captured by the camera 30. Subsequently, the MoPU12 integrates the z coordinate value of the object, indicated by the point cloud data acquired from the Lidar sensor 18, with the above x and y coordinate values to derive the three-dimensional coordinate values (x1, y1, z1) of the object at time t.
[0121] Next, let's discuss time t+1. MoPU12 derives the z-coordinate of an object at time t+1 based on the spatial geometry and the changes in the object's x and y coordinates from time t to time t+1. The spatial geometry includes the shape of the road surface obtained from images captured by the ultra-high-resolution camera equipped on IPU11 and point cloud data from the Lidar sensor 18, as well as the shape of the vehicle 100.
[0122] The geometry representing the shape of the road surface is generated in advance at time t. By using the geometry representing the shape of the vehicle 100 together with the geometry representing the shape of the road surface, MoPU12 can simulate the case of vehicle 100 traveling on the road surface and estimate the amount of movement along the x, y, and z axes.
[0123] Therefore, MoPU12 derives the x and y coordinates of the object at time t+1 from the image of the object captured by camera 30. MoPU12 can derive the z coordinate of the object at time t+1 by calculating the amount of z-axis movement from the simulation when the x and y coordinates of the object change from the x and y coordinates (x1, y1) at time t to the x and y coordinates (x2, y2) at time t+1. Then, MoPU12 integrates the above x and y coordinates and z coordinate to derive the three-dimensional coordinates (x2, y2, z2) of the object at time t+1.
[0124] As shown in Figure 12, since the object moves in the depth direction along with the movement in the planar coordinates (i.e., along the x and y axes), it is also necessary to detect movement in the z-axis direction in order to control the automatic driving of the vehicle 100 with high precision. However, the MoPU 12 may not be able to acquire the z-coordinate value of the object, which can be derived from the point cloud data of the Lidar sensor 18, as quickly as the x and y coordinate values of the object. Therefore, in the fourth embodiment described above, the MoPU 12 derives the z-coordinate value of the object at time t+1 from the x, y, and z coordinate values of the object at time t+1 and the x and y coordinate values of the object at time t+1. As a result, according to the information processing device 10 in the fourth embodiment described above, the MoPU 12 can perform 3D motion detection along with 2D motion detection using high-speed frame shots, with high performance and low data volume.
[0125] Furthermore, while the above description illustrates a case where MoPU12 derives the z-coordinate value of an object as point information from an image of the object captured by camera 30, the disclosed technology is not limited to this embodiment. For example, Central Brain15 may derive the z-coordinate value of an object as point information instead of MoPU12. In this case, Central Brain15 derives the z-coordinate value of an object as point information by performing the same processing that MoPU12 performed in the above description on an image of an object captured by camera 30. As an example, Central Brain15 derives the z-coordinate value of an object as point information from images of an object captured by multiple cameras 30, specifically cameras 30L and 30R. In this case, Central Brain15 uses the principle of a stereo camera to derive the z-coordinate value of an object as point information based on images of an object captured by camera 30L of MoPU12L and camera 30R of MoPU12R, respectively.
[0126] (Tenth embodiment) Next, a tenth embodiment according to this embodiment will be described, omitting or simplifying any parts that overlap with the above embodiments. Figure 13 is the seventh block diagram showing an example of the configuration of the information processing device 10. Note that Figure 13 shows only a portion of the configuration of the information processing device 10.
[0127] As shown in Figure 13, in MoPU12, an image of an object captured by the event camera 30C (hereinafter sometimes referred to as "event image") is input to core 17. Core 17 then outputs point information to Central Brain 15 based on the input event image. The event camera is disclosed, for example, in the reference (https: / / dendenblog.xyz / event-based-camera / ).
[0128] Figure 14 is an explanatory diagram illustrating the image of an object (event image) captured by the event camera 30C. Figure 14(A) shows the object that is the target of the event camera 30C. Figure 14(B) shows an example of an event image. Figure 14(C) shows an example of calculating the centroid of the difference between the image captured at the current time and the image captured at the previous time, as point information. In an event image, the difference between the image captured at the current time and the image captured at the previous time is extracted as points. Therefore, when using the event camera 30C, for example, as shown in Figure 14(B), points are extracted for each moving part of the person area shown in Figure 14(A).
[0129] In contrast, as shown in Figure 14(C), Core 17 extracts the coordinates of feature points representing the area of the person (for example, only one point) after extracting the person, which is an object. This reduces the amount of data transferred to Central Brain 15 and memory 16. Since event images can extract the person, which is an object, at any frame rate, in the case of event camera 30C, it is possible to extract at a frame rate higher than the maximum frame rate of camera 30 mounted on MoPU 12 in the above embodiment (e.g., 1920 frames / second), and the point information of the object can be captured with high accuracy.
[0130] In addition, the information processing device 10 according to the tenth embodiment may also include a visible light camera 30A in addition to the event camera 30C in the MoPU 12, similar to the above embodiment. In this case, the visible light image and event image of the object captured by the visible light camera 30A are input to the core 17 of the MoPU 12. The core 17 then outputs point information to the Central Brain 15 based on at least one of the input visible light image and event image.
[0131] For example, if the core 17 can identify an object from the visible light image captured by the visible light camera 30A, it outputs point information based on the visible light image. On the other hand, if the core 17 cannot capture the object from the visible light image due to predetermined factors, it outputs point information based on an event image. The predetermined factors include at least one of the following: the object's movement speed is greater than or equal to a predetermined value, and the change in ambient light intensity per unit time is greater than or equal to a predetermined value. For example, if the object is moving at high speed and cannot be captured from the visible light image, the core 17 identifies the object based on the event image and outputs the x and y coordinate values of the object as point information. Also, if the object cannot be captured from the visible light image due to a sudden change in ambient light intensity such as backlighting, the core 17 identifies the object based on the event image and outputs the x and y coordinate values of the object as point information. With this configuration, the information processing device 10 can use different cameras 30 to capture objects depending on predetermined factors.
[0132] Figure 15 schematically shows an example of the hardware configuration of a computer 1200 that functions as an information processing device 10 or a cooling execution device 110. A program installed on the computer 1200 can cause the computer 1200 to function as one or more "parts" of the apparatus according to this embodiment, or to cause the computer 1200 to execute operations associated with the apparatus according to this embodiment or such one or more "parts", and / or to cause the computer 1200 to execute a process or a stage of such process according to this embodiment. Such a program may be executed by the CPU 1212 to cause the computer 1200 to execute specific operations associated with some or all of the blocks in the flowcharts and block diagrams described herein.
[0133] The computer 1200 according to this embodiment includes a CPU 1212, RAM 1214, and a graphics controller 1216, which are interconnected by a host controller 1210. The computer 1200 also includes input / output units such as a communication interface 1222, a storage device 1224, a DVD drive, and an IC card drive, which are connected to the host controller 1210 via an input / output controller 1220. The DVD drive may be a DVD-ROM drive and a DVD-RAM drive, etc. The storage device 1224 may be a hard disk drive and a solid-state drive, etc. The computer 1200 also includes legacy input / output units such as a ROM 1230 and a keyboard, which are connected to the input / output controller 1220 via an input / output chip 1240.
[0134] The CPU 1212 operates according to the programs stored in the ROM 1230 and RAM 1214, thereby controlling each unit. The graphics controller 1216 acquires the image data generated by the CPU 1212 and stores it in the frame buffer provided in RAM 1214 or within itself, so that the image data is displayed on the display device 1218.
[0135] The communication interface 1222 communicates with other electronic devices via a network. The storage device 1224 stores programs and data used by the CPU 1212 in the computer 1200. The DVD drive reads programs or data from a DVD-ROM or the like and provides them to the storage device 1224. The IC card drive reads programs and data from an IC card and / or writes programs and data to an IC card.
[0136] The ROM 1230 stores boot programs and / or hardware-dependent programs of the computer 1200, which are executed by the computer 1200 upon activation. The input / output chip 1240 may also connect various input / output units to the input / output controller 1220 via USB ports, parallel ports, serial ports, keyboard ports, mouse ports, etc.
[0137] The program is provided on a computer-readable storage medium such as a DVD-ROM or IC card. The program is read from the computer-readable storage medium and installed on a storage device 1224, RAM 1214, or ROM 1230, which are examples of computer-readable storage media, and executed by the CPU 1212. The information processing described within these programs is read by the computer 1200, resulting in coordination between the program and the various types of hardware resources described above. The apparatus or method may be configured to realize the operation or processing of information in accordance with the use of the computer 1200.
[0138] For example, when communication is performed between a computer 1200 and an external device, the CPU 1212 may execute a communication program loaded into RAM 1214 and, based on the processing described in the communication program, instruct the communication interface 1222 to perform communication processing. Under the control of the CPU 1212, the communication interface 1222 reads transmission data stored in a transmission buffer area provided in a recording medium such as RAM 1214, storage device 1224, DVD-ROM, or IC card, transmits the read transmission data to the network, or writes received data received from the network to a reception buffer area provided on the recording medium.
[0139] Furthermore, the CPU 1212 may read all or necessary parts of a file or database stored on an external recording medium such as the storage device 1224, a DVD drive (DVD-ROM), or an IC card into the RAM 1214, and perform various types of processing on the data in the RAM 1214. The CPU 1212 may then write the processed data back to the external recording medium.
[0140] Various types of information, such as various types of programs, data, tables, and databases, may be stored on the recording medium and subjected to information processing. The CPU 1212 may perform various types of processing on the data read from RAM 1214, including various types of operations, information processing, conditional judgments, conditional branching, unconditional branching, information retrieval / replacement, etc., as described throughout this disclosure and specified by the program instruction sequence, and write the results back to RAM 1214. The CPU 1212 may also retrieve information in files, databases, etc., within the recording medium. For example, if multiple entries are stored in the recording medium, each having an attribute value of a first attribute associated with an attribute value of a second attribute, the CPU 1212 may search among the multiple entries for an entry that matches the specified condition for the attribute value of the first attribute, read the attribute value of the second attribute stored in that entry, and thereby obtain the attribute value of the second attribute associated with the first attribute that satisfies the predetermined condition.
[0141] The program or software module described above may be stored on or near the computer 1200 in a computer-readable storage medium. Alternatively, a recording medium such as a hard disk or RAM provided within a server system connected to a dedicated communication network or the Internet can be used as a computer-readable storage medium, thereby providing the program to the computer 1200 via the network.
[0142] In this embodiment, blocks in the flowchart and block diagram may represent a stage in a process in which an operation is performed or a "part" of a device that has the role of performing an operation. A particular stage and "part" may be implemented by a dedicated circuit, a programmable circuit supplied with computer-readable instructions stored on a computer-readable storage medium, and / or a processor supplied with computer-readable instructions stored on a computer-readable storage medium. The dedicated circuit may include digital and / or analog hardware circuits, and may include integrated circuits (ICs) and / or discrete circuits. The programmable circuit may include reconfigurable hardware circuits, such as field-programmable gate arrays (FPGAs) and programmable logic arrays (PLAs), which include logical AND, logical OR, exclusive OR, negated AND, negated OR, and other logical operations, flip-flops, registers, and memory elements.
[0143] A computer-readable storage medium may include any tangible device capable of storing instructions to be executed by a suitable device, and as a result, a computer-readable storage medium having instructions stored therein will comprise a product that includes instructions that can be executed to create means for performing operations specified in a flowchart or block diagram. Examples of computer-readable storage media may include electronic storage media, magnetic storage media, optical storage media, electromagnetic storage media, semiconductor storage media, etc. More specific examples of computer-readable storage media may include floppy disks, diskettes, hard disks, random access memory (RAM), read-only memory (ROM), erasable programmable read-only memory (EPROM or flash memory), electrically erasable programmable read-only memory (EEPROM), static random access memory (SRAM), compact disk read-only memory (CD-ROM), digital multipurpose disc (DVD), Blu-ray® disc, memory stick, integrated circuit card, etc.
[0144] Computer-readable instructions may include assembler instructions, instruction set architecture (ISA) instructions, machine instructions, machine-dependent instructions, microcode, firmware instructions, state setting data, or source code or object code written in any combination of one or more programming languages, including object-oriented programming languages such as Smalltalk®, Java®, C++, and traditional procedural programming languages such as the C programming language or similar programming languages.
[0145] Computer-readable instructions are used to generate means for a general-purpose computer, special-purpose computer, or other programmable data processing device's processor or programmable circuit to perform an operation specified in a flowchart or block diagram. These instructions are executed locally or via a wide area network (WAN), such as a local area network (LAN) or the internet. It may be provided in a processor or programmable circuit. Examples of processors include computer processors, processing units, microprocessors, digital signal processors, controllers, microcontrollers, and the like.
[0146] Although the present invention has been described above using embodiments, the technical scope of the present invention is not limited to the scope described in the above embodiments. It will be apparent to those skilled in the art that various modifications or improvements can be made to the above embodiments. It will be clear from the claims that such modified or improved forms may also be included in the technical scope of the present invention.
[0147] It should be noted that the execution order of operations, procedures, steps, and stages in the apparatus, systems, programs, and methods shown in the claims, specifications, and drawings is not explicitly stated as "before" or "prior to," and that these can be implemented in any order unless the output of a previous process is used in a later process. Even if the operation flow in the claims, specifications, and drawings is described using phrases such as "first," and "next," for convenience, this does not mean that it is essential to perform the operations in that order.
[0148] In the above embodiment, the processes that each processor (e.g., IPU11, MoPU12, and Central Brain15) is said to perform are merely examples, and the processors that perform each process are not limited. For example, the process that MoPU12 is said to perform in the above embodiment may be performed by Central Brain15 instead of MoPU12, or by other processors other than IPU11, MoPU12, and Central Brain15.
[0149] <Note 1> (1) A first processor outputs point information that captures the captured object as a point from an image of the object taken by the first camera, A second processor outputs identification information that identifies the captured object from an image of the object captured by a second camera facing the same direction as the first camera, Equipped with, Information processing device.
[0150] (2) The system includes a third processor that associates the point information output from the first processor with the identification information output from the second processor. (1) The information processing device described above.
[0151] (3) The frame rate of the first camera is variable, The first processor changes the frame rate of the first camera according to predetermined factors. The information processing device described in (1) or (2).
[0152] (4) The first processor calculates a score relating to the external environment for a predetermined object. (3) The information processing device described above.
[0153] (5) The first processor changes the frame rate of the first camera according to the calculated score regarding the external environment. (4) The information processing device described above.
[0154] (6) From the image of the object captured by the first camera, point information is output, in which the captured object is perceived as a point. From the image of the object captured by the second camera, which is oriented in a direction corresponding to the first camera, identification information is output that identifies the captured object. An information processing method in which a computer performs the processing.
[0155] (7) On the computer, From the image of the object captured by the first camera, point information is output, in which the captured object is perceived as a point. From the image of the object captured by the second camera, which is oriented in a direction corresponding to the first camera, identification information is output that identifies the captured object. An information processing program used to execute a process.
[0156] <Note 2> (1) A first processor outputs coordinate values of at least two coordinate axes in a three-dimensional Cartesian coordinate system of a point indicating the location of an object captured by a first camera, from an image of the object captured by the first camera. A second processor outputs identification information that identifies the captured object from an image of the object captured by a second camera facing the same direction as the first camera, A third processor that associates the coordinate values output from the first processor with the identification information output from the second processor, Equipped with, Information processing device.
[0157] (2) The first processor outputs the coordinate values of at least two diagonal points of the vertices of the polygon surrounding the contour of the object recognized from the image captured by the first camera. (1) The information processing device described above.
[0158] (3) The first processor outputs the coordinate values of multiple vertices of a polygon that surrounds the contour of the object recognized from the image captured by the first camera. (2) The information processing device described in (2).
[0159] (4) From the image of the object captured by the first camera, the coordinate values of at least two coordinate axes in a three-dimensional Cartesian coordinate system of a point indicating the location of the captured object are output. From the image of the object captured by the second camera, which is oriented in a direction corresponding to the first camera, identification information is output that identifies the captured object. The coordinate values and the identification information are associated with each other. An information processing method in which a computer performs the processing.
[0160] (5) On the computer, From the image of the object captured by the first camera, the coordinate values of at least two coordinate axes in a three-dimensional Cartesian coordinate system of a point indicating the location of the captured object are output. From the image of the object captured by the second camera, which is oriented in a direction corresponding to the first camera, identification information is output that identifies the captured object. The coordinate values and the identification information are associated with each other. An information processing program used to execute a process.
[0161] <Note 3> (1) A first processor outputs point information that captures the captured object as a point from an image of the object taken by the first camera, A second processor outputs identification information that identifies the captured object from an image of the object captured by a second camera facing the same direction as the first camera, A third processor associates the point information output from the first processor with the identification information output from the second processor, and controls the automatic driving of the mobile body based on the point information and the identification information. Equipped with, Information processing device.
[0162] (2) The third processor is, Based on the detection information detected by the detection unit, control variables for controlling the automatic driving of the moving object are calculated. Based on the calculated control variables, the point information, and the identification information, the automatic driving of the moving body is controlled. (1) The information processing device described above.
[0163] (3) From the image of the object captured by the first camera, point information is output, in which the captured object is perceived as a point. From the image of the object captured by the second camera, which is oriented in a direction corresponding to the first camera, identification information is output that identifies the captured object. The point information and the identification information are associated, and the automatic driving of the moving object is controlled based on the point information and the identification information. An information processing method in which a computer performs the processing.
[0164] (4) On the computer, From the image of the object captured by the first camera, point information is output, in which the captured object is perceived as a point. From the image of the object captured by the second camera, which is oriented in a direction corresponding to the first camera, identification information is output that identifies the captured object. The point information and the identification information are associated, and the automatic driving of the moving object is controlled based on the point information and the identification information. An information processing program used to execute a process.
[0165] <Note 4> (1) A first processor outputs point information that captures the captured object as a point from an image of the object taken by the first camera, A second processor outputs identification information that identifies the captured object from an image of the object captured by a second camera facing the same direction as the first camera, A third processor that associates the point information output from the first processor with the identification information output from the second processor, Equipped with, The frame rate of the first camera is greater than the frame rate of the second camera. Information processing device.
[0166] (2) The frame rate of the first camera is more than 10 times that of the second camera. (1) The information processing device described above.
[0167] (3) The frame rate of the first camera is 100 frames / second or more, and the frame rate of the second camera is 10 frames / second. (2) The information processing device described in (2).
[0168] (4) From the image of the object captured by the first camera, point information is output, in which the captured object is perceived as a point. From the image of the object captured by the second camera, which has a lower frame rate than the first camera and is oriented in the same direction as the first camera, identification information is output that identifies the captured object. The point information and the identification information are associated with each other. An information processing method in which a computer performs the processing.
[0169] (5) On the computer, From the image of the object captured by the first camera, point information is output, in which the captured object is perceived as a point. From the image of the object captured by the second camera, which has a lower frame rate than the first camera and is oriented in the same direction as the first camera, identification information is output that identifies the captured object. The point information and the identification information are associated with each other. An information processing program used to execute a process.
[0170] <Note 5> (1) A first processor outputs point information that captures the captured object as a point from an image of the object taken by the first camera, A second processor outputs identification information that identifies the captured object from an image of the object captured by a second camera facing the same direction as the first camera, A third processor that associates the point information output from the first processor with the identification information output from the second processor, Equipped with, The first processor calculates a score regarding the external environment for a predetermined moving object, based on the detection information detected by the detection unit and the point information, and determines the degree of risk related to the movement of the moving object. Information processing device.
[0171] (2) The frame rate of the first camera is variable, The first processor changes the frame rate of the first camera according to the calculated risk level. (1) The information processing device described above.
[0172] (3) The aforementioned risk level indicates the degree to which the moving object will travel through dangerous areas in the future. The information processing device described in (1) or (2).
[0173] (4) A first processor outputs point information that captures the captured object as a point from an image of the object taken by the first camera, A second processor outputs identification information that identifies the captured object from an image of the object captured by a second camera facing the same direction as the first camera, A third processor that associates the point information output from the first processor with the identification information output from the second processor, Equipped with, The third processor calculates a score regarding the external environment for a predetermined moving object, based on the detection information detected by the detection unit and the point information, and determines the degree of risk related to the movement of the moving object. Information processing device.
[0174] (5) The frame rate of the first camera is variable, The third processor outputs an instruction to the first processor to change the frame rate of the first camera according to the calculated risk level. (4) The information processing device described above.
[0175] (6) From the image of the object captured by the first camera, point information is output, in which the captured object is perceived as a point. From the image of the object captured by the second camera, which is oriented in a direction corresponding to the first camera, identification information is output that identifies the captured object. The point information and the identification information are associated with each other. Based on the detection information detected by the detection unit and the point information, the system calculates a score for the external environment of a predetermined moving object, which represents the degree of risk related to the movement of the moving object. An information processing method in which a computer performs the processing.
[0176] (7) On the computer, From the image of the object captured by the first camera, point information is output, in which the captured object is perceived as a point. From the image of the object captured by the second camera, which is oriented in a direction corresponding to the first camera, identification information is output that identifies the captured object. The point information and the identification information are associated with each other. Based on the detection information detected by the detection unit and the point information, the system calculates a score for the external environment of a predetermined moving object, which represents the degree of risk related to the movement of the moving object. An information processing program used to execute a process.
[0177] <Note 6> (1) A first processor outputs point information that captures the captured object as a point based on at least one of a visible light image and an infrared image of the object captured by the first camera, A second processor outputs identification information that identifies the captured object from an image of the object captured by a second camera facing the same direction as the first camera, A third processor that associates the point information output from the first processor with the identification information output from the second processor, Equipped with, Information processing device.
[0178] (2) If the first processor cannot capture the object from the visible light image of the object captured by the visible light camera included in the first camera due to predetermined factors, it outputs the point information based on the infrared image of the object captured by the infrared camera included in the first camera. (1) The information processing device described above.
[0179] (3) The first processor synchronizes the timing of capturing the visible light image with the visible light camera and the timing of capturing the infrared image with the infrared camera. (2) The information processing device described in (2).
[0180] (4) Based on at least one of the visible light image and infrared image of the object captured by the first camera, point information is output that captures the captured object as a point. From the image of the object captured by the second camera, which is oriented in a direction corresponding to the first camera, identification information is output that identifies the captured object. The point information and the identification information are associated with each other. An information processing method in which a computer performs the processing.
[0181] (5) On the computer, Based on at least one of the visible light image and infrared image of the object captured by the first camera, point information is output that captures the captured object as a point. From the image of the object captured by the second camera, which is oriented in a direction corresponding to the first camera, identification information is output that identifies the captured object. The point information and the identification information are associated with each other. An information processing program used to execute a process.
[0182] <Note 7> (1) A first processor outputs point information that captures the captured object as a point from an image of the object taken by a first camera and a radar signal based on the reflected waves from the object of electromagnetic waves irradiated onto the object by radar, A second processor outputs identification information that identifies the captured object from an image of the object captured by a second camera facing the same direction as the first camera, A third processor that associates the point information output from the first processor with the identification information output from the second processor, Equipped with, Information processing device.
[0183] (2) The first processor synchronizes the timing of capturing the image with the timing of the radar acquiring three-dimensional point cloud data of the object based on the radar signal. (1) The information processing device described above.
[0184] (3) The number of images captured per unit time by the first camera and the number of 3D point cloud data acquired per unit time by the radar are greater than the number of images captured per unit time by the second camera. The information processing device described in (1) or (2).
[0185] (4) From the image of the object captured by the first camera and the radar signal based on the reflected waves from the object of electromagnetic waves irradiated onto the object by the radar, point information is output that captures the captured object as a point. From the image of the object captured by the second camera, which is oriented in a direction corresponding to the first camera, identification information is output that identifies the captured object. The point information and the identification information are associated with each other. An information processing method in which a computer performs the processing.
[0186] (5) On the computer, From the image of the object captured by the first camera and the radar signal based on the reflected waves from the object of electromagnetic waves irradiated onto the object by the radar, point information is output that captures the captured object as a point. From the image of the object captured by the second camera, which is oriented in a direction corresponding to the first camera, identification information is output that identifies the captured object. The point information and the identification information are associated with each other. An information processing program used to execute a process.
[0187] <Note 8> (1) A first processor outputs point information that captures the captured object as a point from an image of the object taken by the first camera, A second processor outputs label information indicating the type of object captured from an image of the object captured by a second camera facing the same direction as the first camera, A third processor that associates the point information output from the first processor with the label information output from the second processor, Equipped with, Information processing device.
[0188] (2) The third processor associates the position information of the object indicated by the point information with the label information of the object located at the position indicated by the position information. (1) The information processing device described above.
[0189] (3) The third processor associates the point information output from the first processor with the label information at the same time that the second processor outputs the label information. (2) The information processing device described in (2).
[0190] (4) If the third processor receives new point information from the first processor after associating the point information and the label information, it will also associate the new point information with the label information. (2) or (3) the information processing device described above.
[0191] (5) From the image of the object captured by the first camera, point information is output, in which the captured object is perceived as a point. From the image of the object captured by the second camera, which is oriented in a direction corresponding to the first camera, label information indicating the type of the captured object is output. The point information and the label information are associated with each other. An information processing method in which a computer performs the processing.
[0192] (6) On the computer, From the image of the object captured by the first camera, point information is output, in which the captured object is perceived as a point. From the image of the object captured by the second camera, which is oriented in a direction corresponding to the first camera, label information indicating the type of the captured object is output. The point information and the label information are associated with each other. An information processing program used to execute a process.
[0193] <Note 9> (1) An acquisition unit that outputs point information in which the object is captured as a point and identification information in which the object is identified from images of the object taken by multiple cameras facing the corresponding direction, and acquires the detection result of the object by an information processing device that associates the point information and the identification information, Based on the detection results acquired by the acquisition unit, an execution unit causes the information processing device to perform cooling, Equipped with, Cooling execution device.
[0194] (2) The system includes a prediction unit that predicts the operating status of the information processing device based on the detection results acquired by the acquisition unit, The execution unit, based on the prediction result of the prediction unit's prediction of the operating status of the information processing device, causes the information processing device to perform cooling. (1) The cooling device described above.
[0195] (3) The prediction unit predicts the temperature change of the information processing device, The execution unit performs cooling of the information processing device using cooling means corresponding to the temperature change prediction result of the prediction unit. (2) Cooling execution device as described above.
[0196] (4) The detection result acquired by the acquisition unit is the point information. A cooling device described in any one of (1) to (3).
[0197] (5) The system outputs point information that captures the object as a point and identification information that identifies the object from images of the object taken by multiple cameras facing the corresponding direction, and obtains the detection result of the object by an information processing device that associates the point information and the identification information. Based on the acquired detection results, the cooling of the information processing device is performed. A cooling execution method in which a computer performs a process.
[0198] (6) On the computer, The system outputs point information that captures the object as a point and identification information that identifies the object from images of the object taken by multiple cameras facing the corresponding direction, and obtains the detection result of the object by an information processing device that associates the point information and the identification information. Based on the acquired detection results, the cooling of the information processing device is performed. A cooling execution program to run the process.
[0199] <Note 10> (1) A first processor outputs point information that captures the captured object as a point from an image of the object taken by the first camera, A second processor outputs identification information that identifies the captured object from an image of the object captured by a second camera facing the same direction as the first camera, A third processor that associates the point information output from the first processor with the identification information output from the second processor, Equipped with, The first processor derives, from the image of the object captured by the first camera, the coordinate values of the object in the depth direction in a three-dimensional Cartesian coordinate system for a point indicating the location of the object as point information. Information processing device.
[0200] (2) The first processor derives the coordinate values in the depth direction as point information from images of the object captured by a plurality of first cameras. (1) The information processing device described above.
[0201] (3) The first processor derives coordinate values in the width direction, height direction, and depth direction of the object as point information from the image of the object captured by the first camera and the radar signal based on the reflected waves from the object of electromagnetic waves irradiated onto the object by the radar. The information processing device described in (1) or (2).
[0202] (4) The first processor derives coordinate values in the width direction, height direction, and depth direction of the object as point information from the image of the object captured by the first camera and the result of capturing structured light illuminating the object by the illumination device. An information processing device described in any one of (1) to (3).
[0203] (5) The first processor derives the coordinate value in the depth direction at the second time point as point information from the coordinate values in the width direction, height direction, and depth direction of the object in the three-dimensional orthogonal coordinate system at the first time point, and the coordinate values in the width direction and height direction at the second time point, which is the time point following the first time point. An information processing device described in any one of (1) to (4).
[0204] (6) A first processor that outputs point information that captures the photographed object as a point from an image of the object photographed by the first camera, A second processor that outputs identification information that identifies the photographed object from an image of the object photographed by a second camera facing a direction corresponding to the first camera, A third processor that associates the point information output from the first processor and the identification information output from the second processor, Comprising, The third processor derives a coordinate value in the depth direction of the object in a three-dimensional orthogonal coordinate system of a point indicating the existence position of the object as the point information from an image of the object photographed by the first camera, An information processing apparatus.
[0205] (7) Output point information that captures the photographed object as a point from an image of the object photographed by the first camera, Output identification information that identifies the photographed object from an image of the object photographed by a second camera facing a direction corresponding to the first camera, Associate the point information and the identification information, Derive a coordinate value in the depth direction of the object in a three-dimensional orthogonal coordinate system of a point indicating the existence position of the object as the point information from an image of the object photographed by the first camera, An information processing method in which a computer executes processing.
[0206] (8) To a computer, Output point information that captures the photographed object as a point from an image of the object photographed by the first camera, Output identification information that identifies the photographed object from an image of the object photographed by a second camera facing a direction corresponding to the first camera, Associate the point information and the identification information, From the image of the object captured by the first camera, the coordinate values of the object in the depth direction in a three-dimensional Cartesian coordinate system are derived as point information for a point indicating the location of the object. An information processing program used to execute a process.
[0207] <Note 11> (1) A first processor outputs point information from an image of an object captured by an event camera, in which the captured object is treated as a point. A second processor outputs identification information that identifies the captured object from an image of the object taken by a second camera facing the same direction as the event camera, A third processor that associates the point information output from the first processor with the identification information output from the second processor, Equipped with, Information processing device.
[0208] (2) If the first processor cannot capture the object from the visible light image of the object captured by the visible light camera due to predetermined factors, it outputs the point information based on the image of the object captured by the event camera. (1) The information processing device described above.
[0209] (3) The predetermined factors include at least one of the following: the object's moving speed is greater than or equal to a predetermined value, and the change in ambient light intensity per unit time is greater than or equal to a predetermined value. (2) The information processing device described in (2).
[0210] (4) The aforementioned event camera is a camera that outputs an event image representing the difference between an image taken at the current time and an image taken at the previous time. An information processing device described in any one of (1) to (3).
[0211] (5) From the image of the object captured by the event camera, point information is output, in which the captured object is treated as a point. From the image of the object captured by the second camera facing the same direction as the event camera, identification information identifying the captured object is output. The point information and the identification information are associated with each other. An information processing method in which a computer performs the processing.
[0212] (6) On the computer, From the image of the object captured by the event camera, point information is output, in which the captured object is treated as a point. From the image of the object captured by the second camera facing the same direction as the event camera, identification information identifying the captured object is output. The point information and the identification information are associated with each other. An information processing program used to execute a process. [Explanation of Symbols]
[0213] 10 Information Processing Devices 11 IPU (Second Processor) 12 MoPU (First Processor) 15. Central Brain (Third Processor) 110 Cooling execution device
Claims
1. A first processor outputs point information that captures the captured object as a point from an image of the object taken by the first camera, A second processor outputs identification information that identifies the captured object from an image of the object captured by a second camera facing the same direction as the first camera, A third processor that associates the point information output from the first processor with the identification information output from the second processor, Equipped with, The frame rate of the first camera is greater than the frame rate of the second camera. Information processing device.
2. The frame rate of the first camera is 10 times or more the frame rate of the second camera. The information processing apparatus according to claim 1.
3. The frame rate of the first camera is 100 frames / second or more, and the frame rate of the second camera is 10 frames / second. The information processing apparatus according to claim 2.
4. From the image of the object captured by the first camera, point information is output, in which the captured object is perceived as a point. From the image of the object captured by the second camera, which has a lower frame rate than the first camera and is oriented in a direction corresponding to the first camera, identification information is output that identifies the captured object. The point information and the identification information are associated with each other. An information processing method in which a computer performs the processing.
5. On the computer, From the image of the object captured by the first camera, point information is output, in which the captured object is perceived as a point. From the image of the object captured by the second camera, which has a lower frame rate than the first camera and is oriented in a direction corresponding to the first camera, identification information is output that identifies the captured object. The point information and the identification information are associated with each other. An information processing program used to execute a process.
Citation Information
Patent Citations
Simulation system and simulation program therefor
JP2012071394A
Workpiece recognition method and random picking method
JP2017170567A
Moving vehicle, communication system, communication control method, and program
JP2022035198A
Information processing system, information processing device, information processing method, and program
WO2019138835A1
Control loop for navigating a vehicle
WO2021198775A1