Image recognition device and image recognition method

JPWO2025164369A5Pending Publication Date: 2026-05-11
View PDF 0 Cites 0 Cited by

Patent Information

Authority / Receiving Office
JP · JP
Patent Type
Applications
Filing Date
2025-01-17
Publication Date
2026-05-11

AI Technical Summary

Technical Problem

Existing image recognition technologies struggle to balance processing accuracy and speed effectively when performing multitasking, as they primarily rely on data load rather than scene-specific requirements, leading to suboptimal performance in varying environments.

Method used

An image recognition device equipped with a controller unit that dynamically adjusts the processing content of multiple image recognition tasks based on the scene's content, using a machine learning model to optimize network structure and parameters for improved balance between accuracy and speed.

Benefits of technology

Enables more efficient and accurate image recognition by dynamically adapting to scene-specific demands, reducing processing time and power consumption while prioritizing important tasks, thus enhancing multitasking performance.

✦ Generated by Eureka AI based on patent content.
Patent Text Reader

Abstract

This image recognition device comprises: a detector (103) capable of performing multi-task processing for executing a plurality of image recognition tasks on a peripheral image, and capable of adjusting the processing contents of the image recognition tasks; and a controller unit (104) that adjusts the processing contents of the plurality of image recognition tasks in the detector (103). The controller unit (104) receives a peripheral image as an input, and dynamically changes the processing contents of the plurality of image recognition tasks in the detector (103) in accordance with the tendency of the content of the peripheral image.
Need to check novelty before this filing date? Find Prior Art

Description

Image recognition device and image recognition method CROSS-REFERENCE TO RELATED APPLICATIONS

[0001] This application is based on Patent Application No. 2024-14409 filed in Japan on February 1, 2024, and the contents of the original application are incorporated by reference in their entirety.

[0002] The present disclosure relates to an image recognition device and an image recognition method.

[0003] Patent Literature 1 discloses a technology for detecting objects from input image data using three detectors with different object detection accuracies or detection speeds. It also discloses that a controller selects one of the three detectors for each frame of image data and performs object detection. In the technology disclosed in Patent Literature 1, the controller selects one of the three detectors according to a data load, which is information indicating the amount of image data to be processed. When the data load is large, the controller frequently selects a high-speed detector, and when the data load is small, the controller frequently selects a high-precision detector.

[0004] International Publication No. 2021 / 014643

[0005] However, with the technology disclosed in Patent Document 1, problems are thought to arise when multitasking is performed in which a detector is responsible for multiple image recognition tasks (hereinafter referred to as image recognition tasks).

[0006] An image recognition task may be easy or difficult to process depending on the scene depicted in the image. Furthermore, the necessity of each of multiple image recognition tasks may vary depending on the scene depicted in the image. Therefore, when performing multitasking, it is preferable to be able to change the degree of priority given to the processing accuracy or processing speed of each image recognition task depending on the scene depicted in the image. In contrast, the technology disclosed in Patent Document 1 can only select detectors with different object detection accuracy or detection speed depending on the data load, which is information indicating the amount of image data. Therefore, when performing multitasking, it is difficult to perform multiple image recognition tasks with a more desirable balance of processing accuracy and processing speed depending on the scene.

[0007] One object of this disclosure is to provide an image recognition device and an image recognition method that more easily enable image recognition tasks to be performed with a more desirable balance between accuracy and speed depending on the scene.

[0008] The symbols in parentheses in the claims indicate a correspondence with the specific means described in the embodiments described below as one aspect, and do not limit the technical scope of the present disclosure.

[0009] In order to achieve the above-mentioned objective, the image recognition device of the present disclosure is equipped with an image processing unit that is capable of multitasking, performing multiple image recognition tasks on an image, and that is capable of adjusting the processing content of the image recognition tasks, and a controller unit that adjusts the processing content of the multiple image recognition tasks in the image processing unit, and the controller unit takes an image as input and dynamically changes the processing content of the multiple image recognition tasks in the image processing unit depending on the tendency of the content of the image.

[0010] In order to achieve the above object, the image recognition method of the present disclosure includes an image processing step executed by at least one processor, which is capable of multitasking processing in which multiple image recognition tasks are performed on an image and is capable of adjusting the processing content of the image recognition tasks, and a controller step that adjusts the processing content of the multiple image recognition tasks in the image processing step, wherein the controller step takes an image as input and dynamically changes the processing content of the multiple image recognition tasks in the image processing step according to the tendency of the content of the image.

[0011] With the above configuration, the content of multitasking, which executes multiple image recognition tasks on an image, can be dynamically changed according to the trends in the image content. Therefore, it is possible to dynamically change the balance between the processing speed and processing accuracy of the multiple image recognition tasks according to the scene represented by the image content. As a result, it is easier to execute image recognition tasks with a more desirable balance between accuracy and speed according to the scene.

[0012] FIG. 1 is a diagram showing an example of a schematic configuration of an image recognition system; FIG. 2 is a diagram showing an example of a schematic configuration of an image recognition device in embodiment 1; FIG. 3 is a diagram for explaining an example of a learning method of a controller unit; FIG. 4 is a diagram for explaining an example of a case where the NW structure of a detector cannot be dynamically changed; FIG. 5 is a diagram for explaining an example of a case where the NW structure of a detector can be dynamically changed; FIG. 6 is a diagram showing an example of a schematic configuration of an image recognition device in embodiment 2; FIG. 7 is a diagram showing an example of a schematic configuration of an image recognition device in embodiment 3.

[0013] A number of embodiments for the purpose of disclosure will be described with reference to the drawings. For the sake of convenience, parts having the same functions as parts shown in the drawings used in the previous explanations in the number of embodiments will be given the same reference numerals, and their description may be omitted. For parts given the same reference numerals, the explanations in other embodiments may be referred to.

[0014] (Embodiment 1) <Outline of Image Recognition System 1> Hereinafter, a first embodiment of the present disclosure will be described with reference to the drawings. The image recognition system 1 shown in FIG. 1 can be used in a vehicle. As shown in FIG. 1, the image recognition system 1 includes an image recognition device 10, a locator 11, a map database (hereinafter referred to as a map DB) 12, a vehicle state sensor 13, a perimeter monitoring sensor 14, a vehicle control ECU 15, a driving assistance ECU 16, an interior camera 17, a presentation device 18, and an HCU (Human Machine Interface Control Unit) 19. For example, the image recognition device 10, the locator 11, the map DB 12, the vehicle state sensor 13, the perimeter monitoring sensor 14, the vehicle control ECU 15, the driving assistance ECU 16, and the HCU 20 may be configured to be connected to an in-vehicle LAN (LAN) (see the LAN in FIG. 1). Although the vehicle using the image recognition system 1 is not necessarily limited to an automobile, the following description will be given taking an example of use in an automobile.

[0015] A vehicle using the image recognition system 1 may be a vehicle capable of autonomous driving (hereinafter referred to as an autonomous vehicle). There may be multiple levels of autonomous driving (hereinafter referred to as automation levels), as defined by, for example, the SAE. Automation levels are classified, for example, as follows: LV0 to LV5. LV0 is a level where the driver performs all driving tasks without system intervention. The driving tasks may also be referred to as dynamic driving tasks. Examples of driving tasks include steering, acceleration / deceleration, and periphery monitoring. LV0 corresponds to so-called manual driving. LV1 is a level where the system assists with either steering or acceleration / deceleration. LV1 corresponds to so-called driver assistance. LV2 is a level where the system assists with both steering and acceleration / deceleration. LV2 corresponds to so-called partial driving automation. Note that LV1 and LV2 are also considered to be part of autonomous driving. LV3 autonomous driving is a level where the system can perform all driving tasks under certain conditions, and the driver performs driving operations in an emergency. LV4 autonomous driving is a level at which the system can perform all driving tasks except under specific circumstances such as incompatible roads or extreme environments. LV4 corresponds to so-called highly automated driving. LV5 autonomous driving is a level at which the system can perform all driving tasks in any environment. LV5 corresponds to so-called fully automated driving. The following explanation will be given using as an example a case where a vehicle using the image recognition system 1 has an automation level of at least LV1 or higher.

[0016] The locator 11 includes a GNSS (Global Navigation Satellite System) receiver and an inertial sensor. The GNSS receiver receives positioning signals from multiple positioning satellites. The inertial sensor includes, for example, a gyro sensor and an acceleration sensor. The locator 11 sequentially determines the vehicle position of the vehicle (hereinafter referred to as the vehicle position) by combining the positioning signals received by the GNSS receiver with the measurement results of the inertial sensor. The vehicle position may be expressed, for example, in latitude and longitude coordinates. Note that the vehicle position may also be determined using a travel distance calculated from signals sequentially output from a vehicle speed sensor mounted on the vehicle.

[0017] The map DB 12 is a non-volatile memory that stores map data used for route guidance in the navigation device. The map data used for route guidance includes link data, node data, etc. The link data includes a link ID that identifies a link, a link length that indicates the length of the link, a link direction, a link travel time, link shape information, node coordinates (latitude / longitude) of the start and end of the link, and road attributes. The road attributes include a road name, a road type, a road width, and a speed limit. The node data includes a node ID, which is assigned a unique number for each node on the map, node coordinates, a node name, a node type, a connecting link ID that describes the link ID of a link connecting to the node, and an intersection type. The map DB 12 may store high-precision map data. The high-precision map data is map data with higher precision than the map data used for route guidance. The high-precision map data includes information that can be used for driving assistance, such as three-dimensional shape information of roads, information on the number of lanes, and information indicating the allowed traveling direction for each lane.

[0018] The vehicle condition sensor 13 is a group of sensors for detecting various conditions of the vehicle. The vehicle condition sensor 13 includes a vehicle speed sensor. The vehicle speed sensor detects the speed of the vehicle. The vehicle condition sensor 13 outputs the detected sensing information to an in-vehicle LAN. The sensing information detected by the vehicle condition sensor 13 may be configured to be output to the in-vehicle LAN via an ECU installed in the vehicle.

[0019] The perimeter monitoring sensor 14 monitors the environment surrounding the vehicle. As an example, the perimeter monitoring sensor 14 detects obstacles around the vehicle, such as moving objects such as pedestrians and other vehicles, and stationary objects such as fallen objects on the road. The perimeter monitoring sensor 14 also detects road markings such as lane markings around the vehicle. The perimeter monitoring sensor 14 includes a perimeter monitoring camera 141. The perimeter monitoring camera 141 sequentially captures and outputs captured images as sensing information. The captured images sequentially output from the perimeter monitoring camera 141 are, more specifically, image data as captured image data. Hereinafter, the captured images sequentially output from the perimeter monitoring camera 141 are referred to as perimeter image data. The perimeter monitoring camera 141 may be a plurality of cameras with different imaging ranges. The perimeter monitoring sensor 14 may include a search wave sensor in addition to the perimeter monitoring camera 141. Examples of search wave sensors include millimeter-wave radar, sonar, and LiDAR (Light Detection and Ranging / Laser Imaging Detection and Ranging). The exploration wave sensor sequentially outputs, as sensing information, scanning results based on received signals obtained when receiving waves reflected by an obstacle.

[0020] The vehicle control ECU 15 is an electronic control device that controls the driving of the vehicle. Examples of driving control include acceleration / deceleration control and / or steering control. The vehicle control ECU 15 includes a steering ECU that controls steering, a power unit control ECU that controls acceleration / deceleration, and a brake ECU. The vehicle control ECU 15 controls driving by outputting control signals to each driving control device mounted on the vehicle. Examples of driving control devices include an electronically controlled throttle, a brake actuator, and an EPS (Electric Power Steering) motor.

[0021] The driving assistance ECU 16 is an electronic control device that provides driving assistance for the vehicle. The driving assistance ECU 16 executes processes related to driving assistance based on signals input from the various in-vehicle devices described above. The driving assistance ECU 16 executes acceleration / deceleration control, steering control, and the like of the vehicle in cooperation with the vehicle control ECU 15. Examples of driving assistance include adaptive cruise control (ACC), pre-collision safety (PCS), and automatic emergency braking (AEB).

[0022] The interior camera 17 captures an image of a predetermined range within the vehicle interior. The interior camera 17 captures an image of an area including at least the driver's seat of the vehicle. The interior camera 17 is composed of, for example, a near-infrared light source, a near-infrared camera, and a control unit that controls them. The interior camera 17 uses the near-infrared camera to capture an image of the driver illuminated with near-infrared light by the near-infrared light source. The image captured by the near-infrared camera is analyzed by the control unit. The control unit analyzes the captured image to detect the driver's state, such as the direction of the driver's face and line of sight. The interior camera 17 sequentially outputs the detected state of the driver to the HCU 19.

[0023] The presentation device 18 is provided in the vehicle and presents information to the interior of the vehicle. That is, the presentation device 18 presents information to the driver of the vehicle. The presentation device 18 presents information according to instructions from the HCU 19. The presentation device 18 includes a display device 181. The display device 181 presents information by displaying information. Examples of the display device 181 include a meter MID (Multi Information Display), a CID (Center Information Display), and a HUD (Head-Up Display). The meter MID is a display device provided in front of the driver's seat in the interior of the vehicle. As an example, the meter MID may be provided in a meter panel. The CID is a display device located in the center of the instrument panel of the vehicle. The HUD is provided in the interior of the vehicle, for example, on the instrument panel. The HUD projects a display image formed by a projector onto a predetermined projection area on the front windshield, which serves as a projection member. The HUD may be configured to project a display image onto a combiner provided in front of the driver's seat instead of onto the front windshield. The presentation device 18 may include an audio output device that presents information by outputting sound.

[0024] The HCU 19 is an electronic control unit that executes various processes related to the interaction between the occupant and the vehicle's system. The HCU 19 causes the presentation device 18 to present information. The HCU 19 acquires the driver's state detected by the interior camera 17. The HCU 20 may identify the driver's state from an image captured by the interior camera 17. In other words, the HCU 19 may perform part of the function of the control unit of the interior camera 17.

[0025] The image recognition device 10 is mainly composed of a computer including, for example, a processor, volatile memory, nonvolatile memory, I / O, and a bus connecting these. The image recognition device 10 performs image recognition processing by executing a control program stored in the nonvolatile memory. The image recognition device 10 performs an image recognition task (hereinafter referred to as an image recognition task) on an image captured by a perimeter monitoring camera 141 and recognizes an object according to the image recognition task. For example, if the image recognition task is semantic segmentation, class identification is performed to segment the image into regions by class. In this case, the class is a semantic unit, such as "road," "person," or "bicycle." If the image recognition task is traffic light detection, the color and flashing state of the traffic light are recognized. If the image recognition task is branch road detection, the branch road is recognized. The image recognition task may also recognize objects other than those described above from an image. The image recognition device 10 performs multiple image recognition tasks. In other words, the image recognition device 10 performs multitasking processing. The configuration of the image recognition device 10 will be described in detail below.

[0026] <General Configuration of Image Recognition Device 10> Next, the general configuration of the image recognition device 10 will be described with reference to Figure 2. As shown in Figure 2, the image recognition device 10 includes an image acquisition unit 101, a vehicle-related acquisition unit 102, a detector 103, and a controller unit 104 as functional blocks. Execution of processing by each functional block of the image recognition device 10 by a computer corresponds to execution of an image recognition method. Note that some or all of the functions executed by the image recognition device 10 may be configured as hardware using one or more ICs or the like. Furthermore, some or all of the functional blocks included in the image recognition device 10 may be realized by a combination of software execution by a processor and hardware components.

[0027] The image acquisition unit 101 acquires surrounding image data sequentially output from the surrounding monitoring camera 141. In the example of the present embodiment, a case where data of surrounding images captured by the surrounding monitoring camera 141 is used for image recognition will be described as an example, but this is not necessarily limited to this. For example, a configuration may be adopted in which sensing results detected by another surrounding monitoring sensor 14 that can be used for image recognition, such as LiDAR, are used for image recognition. In this case, this sensing result may also be included in the surrounding image data. The vehicle-related acquisition unit 102 acquires information related to the vehicle (hereinafter, vehicle-related information) other than the surrounding image data. Examples of vehicle-related information include information on the vehicle speed of the vehicle, map information, information on the driver's state, and information on sensor characteristics. Information on the vehicle speed of the vehicle will be referred to as vehicle speed information hereinafter.

[0028] The vehicle-related acquisition unit 102 may acquire vehicle speed information from a vehicle speed sensor among the vehicle state sensors 13. The vehicle-related acquisition unit 102 may acquire map information from the map DB 12. The vehicle-related acquisition unit 102 may acquire map information limited to the area around the vehicle's position measured by the locator 11. The vehicle-related acquisition unit 102 may acquire the driver's state from the HCU 19. For example, the driver's state may be the line of sight direction detected using the interior camera 17. The vehicle-related acquisition unit 102 may acquire sensor characteristics from the perimeter monitoring sensor 14. The non-volatile memory of the perimeter monitoring sensor 14 may be configured to store sensor characteristics for each sensor included in the perimeter monitoring sensor 14 in advance. For example, the sensor characteristics may be data indicating difficult objects and difficult situations for each sensor included in the perimeter monitoring sensor 14. The difficult objects are objects that are difficult to detect due to the characteristics of the sensor's detection principle. The difficult situations indicate situations in which object detection performance may deteriorate. Note that the object that is difficult to detect may include an object that is likely to be mistaken for another type of object, or an object for which the detection result is unstable.

[0029] The detector 103 executes a plurality of image recognition tasks on the peripheral image acquired by the image acquisition unit 101. In other words, the detector 103 is capable of multitasking the peripheral image acquired by the image acquisition unit 101. This detector 103 corresponds to an image processing unit. Furthermore, the processing by this detector 103 corresponds to an image processing step. The detector 103 executes a plurality of image recognition tasks on the peripheral image, thereby recognizing a recognition target for each image recognition task from the peripheral image. This recognition may also be referred to as detection.

[0030] The detector 103 may execute multiple image recognition tasks using a machine learning model. This machine learning model is a model generated by performing machine learning so that peripheral images are input and a recognition target for each of the multiple image recognition tasks can be output. The detector 103 may execute multiple image recognition tasks using a neural network (hereinafter referred to as NN), which is one of the machine learning models. The detector 103 may execute multiple image recognition tasks using a machine learning model other than a network structure such as a NN. For example, a random forest, which is a tree-structured machine learning model, may be used. The following description will be continued using an example in which a NN is used as the detector 103. The detector 103 is capable of dynamically changing the processing content of the multiple image recognition tasks. In this embodiment, the detector 103 is capable of dynamically changing the network structure and parameters of the NN. The parameters are, for example, at least one of the weights and biases of each layer in the NN. In this embodiment, the processing content of the multiple image recognition tasks corresponds to the network structure and weights of the NN.

[0031] The controller unit 104 adjusts the processing contents of the multiple image recognition tasks in the detector 103. The controller unit 104 receives a peripheral image as an input and dynamically changes the processing contents of the multiple image recognition tasks in the detector 103 according to the tendency of the content of the peripheral image. The peripheral image input to the controller unit 104 may be the peripheral image acquired by the image acquisition unit 101. The processing in the controller unit 104 corresponds to a controller step.

[0032] According to the above configuration, the content of multitasking, which executes multiple image recognition tasks on a peripheral image, can be dynamically changed in response to trends in the content of the image. Trends in the content of peripheral images are highly correlated with changes in the scene depicted by the content of the peripheral images. Therefore, it is possible to dynamically change the balance between the processing speed and processing accuracy of the multiple image recognition tasks in response to the scene depicted by the image content. As a result, it is more easily possible to perform image recognition tasks with a more desirable balance of accuracy and speed in response to the scene. The controller unit 104 can dynamically change the processing content of the multiple image recognition tasks in the detector 103 by changing at least one of the network structure and parameters of the neural network.

[0033] Furthermore, with the above configuration, the processing content of the image recognition task is automatically switched for the peripheral image, so that the processing time margin when designing the detector 103 can be omitted. Therefore, when the processing accuracy is fixed, faster recognition processing is achieved. In addition, as a secondary effect, power consumption is reduced. In addition, because the processing content of the image recognition task is automatically switched for the peripheral image, unimportant processing can be reduced and more time can be spent on important processing. Therefore, when the processing time is fixed, more accurate recognition processing can be achieved.

[0034] The above-mentioned effects will now be described with reference to FIGS. 3 and 4. FIG. 3 is a diagram illustrating an example in which the content of multitasking cannot be dynamically changed. FIG. 4 is a diagram illustrating an example in which the content of multitasking can be dynamically changed. In FIGS. 3 and 4, an example in which semantic segmentation, traffic light detection, and branch road detection are performed as multiple image recognition tasks in multitasking is described. SS in FIGS. 3 and 4 indicates semantic segmentation among the multiple image recognition tasks. TL in FIGS. 3 and 4 indicates traffic light detection among the multiple image recognition tasks. Br in FIGS. 3 and 4 indicates branch road detection among the multiple image recognition tasks. PC in FIGS. 3 and 4 schematically illustrates the balance of performance and the amount of computation among the multiple image recognition tasks. Here, performance can be rephrased as processing accuracy. The ratio of each patterned area in PC indicates the balance of performance among the multiple image recognition tasks. Furthermore, the size of PC indicates the overall amount of computation among the multiple image recognition tasks. This amount of computation affects the processing speed of the image recognition tasks. NS in Figures 3 and 4 indicates the network structure of the NN. PB in Figures 3 and 4 indicates the processing blocks of the NN. De, IP, and HR in Figure 4 indicate different scenes. De is the default scene. IP is the scene of driving at an intersection. HR is the scene of driving on a highway. In the example of Figure 4, a scene that is neither driving at an intersection nor driving on a highway can be set as the default scene. In Figure 4, unused processing blocks are indicated by dashed lines, and used processing blocks are indicated by solid lines.

[0035] As shown in FIG. 3 , if the content of the multitasking cannot be dynamically changed, the processing speed and processing accuracy of multiple image recognition tasks cannot be changed regardless of the scene. On the other hand, as shown in FIG. 4 , if the content of the multitasking can be dynamically changed, the processing speed and processing accuracy of multiple image recognition tasks can be changed depending on the scene. For example, in a scene of driving at an intersection, it is possible to prioritize and improve the processing accuracy of semantic segmentation and traffic light detection, which are considered more necessary for driving at an intersection, over branch road detection. In addition, in a scene of driving on a highway, as shown in FIG. 4 , it is possible to prioritize and improve the processing accuracy of semantic segmentation and branch road detection over the processing accuracy of traffic light detection, which are considered less necessary for driving on a highway. Furthermore, in a scene of driving on a highway with fewer external disturbances, it is also possible to change the processing speed of multiple image recognition tasks so as to reduce the overall amount of calculation compared to other scenes.

[0036] The controller unit 104 may use a machine learning model to change the processing speed and processing accuracy of multiple image recognition tasks according to the scene. This machine learning model may be a machine learning model that learns a neural network (NN) network structure and parameters that balance the processing speed and processing accuracy of multiple image recognition tasks according to the trends in the content of the surrounding images, based on the trends in the content of the surrounding images. This learning may be performed so as to minimize the loss of accuracy calculated from the detection results of each image recognition task and the amount of calculation calculated from the network configuration. This machine learning model may be realized, for example, by a hypernetwork such as a convolutional neural network (CNN).

[0037] Here, learning for minimizing the accuracy loss calculated from the detection results of each image recognition task and the amount of calculation calculated from the network configuration will be described with reference to Fig. 5. Fig. 5 is a diagram for explaining an example of learning by the controller unit 104. The calculation amount calculation unit 105, calculation amount table 106, accuracy loss calculation unit 107, and correct answer label 108 in Fig. 5 may be provided as functional blocks in the image recognition device 10.

[0038] The computation amount calculation unit 105 calculates the computation amount for the NN of the detector 103 generated by the controller unit 104. The computation amount calculation unit 105 calculates the computation amount by referring to a computation amount table 106. The computation amount table 106 may be a database that stores in advance the computation amount for each unit such as a node or edge of the network structure. This computation amount can also be referred to as the computation amount for each layer of the NN. The computation amount may include the amount of data communication between hardware. The computation amount table 106 may be realized using, for example, a non-volatile memory. The computation amount calculation unit 105 may calculate the computation amount of the NN by referring to the computation amount table 106 and adding up the computation amount for each unit that makes up the network structure that is the target of computation amount calculation.

[0039] The accuracy loss calculation unit 107 calculates the accuracy loss in recognition using the NN of the detector 103 from the detection results of the detector 103. The accuracy loss calculation unit 107 calculates the accuracy loss by referring to the correct answer labels 108. The correct answer labels 108 may be a database that stores in advance the correct recognition results for each surrounding image used for learning. The accuracy loss calculation unit 107 may calculate the accuracy loss by referring to the correct answer labels 108 and depending on how accurate the detection results of the detector 103 were.

[0040] 5, the amount of computation and loss of accuracy of the NN are calculated while changing the network structure and parameters of the NN of the detector 103 generated by the controller unit 104. Then, the network structure and parameters that minimize the amount of computation and loss of accuracy of the NN are learned according to the tendency of the content of the surrounding images used for learning. This enables the controller unit 104 to generate a NN network structure and parameters that can balance the processing speed and processing accuracy of the image recognition task according to the tendency of the content of the surrounding images.

[0041] The controller unit 104 may dynamically change the processing content of the multiple image recognition tasks in the detector 103 in accordance with the tendency of the content of the peripheral image so as to maximize the processing accuracy of each of the multiple image recognition tasks within a given processing time constraint. This may be achieved using learning results obtained by learning the processing content of the image recognition tasks that maximize the processing accuracy of each of the multiple image recognition tasks within a given processing time constraint in accordance with the tendency of the content of the peripheral image. This makes it easier to perform the image recognition tasks in accordance with the scene in such a way that the processing accuracy of each of the multiple image recognition tasks maximizes within a given processing time constraint.

[0042] The controller unit 104 may dynamically change the processing content of the multiple image recognition tasks in the detector 103 in accordance with the tendency of the content of the peripheral image so as to minimize the sum of the processing speeds of the multiple image recognition tasks within a given processing accuracy constraint. This may be achieved using a learning result obtained by learning the processing content of the image recognition tasks that minimize the sum of the processing speeds of the multiple image recognition tasks within a given processing accuracy constraint in accordance with the tendency of the content of the peripheral image. This makes it easier to cause the image recognition tasks to be performed in accordance with the scene in such a way that the sum of the processing speeds of the multiple image recognition tasks within a given processing speed constraint.

[0043] The controller unit 104 may dynamically change the processing content of the multiple image recognition tasks in the detector 103 in accordance with the trends in the content of the peripheral image so as to minimize the total amount of hardware resource usage for each of the multiple image recognition tasks within a given processing accuracy constraint. This may be achieved using a learning result obtained by learning the processing content of the image recognition tasks that minimize the total amount of hardware resource usage for each of the multiple image recognition tasks within a given processing accuracy constraint in accordance with the trends in the content of the peripheral image. This makes it easier to perform the image recognition tasks in accordance with the scene in such a way that the total amount of hardware resource usage for each of the multiple image recognition tasks within a given processing accuracy constraint is minimized. The hardware resource may be, for example, a memory. The hardware resource may include a processor, storage, etc.

[0044] The controller unit 104 preferably includes a scene classification unit 1041 as a sub-functional block. The scene classification unit 1041 may be configured separately from the controller unit 104. The scene classification unit 1041 receives a peripheral image as input and classifies the scene indicated by the content of the peripheral image. The controller unit 104 preferably dynamically changes the processing content of multiple image recognition tasks in the detector 103 according to the scene classified by the scene classification unit 1041 as a tendency of the content of the peripheral image. This enables the image recognition task to be performed with a more accurate balance between accuracy and speed according to the scene. The scene classification unit 1041 may be rule-based or learning-based. If learning-based, the scene may be classified from the peripheral image using a machine learning model. For example, a neural network (NN) trained to classify scenes from the peripheral image may be used as the machine learning model. Examples of scenes to be classified include highways, parking lots, and areas around intersections.

[0045] The controller unit 104 may be configured to perform processing in response to an input other than a peripheral image. Examples of input other than a peripheral image include vehicle-related information and time-series information acquired by the vehicle-related acquisition unit 102. The time-series information may be the detection result of the detector 103 for the peripheral image of the previous frame. The controller unit 104 may also use the vehicle-related information and the time-series information to dynamically change the processing content of the multiple image recognition tasks in the detector 103. The controller unit 104 may also dynamically change the processing content of the multiple image recognition tasks in the detector 103 in response to the vehicle-related information and the time-series information. In this case, the controller unit 104 may dynamically change the processing content of the multiple image recognition tasks in the detector 103 based on the learning result of learning the balance of the processing of the multiple image recognition tasks in response to the vehicle-related information and the time-series information.

[0046] The controller unit 104 may dynamically change the processing content of the multiple image recognition tasks in the detector 103 in accordance with the vehicle speed information acquired by the vehicle-related acquisition unit 102. For example, the controller unit 104 may change the processing speed of the multiple image recognition tasks so that it is faster as the vehicle speed increases than when the vehicle speed is slower. The faster the vehicle speed, the greater the changes in the surrounding image in a short period of time, and therefore a higher processing speed is required. The above configuration makes it easy to meet this demand.

[0047] The controller unit 104 preferably also uses map information acquired by the vehicle-related acquisition unit 102. For example, the controller unit 104 may use the map information to reinforce or correct the scene classification by the scene classification unit 1041. With the above configuration, it is possible to more accurately execute an image recognition task with a more preferable balance between accuracy and speed depending on the scene.

[0048] The controller unit 104 may dynamically change the processing content of the multiple image recognition tasks in the detector 103 in accordance with the driver's state acquired by the vehicle-related acquisition unit 102. The controller unit 104 may change the processing content of the image recognition tasks in directions other than the driver's line of sight to increase the processing accuracy. For example, if the driver's line of sight is to the right, the processing accuracy of the image recognition tasks for recognizing objects to the left and in front may be increased. In addition, when the detector 103 performs image recognition for each imaging direction, the processing content of the multiple image recognition tasks may be changed to increase the processing accuracy of peripheral images in directions other than the driver's line of sight. With the above configuration, it is possible to prioritize increasing the accuracy of image recognition in areas other than the driver's line of sight, making it easier to entrust driving assistance to the vehicle's system.

[0049] The controller unit 104 may dynamically change the processing content of the multiple image recognition tasks in the detector 103 in accordance with the sensor characteristics acquired by the vehicle-related acquisition unit 102. For example, the controller unit 104 may change the processing content of the multiple image recognition tasks in the detector 103 to increase the processing accuracy of the multiple image recognition tasks in a scene that is a poor situation for the perimeter monitoring sensors 14 other than the perimeter monitoring camera 141. The controller unit 104 may determine whether the scene is a poor situation based on the scene classified by the scene classification unit 1041 and the sensor characteristics. With the above configuration, it becomes easier to compensate for deterioration in detection accuracy by the perimeter monitoring sensors 14 that are in a poor situation in sensor fusion.

[0050] The controller unit 104 may dynamically change the processing content of the multiple image recognition tasks in the detector 103 according to the time-series information. The controller unit 104 may dynamically change the processing content of the multiple image recognition tasks in the detector 103 according to the detection result of the detector 103 for the peripheral image of the previous frame. The controller unit 104 may make the change depending on whether the detection result is estimated to be difficult to recognize or easy to recognize. Whether recognition is difficult or easy may be determined by the number of recognition objects, such as pedestrians and vehicles. If the detection result is estimated to be difficult to recognize, the controller unit 104 may change the processing content of the multiple image recognition tasks to be appropriate for cases where recognition is estimated to be difficult. The controller unit 104 may learn the processing content of the multiple image recognition tasks according to the difficulty of recognition through machine learning. The above configuration makes it possible to perform processing of image recognition tasks appropriate for the difficulty of image recognition.

[0051] Additionally, the controller unit 104 may turn off unnecessary image recognition tasks depending on the scene. For example, on a highway where pedestrians are not expected to be present, the controller unit 104 may turn off an image recognition task for detecting pedestrians.

[0052] The controller unit 104 may perform processing related to outputs other than control of the detector 103. An example of this processing will be described below. The controller unit 104 may request the driver to decelerate or perform deceleration control when the processing load of the detector 103 exceeds a specified value. This allows the processing load of the detector 103 to be reduced by decelerating the host vehicle. The processing load of the detector 103 exceeds the specified value when the machine learning model of the detector 103 controlled by the controller unit 104 can no longer satisfy the constraints of processing time and processing accuracy set during learning. The deceleration request may be performed by the presentation device 18. The deceleration control may be performed by the driving assistance ECU 16. When deceleration control is performed, the presentation device 18 may also present the reason for deceleration. This makes it possible to reduce the anxiety of the vehicle occupants regarding the deceleration control.

[0053] The controller unit 104 may instruct the periphery monitoring camera 141 to lengthen the image capturing cycle when the processing load of the controller unit 104 exceeds a specified value. This makes it possible to reduce the processing load of the controller unit 104. The controller unit 104 may change the image capturing cycle of the periphery monitoring camera 141 depending on the scene classified by the scene classification unit 1041. For example, in a simple scene with little disturbance, such as a highway, the controller unit 104 may instruct the periphery monitoring camera 141 to lengthen the image capturing cycle. On the other hand, in a scene where recognition processing is difficult, the controller unit 104 may instruct the periphery monitoring camera 141 to lengthen the image capturing cycle.

[0054] The controller unit 104 may instruct the periphery monitoring camera 141 to lower the resolution of the peripheral image when the processing load of the controller unit 104 exceeds a specified value. This makes it possible to reduce the processing load of the controller unit 104. The controller unit 104 may change the resolution of the periphery monitoring camera 141 depending on the scene classified by the scene classification unit 1041. For example, in a simple scene with few disturbances, such as a highway, the controller unit 104 may instruct the periphery monitoring camera 141 to lower the resolution. On the other hand, in a scene where recognition processing is difficult, the controller unit 104 may instruct the periphery monitoring camera 141 to increase the resolution.

[0055] (Embodiment 2) The configuration of the embodiment is not limited to the above-described embodiment, and may be the following configuration of embodiment 2. An example of the configuration of embodiment 2 will be described below with reference to the drawings. The image recognition system 1 of embodiment 2 is similar to the image recognition system 1 of embodiment 1, except that it includes an image recognition device 10a instead of the image recognition device 10.

[0056] <Schematic Configuration of Image Recognition Device 10a> Next, the schematic configuration of the image recognition device 10a will be described with reference to Fig. 6. As shown in Fig. 6, the image recognition device 10a includes an image acquisition unit 101, a vehicle-related acquisition unit 102, a detector 103, and a controller unit 104a as functional blocks. The image recognition device 10a is similar to the image recognition device 10 of the first embodiment except that the image recognition device 10a includes the controller unit 104a instead of the controller unit 104. Furthermore, the execution of processing by a computer of each functional block of the image recognition device 10a corresponds to the execution of an image recognition method.

[0057] The controller unit 104a has a scene classification unit 1041 and an uncertainty prediction unit 1042 as sub-functional blocks. The controller unit 104a is similar to the controller unit 104 of the first embodiment except that it has the uncertainty prediction unit 1042. Note that the uncertainty prediction unit 1042 may be configured to be provided separately from the controller unit 104a. The uncertainty prediction unit 1042 corresponds to a first uncertainty prediction unit.

[0058] The uncertainty prediction unit 1042 predicts data uncertainty (Aleatoric uncertainty). The uncertainty prediction unit 1042 may predict the data uncertainty using, for example, Bayesian estimation. In a configuration in which the image recognition device 10a is equipped with a scene classification unit 1041, the uncertainty prediction unit 1042 predicts the uncertainty of scenes classified by the scene classification unit 1041. In this case, the data uncertainty becomes the uncertainty of scenes classified by the scene classification unit 1041. Scene uncertainty can be rephrased as the difficulty of scene classification. In a configuration in which the image recognition device 10a is not required to be equipped with a scene classification unit 1041, the uncertainty prediction unit 1042 may predict the uncertainty of the image recognition task performed by the detector 103 controlled by the controller unit 104a from the tendency of the content of the surrounding images.

[0059] The controller unit 104a dynamically changes the processing content of the multiple image recognition tasks in the detector 103, also using the data uncertainty predicted by the uncertainty prediction unit 1042. When the scene classification unit 1041 is a required component, scene uncertainty is used as the data uncertainty. The controller unit 104a may dynamically change the processing content of the multiple image recognition tasks in the detector 103 according to the degree of uncertainty. The degree of uncertainty may be divided into two levels: a high uncertainty level and a low uncertainty level, separated by a predetermined threshold. The controller unit 104a may change the processing content of the multiple image recognition tasks appropriate for each level of uncertainty. The controller unit 104a may learn the processing content of the multiple image recognition tasks appropriate for each level of uncertainty through machine learning. The above configuration makes it possible to perform image recognition task processing appropriate for the data uncertainty. Note that in the second embodiment, the controller unit 104a may be configured without the scene classification unit 1041. In this case, the uncertainty prediction unit 1042 only needs to predict the uncertainty of the data input from the image acquisition unit 101 to the controller unit 104b.

[0060] (Embodiment 3) The configuration is not limited to the above-described embodiments, and may be the following configuration of embodiment 3. An example of the configuration of embodiment 3 will be described below with reference to the drawings. The image recognition system 1 of embodiment 3 is similar to the image recognition system 1 of embodiment 1, except that it includes an image recognition device 10b instead of the image recognition device 10.

[0061] <Schematic Configuration of Image Recognition Device 10b> Next, the schematic configuration of the image recognition device 10b will be described with reference to Fig. 7. As shown in Fig. 7, the image recognition device 10b includes an image acquisition unit 101, a vehicle-related acquisition unit 102, a detector 103, and a controller unit 104b as functional blocks. The image recognition device 10b is similar to the image recognition device 10 of the first embodiment except that the image recognition device 10b includes the controller unit 104b instead of the controller unit 104. Furthermore, the execution of processing by a computer of each functional block of the image recognition device 10b corresponds to the execution of an image recognition method.

[0062] The controller unit 104b has a scene classification unit 1041 and an uncertainty prediction unit 1042b as sub-functional blocks. The controller unit 104b is similar to the controller unit 104 of embodiment 1 except that it has the uncertainty prediction unit 1042b. Note that the uncertainty prediction unit 1042b may be configured to be provided separately from the controller unit 104b. The uncertainty prediction unit 1042b corresponds to a second uncertainty prediction unit. In embodiment 3, it is not essential that the controller unit 104b has the scene classification unit 1041.

[0063] The uncertainty prediction unit 1042b predicts data uncertainty (Aleatoric uncertainty) and model uncertainty (Epistemic uncertainty). The uncertainty prediction unit 1042b may predict data uncertainty in the same manner as the uncertainty prediction unit 1042. If the controller unit 104b includes a scene classification unit 1041, the uncertainty prediction unit 1042b may predict scene uncertainty in the same manner as the uncertainty prediction unit 1042 of the second embodiment. If the controller unit 104b does not include a scene classification unit 1041, the uncertainty prediction unit 1042b may predict the uncertainty of data input to the controller unit 104b from the image acquisition unit 101. The uncertainty prediction unit 1042b may predict model uncertainty using, for example, probabilistic modeling. The model uncertainty is the uncertainty of a machine learning model of the detector 103 controlled by the controller unit 104b. For example, in the example of this embodiment, it may be the uncertainty of semantic segmentation, traffic light detection, and branch road detection in the machine learning model for the image input from the image acquisition unit 101.

[0064] The controller unit 104b dynamically changes the processing content of the multiple image recognition tasks in the detector 103 using the data uncertainty and model uncertainty predicted by the uncertainty prediction unit 1042b. The controller unit 104b may dynamically change the processing content of the multiple image recognition tasks in the detector 103 according to the degree of uncertainty of the data and the model. The degree of uncertainty of the data and the model may be two-level, i.e., a high uncertainty level and a low uncertainty level, as described in the second embodiment. The controller unit 104b may change the processing content of the multiple image recognition tasks appropriate for each combination of the degree of uncertainty of the data and the model. The controller unit 104a may learn, by machine learning, the processing content of the multiple image recognition tasks appropriate for each combination of the degree of uncertainty of the data and the model. With the above configuration, it is possible to perform processing of the image recognition tasks appropriate for the data uncertainty and the model uncertainty.

[0065] The uncertainty prediction unit 1042b may be configured to predict only the model uncertainty out of the data uncertainty and the model uncertainty. In this case, the controller unit 104b may be configured to dynamically change the processing content of multiple image recognition tasks in the detector 103 according to the degree of model uncertainty. This configuration also makes it possible to perform image recognition task processing appropriate for the model uncertainty. An example of image recognition task processing appropriate for the model uncertainty is processing that allocates more resources to the more difficult task among semantic segmentation, traffic light detection, and branch road detection.

[0066] (Fourth Embodiment) In the above-described embodiment, the image recognition devices 10, 10a, and 10b are provided in a vehicle, but this is not necessarily limited to this. The image recognition devices 10, 10a, and 10b may be provided outside the vehicle. For example, they may be provided in a server outside the vehicle. In this case, communication between the vehicle-side system and the image recognition devices 10, 10a, and 10b on the server may be performed via a communication module provided in the vehicle.

[0067] (Embodiment 5) In the above-described embodiments, the image recognition devices 10, 10a, and 10b have been described as being used for image recognition of a peripheral image captured by a vehicle's peripheral monitoring camera 141, but this is not necessarily limited to this. The image recognition devices 10, 10a, and 10b may be configured to be used for image recognition of a peripheral image other than that captured by a vehicle's peripheral monitoring camera 141. For example, they may be used for image recognition of a peripheral image captured by a camera on a moving object such as a drone. Alternatively, they may be used for image recognition of a peripheral image captured by a camera installed in a facility. Furthermore, in the above-described embodiments, an example has been given in which a peripheral image is used as the image used for image recognition, but this is not necessarily limited to this. The image used for image recognition may be an image other than a peripheral image, as long as the content of the image has a tendency to be correlated with the scene.

[0068] Sixth Embodiment In the above-described embodiment, the controller units 104, 104a, and 104b control the network structure of the detector 103 to dynamically change the processing of multiple image recognition tasks in accordance with trends in the content of surrounding images. However, this is not necessarily limited to this. For example, the detector 103 may be configured to be prepared in advance, with multiple detectors 103 having different processing patterns for multiple image recognition tasks, based on human design. The controller units 104, 104a, and 104b may then dynamically change the processing of the multiple image recognition tasks by selecting one of the multiple detectors 103. Note that the network structure and parameters of the multiple detectors 103 prepared in advance may be learned during the learning of the controller unit 104 described in FIG. 5.

[0069] The present disclosure is not limited to the above-described embodiments, and various modifications are possible within the scope of the claims. Embodiments obtained by appropriately combining the technical means disclosed in different embodiments are also within the technical scope of the present disclosure. Furthermore, the control unit and method described in the present disclosure may be implemented by a special-purpose computer comprising a processor programmed to execute one or more functions embodied in a computer program. Alternatively, the apparatus and method described in the present disclosure may be implemented by a special-purpose hardware logic circuit. Alternatively, the apparatus and method described in the present disclosure may be implemented by one or more special-purpose computers configured by combining a processor executing a computer program with one or more hardware logic circuits. Furthermore, the computer program may be stored as instructions executed by a computer on a computer-readable non-transitory tangible recording medium.

[0070] (Disclosed Technical Ideas) This specification discloses multiple technical ideas described in the following multiple clauses. Some clauses may be described in a multiple dependent form, with the subsequent clause alternatively referring to the preceding clause. Furthermore, some clauses may be described in a multiple dependent form, referring to another multiple dependent clause. These multiple dependent clauses define multiple technical ideas.

[0071] (Technical Idea 1) An image recognition device comprising: an image processing unit (103) capable of multitasking to execute a plurality of image recognition tasks on an image and capable of adjusting the processing content of the image recognition tasks; and a controller unit (104, 104a, 104b) that adjusts the processing content of the plurality of image recognition tasks in the image processing unit, wherein the controller unit receives the image as input and dynamically changes the processing content of the plurality of image recognition tasks in the image processing unit according to the tendency of the content of the image.

[0072] (Technical Idea 2) An image recognition device according to Technical Idea 1, wherein the controller unit dynamically changes the processing content of a plurality of image recognition tasks in the image processing unit in accordance with the tendency of the content of the image so as to maximize the processing accuracy of each of the plurality of image recognition tasks within a given processing time constraint.

[0073] (Technical Idea 3) An image recognition device according to Technical Idea 1, wherein the controller unit dynamically changes the processing content of a plurality of image recognition tasks in the image processing unit in accordance with the tendency of the content of the image so as to minimize the sum of the processing speeds of the plurality of image recognition tasks within a given processing accuracy constraint.

[0074] (Technical Idea 4) An image recognition device according to Technical Idea 1, wherein the controller unit dynamically changes the processing content of a plurality of image recognition tasks in the image processing unit in accordance with the tendency of the content of the image so as to minimize the total usage of hardware resources for each of the plurality of image recognition tasks within a given processing accuracy constraint.

[0075] (Technical Idea 5) An image recognition device according to any one of Technical Ideas 1 to 4, wherein the controller unit has a scene classification unit (1041) that receives the image as an input and classifies the scene indicated by the content of the image, and the controller unit dynamically changes the processing content of a plurality of image recognition tasks in the image processing unit according to the scene classified by the scene classification unit as a tendency of the content of the image.

[0076] (Technical Idea 6) An image recognition device according to Technical Idea 5, wherein the controller unit further has a first uncertainty prediction unit (1042) that predicts the uncertainty of a scene to be classified by the scene classification unit, and the controller unit dynamically changes the processing content of multiple image recognition tasks in the image processing unit, also using the uncertainty of the scene predicted by the first uncertainty prediction unit.

[0077] (Technical Idea 7) An image recognition device according to any one of Technical Ideas 1 to 5, wherein the image processing unit performs the multitasking processing using a machine learning model, the controller unit has a second uncertainty prediction unit (1042b) that predicts at least one of uncertainty regarding the image input to the controller unit and uncertainty of the machine learning model, and the controller unit dynamically changes the processing content of multiple image recognition tasks in the image processing unit using the uncertainty predicted by the second uncertainty prediction unit.

[0078] (Technical Idea 8) An image recognition device according to any one of Technical Ideas 1 to 7, wherein the image processing unit performs the multitasking processing using a neural network of a machine learning model, and the controller unit dynamically changes the processing content of a plurality of image recognition tasks in the image processing unit by changing at least one of the network structure and parameters of the neural network.

[0079] (Technical Idea 9) An image recognition device according to any one of Technical Ideas 1 to 8, wherein the image processing unit is capable of multitasking to execute a plurality of image recognition tasks on a peripheral image, which is an image captured by a peripheral monitoring camera (141) that captures an image of the periphery of the vehicle, and the controller unit receives the peripheral image as an input and dynamically changes the processing content of the plurality of image recognition tasks in the image processing unit according to the tendency of the content of the peripheral image.

Claims

1. An image processing unit (103) capable of multitasking, which executes multiple image recognition tasks on an image, and which can adjust the processing content of the image recognition tasks, The system includes controller units (104, 104a, 104b) that adjust the processing content of multiple image recognition tasks in the image processing unit, The controller unit takes the image as input and dynamically changes the processing content of multiple image recognition tasks in the image processing unit according to the trends in the content of the image. The controller unit is A scene classification unit (1041) takes the aforementioned image as input and classifies the scene indicated by the content of the image, The scene classification unit has a first uncertainty prediction unit (1042) that predicts the uncertainty of the scenes to be classified, As a trend in the content of the aforementioned images, the processing content of multiple image recognition tasks in the image processing unit is dynamically changed according to the scene classified by the scene classification unit. An image recognition device that dynamically changes the processing content of multiple image recognition tasks in the image processing unit, using the uncertainty of the scene predicted by the first uncertainty prediction unit.

2. A device that can be used in a vehicle, An image processing unit (103) capable of multitasking, which executes multiple image recognition tasks on an image, and which can adjust the processing content of the image recognition tasks, The system includes controller units (104, 104a, 104b) that adjust the processing content of multiple image recognition tasks in the image processing unit, The controller unit takes the image as input and dynamically changes the processing content of multiple image recognition tasks in the image processing unit according to the trends in the content of the image. The system includes a vehicle-related information acquisition unit that acquires vehicle-related information other than the aforementioned image, The aforementioned vehicle-related information acquisition unit acquires at least the vehicle speed information of the vehicle, The controller unit is an image recognition device that, taking the image as input, dynamically changes the processing content of multiple image recognition tasks in the image processing unit in accordance with the trends in the content of the image, as well as the vehicle speed information acquired by the vehicle-related information acquisition unit.

3. A device that can be used in a vehicle, An image processing unit (103) capable of multitasking, which executes multiple image recognition tasks on an image, and which can adjust the processing content of the image recognition tasks, The system includes controller units (104, 104a, 104b) that adjust the processing content of multiple image recognition tasks in the image processing unit, The controller unit takes the image as input and dynamically changes the processing content of multiple image recognition tasks in the image processing unit according to the trends in the content of the image. The system includes a vehicle-related information acquisition unit that acquires vehicle-related information other than the aforementioned image, The aforementioned vehicle-related information acquisition unit acquires at least information on the driver status of the vehicle, The controller unit is an image recognition device that, taking the image as input, dynamically changes the processing content of multiple image recognition tasks in the image processing unit in accordance with the trends in the content of the image, as well as the driver status information acquired by the vehicle-related information acquisition unit.

4. An image recognition device according to claim 2 or 3, The controller unit has a scene classification unit (1041) that takes the image as input and classifies the scene indicated by the content of the image, The controller unit is an image recognition device that dynamically changes the processing content of multiple image recognition tasks in the image processing unit according to the scene classified by the scene classification unit as a trend in the content of the image.

5. The image recognition device according to claim 4, The controller unit further includes a first uncertainty prediction unit (1042) that predicts the uncertainty of the scenes to be classified by the scene classification unit, The controller unit is an image recognition device that dynamically changes the processing content of multiple image recognition tasks in the image processing unit, using the uncertainty of the scene predicted by the first uncertainty prediction unit.

6. An image recognition device according to any one of claims 1 to 3, The controller unit dynamically changes the processing content of multiple image recognition tasks in the image processing unit in accordance with the trends of the image content, so as to maximize the processing accuracy of each of the multiple image recognition tasks within the given processing time constraints.

7. An image recognition device according to any one of claims 1 to 3, The controller unit dynamically changes the processing content of multiple image recognition tasks in the image processing unit in accordance with the trends of the content of the image, so as to minimize the sum of the processing speeds of each of the multiple image recognition tasks with respect to a given processing accuracy constraint.

8. An image recognition device according to any one of claims 1 to 3, The controller unit dynamically changes the processing content of multiple image recognition tasks in the image processing unit in accordance with the trends of the image content, so as to minimize the total amount of hardware resources used by each of the multiple image recognition tasks in relation to a given processing accuracy constraint.

9. An image recognition device according to any one of claims 1 to 3, The image processing unit performs the multitasking processing using a machine learning model. The controller unit includes a second uncertainty prediction unit (1042b) that predicts uncertainty in at least one of the uncertainty regarding the image input to the controller unit and the uncertainty of the machine learning model. The controller unit is an image recognition device that dynamically changes the processing content of multiple image recognition tasks in the image processing unit, using the uncertainty predicted by the second uncertainty prediction unit.

10. An image recognition device according to any one of claims 1 to 3, The image processing unit performs the multitasking processing using a neural network, which is a type of machine learning model. The controller unit dynamically changes the processing content of multiple image recognition tasks in the image processing unit by changing at least one of the network structure and parameters of the neural network.

11. An image recognition device according to any one of claims 1 to 3, The image processing unit is capable of multitasking, performing multiple image recognition tasks on the surrounding image, which is an image captured by a surrounding surveillance camera (141) that captures images of the area around the vehicle. The controller unit is an image recognition device that takes the surrounding image as input and dynamically changes the processing content of multiple image recognition tasks in the image processing unit according to the trends in the content of the surrounding image.

12. Run by at least one processor, An image processing process that enables multitasking, where multiple image recognition tasks are performed on an image, and where the processing content of the image recognition tasks can be adjusted. The process includes a controller step that adjusts the processing content of multiple image recognition tasks in the image processing step, In the controller step, the image is taken as input, and the processing content of multiple image recognition tasks in the image processing step is dynamically changed according to the trends in the content of the image. The controller process described above is: A scene classification step that uses the aforementioned image as input to classify the scene represented by the content of the aforementioned image, This includes a first uncertainty prediction step for predicting the uncertainty of the scenes to be classified in the scene classification step, As a trend in the content of the aforementioned images, the processing content of multiple image recognition tasks in the image processing step is dynamically changed according to the scene classified in the scene classification step. An image recognition method that dynamically changes the processing content of multiple image recognition tasks in the image processing unit, using the uncertainty of the scene predicted in the first uncertainty prediction step.

13. A method that can be used in a vehicle, Run by at least one processor, An image processing process that enables multitasking, where multiple image recognition tasks are performed on an image, and where the processing content of the image recognition tasks can be adjusted. The process includes a controller step that adjusts the processing content of multiple image recognition tasks in the image processing step, In the controller step, the image is taken as input, and the processing content of multiple image recognition tasks in the image processing step is dynamically changed according to the trends in the content of the image. This includes a vehicle-related information acquisition step that acquires vehicle-related information other than the aforementioned image, The aforementioned vehicle-related information acquisition step involves acquiring at least the vehicle speed information of the vehicle, The controller step includes an image recognition method that, in addition to the trends in the content of the image, dynamically changes the processing content of multiple image recognition tasks in the image processing step, based on the image as input.

14. A method that can be used in a vehicle, Run by at least one processor, An image processing process that enables multitasking, where multiple image recognition tasks are performed on an image, and where the processing content of the image recognition tasks can be adjusted. The process includes a controller step that adjusts the processing content of multiple image recognition tasks in the image processing step, In the controller step, the image is taken as input, and the processing content of multiple image recognition tasks in the image processing step is dynamically changed according to the trends in the content of the image. This includes a vehicle-related information acquisition step that acquires vehicle-related information other than the aforementioned image, The aforementioned vehicle-related information acquisition step involves acquiring information on the driver status of the vehicle, The controller step includes an image recognition method that, in addition to the trends in the content of the image, dynamically changes the processing content of multiple image recognition tasks in the image processing step, based on the image as input.