Evaluation method, program, and evaluation system
The evaluation method employs multiple trained models to address accuracy issues in predicting movement directions by using simple and detailed models based on image resolution, improving estimation accuracy.
Patent Information
- Application Number
- JP2022079051
- Authority / Receiving Office
- JP · JP
- Patent Type
- Patents
- Current Assignee / Owner
- Filing Date
- 2022-05-12
- Publication Date
- 2026-01-14
- Estimated Expiration
- 2042-05-12
AI Technical Summary
Existing methods for predicting the movement direction of moving objects, such as pedestrians, face accuracy challenges due to the reliance on varying information quality, which may not necessarily improve accuracy when increased.
An evaluation method utilizing multiple trained models - a simple model for low-resolution images and a detailed model for high-resolution images - to enhance the accuracy of movement direction estimation by selectively using overall and part information based on image resolution.
Improves the accuracy of estimating the movement direction of moving objects by adaptively selecting appropriate models based on image resolution, thereby enhancing prediction precision.
Smart Images

Figure 0007797959000001 
Figure 0007797959000002 
Figure 0007797959000003
Abstract
Description
[Technical Field]
[0001] The present disclosure relates to an evaluation method, a program (computer program), and an evaluation system. [Background technology]
[0002] Non-Patent Document 1 discloses a technology for predicting a pedestrian's intention to cross the street using a deep neural network model. In Non-Patent Document 1, non-visual-based features are extracted from key points of vehicle speed, pedestrian bounding box, and pedestrian posture, visual-based features are extracted from the local context of the image and the overall context of the image, and the pedestrian's intention to cross the street is predicted using features that combine the non-visual-based features and the visual-based features. [Prior art documents] [Non-patent literature]
[0003] [Non-Patent Document 1] Dongfang Yang, et al., “Predicting Pedestrian Crossing Intention with Feature Fusion and Spatio-Temporal Attention”, [online], April 12, 2021, [Retrieved April 11, 2022], Cornell University, Internet<URL:https: / / arxiv.org / abs / 2104.05485> Summary of the Invention [Problem to be solved by the invention]
[0004] The technology disclosed in Non-Patent Document 1 is expected to be able to predict pedestrians' crossing intentions with high accuracy by using various information such as vehicle speed, pedestrian bounding boxes, key points of pedestrian posture, local context of the image, and global context of the image.
[0005] In evaluating the moving direction of a moving object, such as predicting a pedestrian's intention to cross the street, the accuracy of the evaluation of the moving direction of the moving object may also be affected by the accuracy of various pieces of information themselves, such as those disclosed in Non-Patent Document 1. Therefore, simply increasing the amount of information used to evaluate the moving direction of the moving object may not necessarily lead to an improvement in the accuracy of the evaluation of the moving direction of the moving object.
[0006] The present disclosure provides an evaluation method, a program, and an evaluation system that enable improvement in the accuracy of evaluation of the movement direction of a moving object. [Means for solving the problem]
[0007] An evaluation method according to one aspect of the present disclosure is an evaluation method executed by an arithmetic circuit that can access a storage device that stores multiple trained models. The multiple trained models include a simple model that has been trained to output an evaluation of the movement direction of a moving object in response to input of one or more pieces of overall information based on the entire moving object in image information, and a detailed model that has been trained to output an evaluation of the movement direction of the moving object in response to input of at least one of the one or more pieces of overall information and one or more pieces of part information based on one or more parts of the moving object. When the resolution of an image of the moving object detected from the target image information is less than a threshold, the evaluation method selects a simple model from the multiple trained models, inputs the one or more pieces of overall information of the moving object detected from the target image information into the simple model, and causes the simple model to output an evaluation of the movement direction of the moving object detected from the target image information. When the resolution is equal to or greater than the threshold, the evaluation method selects a detailed model from the multiple trained models, inputs at least one of the one or more pieces of overall information of the moving object detected from the target image information and one or more pieces of part information into the detailed model, and causes the detailed model to output an evaluation of the movement direction of the moving object detected from the target image information.
[0008] An evaluation method according to one aspect of the present disclosure is an evaluation method executed by an arithmetic circuit that can access a storage device that stores multiple trained models. The multiple trained models include a simple model that has been trained to output an evaluation of the movement direction of a moving object in response to input of one or more pieces of overall information based on the entire moving object in image information, and a detailed model that has been trained to output an evaluation of the movement direction of the moving object in response to input of at least one of the one or more pieces of overall information and one or more pieces of part information based on one or more parts of the moving object. The evaluation method confirms that the resolution of an image of the moving object detected from the target image information is less than a threshold, selects a simple model from the multiple trained models in response to confirming that the resolution is less than the threshold, inputs the one or more pieces of overall information of the moving object detected from the target image information to the simple model, and causes the simple model to output an evaluation of the movement direction of the moving object detected from the target image information.
[0009] An evaluation method according to one aspect of the present disclosure is an evaluation method executed by an arithmetic circuit that can access a storage device that stores multiple trained models. The multiple trained models include a simple model that has been trained to output an evaluation of the movement direction of a moving object in response to input of one or more pieces of overall information based on the entire moving object in image information, and a detailed model that has been trained to output an evaluation of the movement direction of the moving object in response to input of at least one of the one or more pieces of overall information and one or more pieces of part information based on one or more parts of the moving object. The evaluation method confirms that the resolution of an image of the moving object detected from the target image information is equal to or greater than a threshold, selects a detailed model from the multiple trained models in response to confirming that the resolution is equal to or greater than the threshold, inputs at least one of the one or more pieces of overall information and one or more pieces of part information of the moving object detected from the target image information to the detailed model, and causes the detailed model to output an evaluation of the movement direction of the moving object detected from the target image information.
[0010] An evaluation method according to one aspect of the present disclosure is an evaluation method executed by an arithmetic circuit that can access a storage device that stores multiple trained models. The multiple trained models include a first model that has been trained to output an evaluation of the moving direction of a moving object in response to input of one or more pieces of overall information based on the entire moving object in image information and one or more pieces of first part information based on one or more first parts of the moving object, and a second model that has been trained to output an evaluation of the moving direction of the moving object in response to input of at least one of the one or more pieces of overall information, at least one of the one or more pieces of first part information, and one or more pieces of second part information based on one or more second parts of the moving object. The one or more second parts are smaller than the one or more first parts. The evaluation method includes, when the resolution of an image of the moving object detected from the target image information is less than a threshold, selecting a first model from the multiple trained models, inputting the one or more pieces of overall information and the one or more pieces of first part information of the moving object detected from the target image information into the first model, and causing the first model to output an evaluation of the moving direction of the moving object detected from the target image information. The evaluation method involves selecting a second model from multiple trained models when the resolution is above a threshold, inputting at least one of one or more pieces of overall information of the moving body detected from the target image information, at least one of one or more pieces of first part information, and one or more pieces of second part information into the second model, and causing the second model to output an evaluation of the movement direction of the moving body detected from the target image information.
[0011] An evaluation method according to one aspect of the present disclosure is an evaluation method executed by an arithmetic circuit that can access a storage device that stores multiple trained models. The multiple trained models include a first model that is trained to output an evaluation of the moving direction of a moving object in response to input of one or more pieces of overall information based on the entire moving object in image information and one or more pieces of first part information based on one or more first parts of the moving object, and a second model that is trained to output an evaluation of the moving direction of the moving object in response to input of at least one of the one or more pieces of overall information, at least one of the one or more pieces of first part information, and one or more pieces of second part information based on one or more second parts of the moving object. The one or more second parts are smaller than the one or more first parts. The evaluation method includes confirming that the resolution of an image of the moving object detected from target image information is less than a threshold, selecting a first model from the multiple trained models in response to confirming that the resolution is less than the threshold, inputting the one or more pieces of overall information and the one or more pieces of first part information of the moving object detected from the target image information into the first model, and causing the first model to output an evaluation of the moving direction of the moving object detected from the target image information.
[0012] An evaluation method according to one aspect of the present disclosure is an evaluation method executed by an arithmetic circuit that can access a storage device that stores multiple trained models. The multiple trained models include a first model that is trained to output an evaluation of the moving direction of a moving object in response to input of one or more pieces of overall information based on the entire moving object in image information and one or more pieces of first part information based on one or more first parts of the moving object, and a second model that is trained to output an evaluation of the moving direction of the moving object in response to input of at least one of the one or more pieces of overall information, at least one of the one or more pieces of first part information, and one or more pieces of second part information based on one or more second parts of the moving object. The one or more second parts are smaller than the one or more first parts. The evaluation method confirms that the resolution of the image of the moving body detected from the target image information is equal to or greater than a threshold, and in response to confirming that the resolution is equal to or greater than the threshold, selects a second model from a plurality of trained models, inputs at least one of one or more pieces of overall information of the moving body detected from the target image information, at least one of one or more pieces of first part information, and one or more pieces of second part information into the second model, and causes the second model to output an evaluation of the movement direction of the moving body detected from the target image information.
[0013] A program according to one aspect of the present disclosure is a program for causing an arithmetic circuit to execute any one of the above evaluation methods.
[0014] An evaluation system according to one aspect of the present disclosure includes a storage device storing multiple trained models and an arithmetic circuit accessible to the storage device. The multiple trained models include a simple model trained to output an evaluation of the movement direction of a moving object in response to input of one or more pieces of overall information based on the entire moving object in image information, and a detailed model trained to output an evaluation of the movement direction of the moving object in response to input of at least one of the one or more pieces of overall information and one or more pieces of part information based on one or more parts of the moving object. When the resolution of an image of the moving object detected from the target image information is less than a threshold, the arithmetic circuit selects a simple model from the multiple trained models, inputs the one or more pieces of overall information of the moving object detected from the target image information to the simple model, and causes the simple model to output an evaluation of the movement direction of the moving object detected from the target image information. When the resolution is equal to or greater than the threshold, the arithmetic circuit selects a detailed model from the multiple trained models, inputs at least one of the one or more pieces of overall information of the moving object detected from the target image information and one or more pieces of part information to the detailed model, and causes the detailed model to output an evaluation of the movement direction of the moving object detected from the target image information.
[0015] An evaluation system according to one aspect of the present disclosure includes a storage device storing multiple trained models and an arithmetic circuit accessible to the storage device. The multiple trained models include a first model trained to output an evaluation of the moving direction of a moving object in response to input of one or more pieces of overall information based on the entire moving object in image information and one or more pieces of first part information based on one or more first parts of the moving object, and a second model trained to output an evaluation of the moving direction of the moving object in response to input of at least one of the one or more pieces of overall information, at least one of the one or more pieces of first part information, and one or more pieces of second part information based on one or more second parts of the moving object. The one or more second parts are smaller than the one or more first parts. When the resolution of the image of the moving object detected from the target image information is less than a threshold, the arithmetic circuit selects the first model from the multiple trained models, inputs the one or more pieces of overall information and the one or more pieces of first part information of the moving object detected from the target image information to the first model, and causes the first model to output an evaluation of the moving direction of the moving object detected from the target image information. When the resolution is equal to or greater than a threshold, the calculation circuit selects a second model from the plurality of trained models, inputs at least one of one or more pieces of overall information of the moving body detected from the target image information, at least one of one or more pieces of first part information, and one or more pieces of second part information into the second model, and causes the second model to output an evaluation of the movement direction of the moving body detected from the target image information. [Effects of the Invention]
[0016] Aspects of the present disclosure enable improved accuracy in estimating the direction of movement of a moving object. [Brief explanation of the drawings]
[0017] [Figure 1] FIG. 1 is a block diagram illustrating a configuration example of a mobile device equipped with an evaluation system according to an embodiment. [Figure 2] 1 is a flowchart illustrating an example of a process executed by the evaluation system of FIG. 1. [Figure 3] Schematic diagram of the first example of an image showing a moving object [Figure 4] Schematic diagram of a second example of an image showing a moving object [Figure 5] FIG. 2 is an explanatory diagram of the positional relationship between the moving device and the moving body shown in FIG. 1; [Figure 6] Schematic diagram of an example image showing a moving object in the distance [Figure 7] Schematic diagram of an example image showing a moving object in the middle distance [Figure 8] Schematic diagram of an example image showing a moving object in the near distance [Figure 9] A schematic diagram of an example of the configuration of a low-resolution model included in the evaluation system of Figure 1. [Figure 10] A schematic diagram of an example of the configuration of a medium-resolution model included in the evaluation system of Figure 1. [Figure 11] Schematic diagram of a high-resolution model configuration example included in the evaluation system of Figure 1. [Figure 12] FIG. 2 is an explanatory diagram of an example of the operation of the mobile device of FIG. 1; [Figure 13] FIG. 2 is an explanatory diagram of an example of the operation of the mobile device of FIG. 1; [Figure 14] FIG. 2 is an explanatory diagram of an example of the operation of the mobile device of FIG. 1; DETAILED DESCRIPTION OF THE INVENTION
[0018] [1. Embodiment] [1.1 Configuration] 1 is a block diagram of a configuration example of a mobile device 1 according to an embodiment. The mobile device 1 is, for example, a robot capable of autonomous movement. The mobile device 1 includes an evaluation system 2, an imaging system 3, and a control system 4.
[0019] The evaluation system 2 is a system for evaluating the movement direction (movement intention) of the moving body 10. In other words, the evaluation system 2 is a system for predicting the movement of the moving body 10. In this embodiment, the evaluation system 2 acquires image information of the moving body 10 from the imaging system 3, provides the result of evaluation of the movement direction of the moving body 10 to the control system 4, and is used to cause the control system 4 to execute an operation according to the result of evaluation of the movement direction of the moving body 10.
[0020] The moving body 10 is a visible object whose moving direction is evaluated by the evaluation system 2. The moving body 10 is an object that can move autonomously. That is, the state of the moving body 10 can include a moving state and a stopped state. In this embodiment, the moving body 10 is a person (for example, a pedestrian). In the following, the "moving body" may be written as a "person." However, this is not intended to limit the moving body to a person, but simply to avoid complexity and make the explanation easier to understand. Therefore, the moving body 10 is not limited to a person, but may also be a living thing other than a human, such as an animal. The moving body 10 is not limited to a living thing, but may also be an inanimate object. Examples of inanimate objects include vehicles such as motorcycles, automobiles, ships, and aircraft, as well as drones. The moving body 10 is not limited to the entire object, but may also be a part of an object.
[0021] The imaging system 3 is a system for generating target image information. The target image information may include data of one or more images in which the moving object 10 is captured. The target image information may be data of a still image in which the moving object 10 is captured, or data of a moving image in which the moving object 10 is captured. In this embodiment, "in which the moving object 10 is captured" includes capturing at least a part of the moving object 10, but not all of it. For example, if the moving object 10 is a person, an image in which the person's face is captured can be said to be an image in which the moving object 10 is captured. The imaging system 3 is communicatively connected to the evaluation system 2 and is capable of providing target image information to the evaluation system 2. The imaging system 3 includes one or more cameras (digital cameras).
[0022] The control system 4 is a system having a function of executing an operation according to the result of evaluation of the movement direction of the mobile object 10. In one example, the control system 4 can be used to control the movement (behavior) of the mobile device 1. The control system 4 can determine the movement of the mobile device 1 according to the result of evaluation of the movement direction of the mobile object 10. The control system 4 is communicatively connected to the evaluation system 2 and can receive the result of evaluation of the movement direction of the mobile object 10 from the evaluation system 2. The control system 4 includes a computer system including one or more memories, one or more processors, etc.
[0023] As shown in FIG. 1, the evaluation system 2 includes an interface 21, a storage device 22, and an arithmetic circuit .
[0024] The interface 21 is used to input information to the evaluation system 2 and output information from the evaluation system 2. The interface 21 includes an input / output device 211 and a communication device 212. The input / output device 211 functions as an input device for inputting information from a user and as an output device for outputting information to a user. The input / output device 211 includes one or more human-machine interfaces. Examples of human-machine interfaces include input devices such as a keyboard, a pointing device (mouse, trackball, etc.), and a touchpad, output devices such as a display and a speaker, and input / output devices such as a touch panel. The communication device 212 is communicatively connected to an external device or system. In this embodiment, the communication device 212 is used for communication with the imaging system 3 and the control system 4 via a communication network. The communication device 212 includes one or more communication interfaces. The communication device 212 is connectable to a communication network and has a function for communicating via the communication network. The communication device 212 complies with a predetermined communication protocol. The predetermined communication protocol may be selected from a variety of well-known wired and wireless communication standards.
[0025] The storage device 22 is used to store information used by the arithmetic circuit 23 and information generated by the arithmetic circuit 23. The storage device 22 includes one or more storages (non-transitory storage media). The storage may be, for example, a hard disk drive, an optical drive, or a solid-state drive (SSD). The storage may also be an internal type, an external type, or a network-attached storage (NAS) type.
[0026] The information stored in the storage device 22 includes multiple trained models. The multiple trained models include a low-resolution model 5, a medium-resolution model 6, and a high-resolution model 7. Note that the terms "low," "medium," and "high" are used with respect to the low-resolution model 5, the medium-resolution model 6, and the high-resolution model 7 to easily distinguish them. FIG. 1 shows a state in which the storage device 22 stores all of the low-resolution model 5, the medium-resolution model 6, and the high-resolution model 7. The low-resolution model 5, the medium-resolution model 6, and the high-resolution model 7 do not need to be stored in the storage device 22 at all times; they only need to be stored in the storage device 22 when needed by the arithmetic circuit 23.
[0027] The low-resolution model 5, the medium-resolution model 6, and the high-resolution model 7 are used to evaluate the moving direction of the moving object 10. The low-resolution model 5, the medium-resolution model 6, and the high-resolution model 7 will be described in detail later.
[0028] The arithmetic circuit 23 is a circuit that controls the operation of the evaluation system 2. The arithmetic circuit 23 is connected to the interface 21 and can access the storage device 22 (i.e., can access the low-resolution model 5, the medium-resolution model 6, and the high-resolution model 7). The arithmetic circuit 23 can be realized, for example, by a computer system including one or more processors (microprocessors) and one or more memories. The one or more processors execute a program (stored in one or more memories or the storage device 22) to realize the function of the arithmetic circuit 23. Here, the program is pre-recorded in the storage device 22, but it may also be provided via a telecommunications line such as the Internet or recorded on a non-transitory recording medium such as a memory card.
[0029] The arithmetic circuit 23 executes the evaluation method shown in Fig. 2 to evaluate the moving direction of the moving object 10. Fig. 2 is a flowchart of an example of the processing of the evaluation system of Fig. 1.
[0030] The arithmetic circuit 23 executes a process of acquiring target image information (S11). In this embodiment, the arithmetic circuit 23 acquires the target image information from the imaging system 3 through the interface 21.
[0031] FIG. 3 is a schematic diagram showing a first example of an image of the moving body 10 included in the target image information. FIG. 4 is a schematic diagram showing a second example of an image of the moving body 10 included in the target image information. In FIGS. 3 and 4, the moving body 10 is moving forward. In other words, the moving body 10 is approaching the mobile device 1. In the image of FIG. 3, the moving body 10 is walking while facing forward. In the image of FIG. 4, the moving body 10 is looking down and is not walking while facing forward.
[0032] The arithmetic circuit 23 executes a process of detecting the moving object 10 from the image included in the target image information (S12). The detection of the moving object 10 from the image may be realized by a conventionally known method. The detection of the moving object 10 from the image may be realized, for example, by an image processing technique such as edge detection, or by utilizing a trained model. As shown in FIGS. 3 and 4, in this embodiment, the arithmetic circuit 23 extracts a bounding box B representing the moving object 10 from the image. The bounding box B is a rectangular box (rectangular boundary line) that surrounds the moving object 10. More specifically, the bounding box B is a rectangular box of a size that just surrounds the moving object 10. Therefore, the width W and height H of the bounding box B change depending on the size of the moving object 10 on the image. The bounding box B may be set by a conventionally known technique. For example, the bounding box B may be set based on whether or not the moving object 10 is included in a rectangular area of interest in the image (see, for example, Patent Publication No. 4447245), or may be set to include the human skeleton detected using human skeleton detection technology (for example, a pose estimation model such as OpenPose).
[0033] The arithmetic circuit 23 performs a process of determining the resolution of the image of the moving body 10 detected from the image included in the target image information (S13). In this embodiment, the arithmetic circuit 23 determines whether the resolution of the image of the moving body 10 corresponds to low resolution, medium resolution, or high resolution.
[0034] The resolution of the image of the moving object 10 corresponds to the number of pixels used to display the moving object 10 in the image. The resolution of the image of the moving object 10 can vary depending on, for example, the distance from the imaging system 3 to the moving object 10, the performance (resolution) of the camera of the imaging system 3, and the size of the moving object 10 in real space. As the distance from the imaging system 3 to the moving object 10 increases, the size of the moving object 10 decreases in the image obtained from the imaging system 3, and the number of pixels in the image of the moving object 10 decreases, resulting in a lower resolution of the image of the moving object 10. If the performance (resolution) of the camera of the imaging system 3 is low, the resolution of the image obtained from the imaging system 3 itself is low, resulting in a lower number of pixels in the image of the moving object 10 and a lower resolution. If the performance (resolution) of the camera of the imaging system 3 is high, the resolution of the image obtained from the imaging system 3 itself is high, resulting in a higher number of pixels in the image of the moving object 10 and a higher resolution. Even if the distance from the imaging system 3 to the moving object 10 is the same, the resolution of the image of the moving object 10 varies depending on the resolution of the camera of the imaging system 3. The size of the moving body 10 in real space may be, for example, the physique (height, etc.) of a person, or the height of a vehicle, etc. For example, between an adult who is 180 cm tall and a child who is 80 cm tall, the child will have fewer pixels, and therefore the resolution will tend to be lower.
[0035] 5 is an explanatory diagram of the positional relationship between mobile device 1 (imaging system 3) and mobile bodies 11, 12, and 13. In this embodiment, mobile device 1 is equipped with imaging system 3, so the distances from imaging system 3 to mobile bodies 11, 12, and 13 are equal to the distances from mobile device 1 to mobile bodies 11, 12, and 13. In FIG. 5, mobile body 11 is at a long distance, mobile body 12 is at a medium distance, and mobile body 13 is at a short distance. For example, distance D1 from mobile device 1 to mobile body 11 is 7 m, distance D2 from mobile device 1 to mobile body 12 is 5 m, and distance D3 from mobile device 1 to mobile body 13 is 3 m.
[0036] Fig. 6 is a schematic diagram of an example of an image showing a moving object 11 at a long distance, Fig. 7 is a schematic diagram of an example of an image showing a moving object 12 at a medium distance, and Fig. 8 is a schematic diagram of an example of an image showing a moving object 13 at a short distance.
[0037] Among the bounding box B1 of the moving object 11 in Figure 6, the bounding box B2 of the moving object 12 in Figure 7, and the bounding box B3 of the moving object 13 in Figure 8, the area of the bounding box B1 is the smallest and the area of the bounding box B3 is the largest. The areas of the bounding boxes B1, B2, and B3 are determined by the size of the bounding boxes B1, B2, and B3 relative to the image. As shown in Figures 6, 7, and 8, the areas of the bounding boxes B1, B2, and B3 are expressed as the product of the number of pixels corresponding to the widths W1, W2, and W3 of the bounding boxes B1, B2, and B3 in the image and the number of pixels corresponding to the heights H1, H2, and H3 of the bounding boxes B1, B2, and B3.
[0038] The number of pixels that make up the bounding box is directly proportional to the area of the bounding box. A larger area of the bounding box means that the image resolution of the moving body 10 is high. A smaller area of the bounding box means that the image resolution of the moving body 10 is low. In Figures 6, 7, and 8, the image resolution of moving body 11 is the lowest, the image resolution of moving body 13 is the highest, and the resolution of moving body 12 is between the image resolution of moving body 11 and the image resolution of moving body 13.
[0039] As is clear from Figures 6, 7, and 8, the closer the moving object 13 is to the image, the larger the area occupied by the moving object 13 in the image. As the area occupied by the moving object 13 in the image increases, the number of pixels constituting the image of the moving object 13 increases, and therefore the resolution of the moving object 13 increases. As is clear from Figures 6, 7, and 8, the farther the moving object 11 is from the image, the smaller the area occupied by the moving object 11 in the image. As the area occupied by the moving object 11 in the image decreases, the number of pixels constituting the image of the moving object 11 decreases, and therefore the resolution of the moving object 11 decreases. In other words, the farther the distance from the imaging system 3 to the moving object 10 is, the lower the resolution of the image of the moving object 10 is. The closer the distance from the imaging system 3 to the moving object 10 is, the higher the resolution of the image of the moving object 10 is.
[0040] In this embodiment, the arithmetic circuit 23 determines the resolution of the image of the moving body 10 based on the area of the bounding box of the moving body 10. As an example, the arithmetic circuit 23 may determine the area of the bounding box of the moving body 10 as the resolution of the image of the moving body 10. The arithmetic circuit 23 determines the resolution of the image of the moving body 10 by comparing the area of the bounding box of the moving body 10 with a threshold. For example, if the area of the bounding box is less than a first threshold, the arithmetic circuit 23 may determine that the resolution of the image of the moving body 10 is low. In this embodiment, if the area of the bounding box is equal to or greater than the first threshold but less than a second threshold, the arithmetic circuit 23 may determine that the resolution of the image of the moving body 10 is medium. If the area of the bounding box is equal to or greater than the second threshold, the arithmetic circuit 23 may determine that the resolution of the image of the moving body 10 is high.
[0041] The first threshold is a threshold for distinguishing between low resolution and medium resolution. For example, if the number of pixels in an image is 360 x 270, and the number of pixels in the bounding box is less than 30 x 90, the resolution of the image of the moving body 10 may be considered to be low resolution. In this case, the first threshold is 2700. The second threshold is a threshold for distinguishing between medium resolution and high resolution. For example, if the number of pixels in an image is 360 x 270, and the number of pixels in the bounding box is 60 x 180 or more, the resolution of the image of the moving body 10 may be considered to be high resolution. In this case, the second threshold is 10800.
[0042] When the target image information includes multiple images, a representative value of the bounding box of the moving object 10 appearing in each of the multiple images may be compared with thresholds (first and second thresholds). Examples of representative values include the average, minimum, maximum, median, and mode. This reduces the possibility that the determination of whether the image resolution of the moving object 10 corresponds to low, medium, or high resolution will change in a short period of time due to changes in the bounding boxes of the multiple images.
[0043] Comparing the moving bodies 11 to 13 in Figures 6 to 8, in the image of the moving body 11 in Figure 6, the entire moving body 11 is relatively clear, but parts of the moving body 11, such as the head, are relatively unclear. For example, in Figure 6, it is difficult to accurately determine the facial orientation of the moving body 11. In the image of the moving body 12 in Figure 7, the entire moving body 12 is relatively clear, and parts of the moving body 11, such as the head R, are also relatively clear. However, parts smaller than the head, such as the eyes or skeleton points of the moving body 11, are relatively unclear. For example, in Figure 7, it is difficult to accurately determine the line of sight or posture of the moving body 11. In the image of the moving body 13 in Figure 8, the entire moving body 13 is relatively clear, and parts (first parts) such as the head R of the moving body 13 are also relatively clear. Furthermore, parts (second parts) smaller than the head, such as the eyes or skeleton points P of the moving body 13, are also relatively clear. Therefore, for the moving body 13 in FIG. 8, the line of sight or the attitude of the moving body 11 can also be determined with high accuracy.
[0044] The lower the resolution of the image of the moving object 10, the lower the accuracy of the information obtained from parts of the moving object 10 rather than the entire moving object 10. The use of information with low accuracy can be one factor in reducing the accuracy of the evaluation of the moving direction of the moving object 10.
[0045] From this perspective, the evaluation system 2 includes a plurality of trained models (a low-resolution model 5, a medium-resolution model 6, and a high-resolution model 7) according to the resolution of the image of the moving body 10. The arithmetic circuit 23 selects a trained model to be used for evaluating the movement direction of the moving body 10 from the plurality of trained models (the low-resolution model 5, the medium-resolution model 6, and the high-resolution model 7) according to the resolution of the image of the moving body 10.
[0046] More specifically, the arithmetic circuit 23 confirms that the resolution of the image of the moving body 10 detected from the target image information is less than a first threshold, and selects a low-resolution model 5 from the multiple trained models in response to confirming that the resolution is less than the first threshold. The arithmetic circuit 23 confirms that the resolution of the image of the moving body 10 detected from the target image information is equal to or greater than a first threshold, and selects a medium-resolution model 6 or a high-resolution model 7 from the multiple trained models in response to confirming that the resolution is equal to or greater than the first threshold. In this case, the arithmetic circuit 23 confirms that the resolution of the image of the moving body 10 detected from the target image information is less than a second threshold, and selects a medium-resolution model 6 from the multiple trained models (the medium-resolution model 6 and the high-resolution model 7) in response to confirming that the resolution is less than the second threshold. On the other hand, the arithmetic circuit 23 confirms that the resolution of the image of the moving body 10 detected from the target image information is equal to or greater than a second threshold, and selects a high-resolution model 7 from the multiple trained models (the medium-resolution model 6 and the high-resolution model 7) in response to confirming that the resolution is equal to or greater than the second threshold.
[0047] As shown in FIG. 2, when the resolution of the image of the moving object 10 is low (S13; low resolution), the arithmetic circuit 23 selects the low-resolution model 5 (S21).
[0048] The low-resolution model 5 is used when the image of the moving object 10 has a low resolution. The low-resolution model 5 is a simple model that is trained to output an evaluation of the moving direction of the moving object 10 in response to input of one or more pieces of overall information based on the entire moving object 10 in the image information.
[0049] The one or more pieces of overall information include at least one of the position of the moving body 10, the speed of the moving body 10, and an image of the entire moving body 10. The position of the moving body 10 is, for example, the absolute position of the moving body 10 in real space, or the position (relative position) of the moving body 10 with respect to the mobile device 1 in real space. The speed of the moving body 10 is, for example, the absolute speed of the moving body 10, or the speed (relative speed) of the moving body 10 with respect to the mobile device 1. The image of the entire moving body 10 is, for example, an image surrounded by a bounding box corresponding to the moving body 10.
[0050] In this embodiment, the evaluation of the moving direction of the moving body 10 includes the probability that the moving direction of the moving body 10 is a first moving direction and the probability that the moving direction of the moving body 10 is a second moving direction. The first moving direction is the direction in which the moving body 10 moves toward an object. The second moving direction is the direction in which the moving body 10 avoids the object. In this embodiment, the object is, for example, the moving device 1. In other words, the evaluation of the moving direction of the moving body 10 can be said to represent the probability of whether the moving body 10 will attempt to avoid the moving device 1. From this perspective, the evaluation of the moving direction of the moving body 10 can be said to be an avoidance intention score that indicates the probability that the moving body 10 will attempt to avoid the moving device 1. Therefore, it can be said that the evaluation system 2 uses the trained model to predict the avoidance intention of a pedestrian ahead based on the configuration resolution of the human area image.
[0051] FIG. 9 is a schematic explanatory diagram of an example configuration of a low-resolution model 5. In FIG. 9, the target image information includes multiple images G in chronological order. A bounding box B of the moving object 10 is detected from the multiple images G. Multiple pieces of overall information I11, I12, and I13 are obtained from each bounding box B. As a result, overall information I11 in chronological order, overall information I12 in chronological order, and overall information I13 in chronological order are obtained. In FIG. 9, overall information I11 is the position of the moving object 10, overall information I12 is the speed of the moving object 10, and overall information I13 is an image of the entire moving object 10.
[0052] 9 receives input of overall information I11, I12, and I13 and outputs an evaluation O1 of the moving direction of the moving object 10. The low-resolution model 5 in FIG. 9 includes a first network mechanism 511, a first attention mechanism 512, a second network mechanism 521, a second attention mechanism 522, a third network mechanism 53, a combining unit 54, and an output unit 55.
[0053] The first network mechanism 511 extracts features from the overall information I11 in time series. The first network mechanism 511 includes a recurrent neural network (RNN) architecture such as a long short-term memory (LSTM) or a gated recurrent unit (GRU). The first attention mechanism 512 includes a temporal attention mechanism that determines which piece of the overall information I11 in time series should be focused on. This allows the feature of the overall information I11 that is considered most important to be extracted from the overall information I11 in time series, and inputs it to the combination unit 54.
[0054] The second network mechanism 521 extracts features from the overall information I12 in time series. The second network mechanism 521 includes a recurrent neural network (RNN) architecture such as a long-short-term memory (LSTM) or a gated recurrent unit (GRU). The second attention mechanism 522 includes a temporal attention mechanism that determines which piece of the overall information I12 in time series should be focused on. This allows the feature of the overall information I12 that is considered most important to be extracted from the overall information I12 in time series, and inputs it to the combination unit 54.
[0055] The third network mechanism 53 extracts features from the overall information I13 in chronological order. As described above, the overall information I13 is an image of the entire moving object 10. The third network mechanism 53 includes, for example, a 3D-convolutional neural network (3DCNN) or an architecture that combines a convolutional neural network (CNN) and a recurrent neural network (RNN). This allows features to be extracted from the overall information I13 in chronological order and input to the combining unit 54.
[0056] The combination unit 54 collectively inputs the features of the overall information I11, I12, and I13 to the output unit 55. In the case of Fig. 9, the combination unit 54 inputs to the output unit 55 the feature of the overall information I11 that is considered to be the most important among the overall information I11 in chronological order from the first attention mechanism 512, the feature of the overall information I12 that is considered to be the most important among the overall information I12 in chronological order from the second attention mechanism 522, and the feature of the overall information I13 in chronological order from the third network mechanism 53.
[0057] The output unit 55 outputs an evaluation O1 of the moving direction of the moving object 10 from the feature amount from the combination unit 54. The output unit 55 includes, for example, a fully connected layer. The fully connected layer includes, for example, a softmax function.
[0058] As shown in FIG. 2, when the resolution of the image of the moving object 10 detected from the target image information is low (S13; low resolution), the arithmetic circuit 23 selects a simple model (low-resolution model 5) from multiple trained models (S21) and performs evaluation using the low-resolution model 5 (S22). More specifically, the arithmetic circuit 23 inputs overall information I11, I12, and I13 of the moving object 10 detected from the target image information to the low-resolution model 5, and causes the low-resolution model 5 to output an evaluation O1 of the movement direction of the moving object 10 detected from the target image information. This allows the evaluation system 2 to evaluate the movement direction of the moving object 10 without using information (region information, described below) that is likely to be low in accuracy in the case of low resolution. As a result, the evaluation system 2 enables improvement in the accuracy of the evaluation of the movement direction of the moving object 10.
[0059] If the resolution of the image of the moving object 10 is medium resolution (S13; medium resolution), the arithmetic circuit 23 selects the medium resolution model 6 (S31). As shown in FIG. 2, if the resolution of the image of the moving object 10 is high resolution (S13; high resolution), the arithmetic circuit 23 selects the high resolution model 7 (S41).
[0060] Each of the medium resolution model 6 and the high resolution model 7 is used when the resolution of the image of the moving body 10 is medium. The medium resolution model 6 is a detailed model trained to output an evaluation of the direction of movement of the moving body 10 in response to input of at least one piece of overall information based on the entire moving body 10 in the image information and one or more pieces of part information based on one or more parts of the moving body 10 in the image information.
[0061] The one or more pieces of part information include, for example, at least one of the positions of one or more parts of the moving body 10, the orientation of one or more parts of the moving body 10, an image of one or more parts of the moving body 10, and information based on the correlation of the multiple parts of the moving body 10. An example of the position of a part is the position of the face of the moving body 10. Examples of the orientation of a part are the orientation of the face of the moving body and the line of sight of the moving body. An example of the image of a part is an image of the face of the moving body. An example of information based on the correlation of the multiple parts is information about the posture of the moving body.
[0062] The parts of the moving body 10 can be classified into first parts and second parts that are smaller than the first parts. The first part is a part of the body of the moving body 10, for example, the face R shown in FIG. 7. The first part is not limited to the face, but may be the head, arms, torso, legs, etc. The second part is a part that is smaller than the part of the body of the moving body 10. In comparison with the first part, the second part can be regarded as a point on the moving body 10. An example of the second part is, for example, a skeleton point P shown in FIG. 8. The second part is not limited to a skeleton point, but may be the eyes, nose, fingers, joints, etc.
[0063] From this point of view, the part information can be classified into first part information based on the first part and second part information based on the second part. In order to obtain the second part information with high accuracy, the resolution of the image of the moving object 10 needs to be higher than that of the first part information.
[0064] In this embodiment, the medium resolution model 6 uses first region information in addition to overall information. That is, the medium resolution model 6 is a first model that has been trained to output an evaluation of the moving direction of the moving object 10 in response to input of one or more pieces of overall information based on the entire moving object 10 in the image information and one or more pieces of first region information based on one or more first regions of the moving object 10.
[0065] In this embodiment, the high-resolution model 7 uses second region information in addition to the overall information and the first region information. That is, the high-resolution model 7 is a second model that has been trained to output an evaluation of the moving direction of the moving object 10 in response to input of at least one of one or more pieces of overall information based on the entire moving object 10 in the image information, at least one of one or more pieces of first region information, and one or more pieces of second region information.
[0066] FIG. 10 is a schematic explanatory diagram of an example configuration of a medium resolution model 6. In FIG. 10, the target image information includes multiple images G in chronological order. A bounding box B of a moving object 10 is detected from the multiple images G. Multiple pieces of overall information I11, I12, and I13 are obtained from each bounding box B. As a result, overall information I11 in chronological order, overall information I12 in chronological order, and overall information I13 in chronological order are obtained. In FIG. 10, overall information I11 is the position of the moving object 10, overall information I12 is the speed of the moving object 10, and overall information I13 is an image of the entire moving object 10.
[0067] In FIG. 10, a bounding box R of the face of moving object 10 is detected from multiple images G. Multiple pieces of first region information I21, I22, and I23 are obtained from each bounding box R. As a result, first region information I21 in chronological order, first region information I22 in chronological order, and first region information I23 in chronological order are obtained. In FIG. 10, first region information I21 is the position of the face of moving object 10, first region information I22 is the direction of the face of moving object 10, and first region information I23 is an image of the face of moving object 10.
[0068] The medium resolution model 6 in Figure 10 outputs an evaluation O2 of the moving direction of the moving object 10 in response to input of overall information I11, I12, and I13 and part information I21, I22, and I23. The medium resolution model 6 in Figure 10 includes a first network mechanism 611, a first attention mechanism 612, a second network mechanism 621, a second attention mechanism 622, a third network mechanism 63, a fourth network mechanism 641, a fourth attention mechanism 642, a fifth network mechanism 651, a fifth attention mechanism 652, a sixth network mechanism 66, a combination unit 67, and an output unit 68.
[0069] The first network mechanism 611, the first attention mechanism 612, the second network mechanism 621, the second attention mechanism 622, and the third network mechanism 63 are similar to the first network mechanism 511, the first attention mechanism 512, the second network mechanism 521, the second attention mechanism 522, and the third network mechanism 53, respectively, of the low-resolution model 5 in Figure 9.
[0070] The fourth network mechanism 641 extracts features from the chronologically ordered first part information I21. The fourth network mechanism 641 includes a recurrent neural network (RNN) architecture such as a long-short-term memory (LSTM) or a gated recurrent unit (GRU). The fourth attention mechanism 642 includes a temporal attention mechanism that determines which piece of first part information I21 in the chronologically ordered first part information I21 should be focused on. This allows the feature of the first part information I21 that is considered most important to be extracted from the chronologically ordered first part information I21 and input to the combination unit 67.
[0071] The fifth network mechanism 651 extracts features from the chronologically ordered first part information I22. The fifth network mechanism 651 includes a recurrent neural network (RNN) architecture such as a long-short-term memory (LSTM) or a gated recurrent unit (GRU). The fifth attention mechanism 652 includes a temporal attention mechanism that determines which of the chronologically ordered first part information I22 should be focused on. This extracts features of the first part information I22 that is considered most important from the chronologically ordered first part information I22, and inputs them to the combination unit 67.
[0072] The sixth network mechanism 66 extracts features from the chronologically ordered first region information I23. As described above, the first region information I23 is an image of the face (first region) of the moving object 10. The sixth network mechanism 66 includes, for example, a three-dimensional convolutional neural network (3DCNN) or an architecture that combines a convolutional neural network (CNN) and a recurrent neural network (RNN). As a result, features are extracted from the chronologically ordered first region information I23 and input to the combination unit 67.
[0073] The combination unit 67 collectively inputs the features of the overall information I11, I12, and I13 and the features of the part information (first part information) I21, I22, and I23 to the output unit 68. In the case of Fig. 10 , the combination unit 67 inputs to the output unit 68 the feature of the overall information I11 considered to be the most important of the chronologically ordered overall information I11 from the first attention mechanism 612, the feature of the overall information I12 considered to be the most important of the chronologically ordered overall information I12 from the second attention mechanism 622, the feature of the overall information I13 in chronological order from the third network mechanism 63, the feature of the first part information I21 considered to be the most important of the chronologically ordered first part information I21 from the fourth attention mechanism 642, the feature of the first part information I22 considered to be the most important of the chronologically ordered first part information I22 from the fifth attention mechanism 652, and the feature of the first part information I23 in chronological order from the sixth network mechanism 66.
[0074] The output unit 68 outputs an evaluation O2 of the moving direction of the moving object 10 from the feature amount from the combination unit 67. The output unit 68 includes, for example, a fully connected layer. The fully connected layer includes, for example, a softmax function.
[0075] As shown in FIG. 2, when the resolution of the image of the moving object 10 detected from the target image information is medium resolution (S13; medium resolution), the arithmetic circuit 23 selects a medium resolution model 6 (detailed model, first model) from multiple trained models (S31) and performs evaluation using the medium resolution model 6 (S32). More specifically, the arithmetic circuit 23 inputs the overall information I11, I12, and I13 of the moving object 10 detected from the target image information and the part information (first part information) I21, I22, and I23 to the medium resolution model 6, and causes the medium resolution model 6 to output an evaluation O2 of the movement direction of the moving object 10 detected from the target image information. This allows the evaluation system 2 to evaluate the movement direction of the moving object 10 without using information (second part information) that is likely to be low in accuracy in the case of medium resolution. As a result, the evaluation system 2 enables improved accuracy in evaluating the movement direction of the moving object 10.
[0076] FIG. 11 is a schematic explanatory diagram of an example configuration of the high-resolution model 7. In FIG. 11, the target image information includes multiple images G in chronological order. A bounding box B of the moving object 10 is detected from the multiple images G. Multiple pieces of overall information I11, I12, and I13 are obtained from each bounding box B. As a result, overall information I11 in chronological order, overall information I12 in chronological order, and overall information I13 in chronological order are obtained. In FIG. 11, overall information I11 is the position of the moving object 10, overall information I12 is the speed of the moving object 10, and overall information I13 is an image of the entire moving object 10.
[0077] In FIG. 11, a bounding box R of the face (first region) of moving object 10 is detected from multiple images G. Multiple pieces of first region information I21, I22, and I23 are obtained from each bounding box R. As a result, first region information I21 in chronological order, first region information I22 in chronological order, and first region information I23 in chronological order are obtained. In FIG. 11, first region information I21 is the position of the face of moving object 10, first region information I22 is the direction of the face of moving object 10, and first region information I23 is an image of the face of moving object 10.
[0078] In FIG. 11, multiple skeleton points (second regions) P of the moving object 10 are detected from multiple images G. Second region information I31 is obtained based on the associations between the multiple skeleton points P. This results in the second region information I31 being in chronological order. In FIG. 11, the second region information I31 is, for example, information on the posture of the moving object 10. In some examples, the second region information may also include information on the skeleton points, the positions of the skeleton points, and neighboring information of the skeleton points.
[0079] The high-resolution model 7 in Fig. 11 outputs an evaluation O3 of the movement direction of the moving object 10 in response to input of overall information I11, I12, I13, first region information I21, I22, I23, and second region information I31. The high-resolution model 7 in Fig. 11 includes a first network mechanism 711, a first attention mechanism 712, a second network mechanism 721, a second attention mechanism 722, a third network mechanism 73, a fourth network mechanism 741, a fourth attention mechanism 742, a fifth network mechanism 751, a fifth attention mechanism 752, a sixth network mechanism 76, a seventh network mechanism 771, a seventh attention mechanism 772, a combination unit 78, and an output unit 79.
[0080] The first network mechanism 711, the first attention mechanism 712, the second network mechanism 721, the second attention mechanism 722, and the third network mechanism 73 are similar to the first network mechanism 511, the first attention mechanism 512, the second network mechanism 521, the second attention mechanism 522, and the third network mechanism 53, respectively, of the low-resolution model 5 in Figure 9.
[0081] The fourth network mechanism 741, the fourth attention mechanism 742, the fifth network mechanism 751, the fifth attention mechanism 752, and the sixth network mechanism 76 are similar to the fourth network mechanism 641, the fourth attention mechanism 642, the fifth network mechanism 651, the fifth attention mechanism 652, and the sixth network mechanism 66, respectively, of the medium resolution model 6 in Figure 10.
[0082] The seventh network mechanism 771 extracts features from the chronologically ordered second body part information I31. The seventh network mechanism 771 includes a recurrent neural network (RNN) architecture such as a long-short-term memory (LSTM) or a gated recurrent unit (GRU). The seventh attention mechanism 772 includes a temporal attention mechanism that determines which piece of second body part information I31 should be focused on among the chronologically ordered second body part information I31. This extracts features of the second body part information I31 that is considered most important from the chronologically ordered second body part information I31, and inputs them to the combination unit 78.
[0083] The combining unit 78 inputs to the output unit 79 the feature amounts of the overall information I11, I12, and I13, the feature amounts of the region information (first region information) I21, I22, and I23, and the feature amount of the region information (second region information) I31 all together. In the case of Figure 11, the combination unit 78 inputs to the output unit 79 the features of the overall information I11 that is considered to be the most important among the overall information I11 in chronological order from the first attention mechanism 712, the features of the overall information I12 that is considered to be the most important among the overall information I12 in chronological order from the second attention mechanism 722, the features of the overall information I13 in chronological order from the third network mechanism 73, the features of the first part information I21 that is considered to be the most important among the first part information I21 in chronological order from the fourth attention mechanism 742, the features of the first part information I22 that is considered to be the most important among the first part information I22 in chronological order from the fifth attention mechanism 752, the features of the first part information I23 in chronological order from the sixth network mechanism 66, and the features of the second part information I31 that is considered to be the most important among the second part information I31 in chronological order from the seventh attention mechanism 772.
[0084] The output unit 79 outputs an evaluation O3 of the moving direction of the moving object 10 from the feature amount from the combination unit 78. The output unit 79 includes, for example, a fully connected layer. The fully connected layer includes, for example, a softmax function.
[0085] As shown in FIG. 2, when the resolution of the image of the moving object 10 detected from the target image information is high resolution (S13; high resolution), the arithmetic circuit 23 selects a high-resolution model 7 (detailed model, second model) from multiple trained models (S41) and performs evaluation using the high-resolution model 7 (S42). More specifically, the arithmetic circuit 23 inputs the overall information I11, I12, and I13 of the moving object 10 detected from the target image information, the part information (first part information) I21, I22, and I23, and the part information (second part information) I31, to the high-resolution model 7, and causes the high-resolution model 7 to output an evaluation O3 of the movement direction of the moving object 10 detected from the target image information. In this way, the evaluation system 2 evaluates the movement direction of the moving object 10 using more information in the case of high resolution than in the case of medium resolution. As a result, the evaluation system 2 enables improved accuracy in evaluating the movement direction of the moving object 10.
[0086] The result of the evaluation of the moving direction of the moving object 10 by the evaluation system 2 described above can be used in the control system 4 of the moving device 1 to determine the operation of the moving device 1 .
[0087] 12 to 14 are explanatory diagrams of an example of the operation of the mobile apparatus 1. In Fig. 12, a moving object 10 is present on a planned travel route M1 of the mobile apparatus 1. The moving object 10 is moving in a moving direction V0. In Fig. 12, the moving direction V0 is represented by a relative velocity vector of the moving object 10 with respect to the mobile apparatus 1.
[0088] In the mobile device 1, the evaluation system 2 evaluates the movement direction of the mobile object 10. The evaluation of the movement direction of the mobile object 10 includes the probability that the movement direction of the mobile object 10 is a first movement direction V1 and the probability that the movement direction of the mobile object 10 is a second movement direction V2. The first movement direction V1 is the direction in which the mobile object 10 moves toward the object (mobile device 1). The second movement direction V2 is the direction in which the mobile object 10 avoids the object (mobile device 1).
[0089] If there is a high probability that the moving direction of the moving object 10 is the first moving direction V1, there is a high possibility that the moving object 10 will reach the proximity area 1a of the moving device 1 and collide with the moving device 1. In this case, as shown in Fig. 13, the control system 4 of the moving device 1 can take action to avoid a collision with the moving object 10 by changing the planned travel path M1 of the moving device 1 to a planned travel path M2.
[0090] When there is a high probability that the moving direction of the moving object 10 is the second moving direction V2, there is a low possibility that the moving object 10 will collide with the moving apparatus 1. In this case, as shown in Fig. 14, in the moving apparatus 1, the control system 4 takes action to maintain the planned travel path M1 of the moving apparatus 1.
[0091] In this way, in the mobile device 1, the control system 4 controls the behavior of the mobile device 1 based on the evaluation of the moving direction of the moving body 10 obtained from the evaluation system 2, thereby reducing the possibility of a collision between the mobile device 1 and the moving body 10. In particular, the mobile device 1 can set a planned route of travel based on a movement policy that takes into account mutual yielding and considers whether the moving body 10 will take action to avoid a collision with the mobile device 1, rather than a movement policy that prioritizes the movement of the moving body 10 so as to simply avoid a collision with the moving body 10 completely. This may enable the mobile device 1 to achieve smooth passing by the moving body 10.
[0092] [1.2 Effects, etc.] The evaluation method described above is executed by an arithmetic circuit 23 that can access a storage device 22 that stores multiple trained models. The multiple trained models include a simple model (low-resolution model 5) that has been trained to output an evaluation of the movement direction of the moving object 10 in response to input of one or more pieces of overall information I11 to I13 based on the entire moving object 10 to 13 (bounding boxes B, B1 to B3) in the image information, and detailed models (medium-resolution model 6, high-resolution model 7) that have been trained to output an evaluation of the movement direction of the moving object 10 in response to input of at least one of the one or more pieces of overall information I11 to I13 and one or more pieces of part information I21 to I23, I31 based on one or more parts R, P of the moving object 10.
[0093] Here, the evaluation method confirms that the resolution of the images of the moving bodies 10-13 detected from the target image information is less than a threshold (first threshold), selects a simple model (low-resolution model 5) from the plurality of trained models in response to confirmation that the resolution is less than the threshold (first threshold), inputs one or more pieces of overall information I11-I13 of the moving body 10 detected from the target image information to the simple model (low-resolution model 5), and causes the simple model (low-resolution model 5) to output an evaluation O1 of the movement direction of the moving body 10 detected from the target image information. This configuration enables improvement in the accuracy of the evaluation of the movement direction of the moving bodies 10-13.
[0094] Alternatively, the evaluation method confirms that the resolution of the images of the moving objects 10-13 detected from the target image information is equal to or greater than a threshold (first threshold), and in response to confirming that the resolution is equal to or greater than the threshold (first threshold), selects detailed models (medium resolution model 6, high resolution model 7) from the plurality of trained models, inputs at least one of one or more pieces of overall information I11-I13 of the moving objects 10-13 detected from the target image information and one or more pieces of part information I21-I23, I31 into the detailed models (medium resolution model 6, high resolution model 7), and causes the detailed models (medium resolution model 6, high resolution model 7) to output evaluations O2, O3 of the movement direction of the moving object 10 detected from the target image information. This configuration enables improvement in the accuracy of the evaluation of the movement direction of the moving objects 10-13.
[0095] In the evaluation method, the resolution is determined based on the area of the bounding boxes B, B1 to B3 of the moving objects 10 to 13 detected from the target image information. This configuration makes it possible to improve the accuracy of evaluation of the moving directions of the moving objects 10 to 13.
[0096] In the evaluation method, one or more pieces of overall information I11 to I13 include at least one of the positions of the moving bodies 10 to 13, the speeds of the moving bodies 10 to 13, and images of the entire moving bodies 10 to 13 (bounding boxes B, B1 to B3). This configuration enables improvement in the accuracy of evaluation of the moving directions of the moving bodies 10 to 13.
[0097] In the evaluation method, the one or more pieces of part information I21, I22, I23, I31 include at least one of the positions of the one or more parts R, P of the moving objects 10 to 13, the orientation of the one or more parts R, P of the moving objects 10 to 13, an image of the one or more parts R, P of the moving objects 10 to 13, and information based on the correlation between the multiple parts R, P of the moving objects 10 to 13. This configuration enables improvement in the accuracy of evaluation of the movement direction of the moving objects 10 to 13.
[0098] In the evaluation method, the positions of one or more parts R, P include the positions of the faces of the moving objects 10 to 13. This configuration makes it possible to improve the accuracy of evaluation of the moving directions of the moving objects 10 to 13.
[0099] In the evaluation method, the orientation of one or more parts R, P includes at least one of the orientation of the face of the moving object 10-13 and the line of sight of the moving object 10-13. This configuration enables improvement in the accuracy of evaluation of the moving direction of the moving object 10-13.
[0100] In the evaluation method, the image of one or more parts R, P includes an image of the face of the moving object 10 to 13. This configuration makes it possible to improve the accuracy of evaluation of the moving direction of the moving object 10 to 13.
[0101] In the evaluation method, the information based on the correlation between the multiple parts R and P includes information about the posture of the moving objects 10 to 13. This configuration makes it possible to improve the accuracy of evaluation of the moving directions of the moving objects 10 to 13.
[0102] In the evaluation method, the evaluation of the movement direction of the moving bodies 10-13 includes the probability that the movement direction of the moving bodies 10-13 is a first movement direction and the probability that the movement direction of the moving bodies 10-13 is a second movement direction. The first movement direction is the direction in which the moving bodies 10-13 move toward the object (the moving device 1). The second movement direction is the direction in which the moving bodies 10-13 avoid the object (the moving device 1). This configuration enables efficient movement of the object taking into account the movement direction of the moving bodies 10-13.
[0103] From another perspective, the evaluation method described above is executed by an arithmetic circuit 23 that can access a storage device 22 that stores multiple trained models. The multiple trained models include a first model (medium-resolution model 6) that has been trained to output an evaluation O2 of the movement direction of the moving objects 10-13 in response to input of one or more pieces of overall information I11-I13 based on the entire moving objects 10-13 (bounding boxes B, B1-B3) in the image information and one or more pieces of first region information I21-I23 based on one or more first regions R of the moving objects 10-13, and a second model (high-resolution model 7) that has been trained to output an evaluation O3 of the movement direction of the moving objects 10-13 in response to input of at least one of the one or more pieces of overall information I11-I13, at least one of the one or more pieces of first region information I21-I23, and one or more pieces of second region information I31 based on one or more second regions P of the moving objects 10-13. The one or more second portions P are smaller than the one or more first portions R.
[0104] Here, the evaluation method confirms that the resolution of the images of the moving objects 10-13 detected from the target image information is less than a threshold (second threshold), and in response to confirming that the resolution is less than the threshold (second threshold), selects a first model (medium resolution model 6) from the multiple trained models, inputs one or more pieces of overall information I11-I13 and one or more pieces of first part information I21-I23 of the moving objects 10-13 detected from the target image information to the first model (medium resolution model 6), and causes the first model (medium resolution model 6) to output an evaluation O2 of the movement direction of the moving objects 10-13 detected from the target image information. This configuration enables improvement in the accuracy of the evaluation of the movement direction of the moving objects 10-13.
[0105] Alternatively, the evaluation method may include confirming that the resolution of the images of the moving objects 10-13 detected from the target image information is equal to or greater than a threshold (second threshold), and selecting a second model (high-resolution model 7) from the plurality of trained models in response to confirming that the distances D1-D3 are equal to or greater than 1 / 2, inputting at least one of one or more pieces of overall information I11-I13 of the moving objects 10-13 detected from the target image information, at least one of one or more pieces of first portion information I21-I23, and one or more pieces of second portion information I31 into the second model (high-resolution model 7), and causing the second model (high-resolution model 7) to output an evaluation O3 of the movement direction of the moving objects 10-13 detected from the target image information. This configuration enables improvement in the accuracy of the evaluation of the movement direction of the moving objects 10-13.
[0106] The above-described program is a program for causing the arithmetic circuit 23 to execute the above-described evaluation method. This configuration makes it possible to improve the accuracy of evaluation of the moving directions of the moving objects 10-13.
[0107] The evaluation system 2 described above includes a storage device 22 that stores multiple trained models, and an arithmetic circuit 23 that can access the storage device 22. The multiple trained models include a simple model (low-resolution model 5) that has been trained to output an evaluation of the movement direction of the moving object 10 in response to input of one or more pieces of overall information I11 to I13 based on the entire moving object 10 to 13 (bounding boxes B, B1 to B3) in the image information, and detailed models (medium-resolution model 6, high-resolution model 7) that have been trained to output an evaluation of the movement direction of the moving object 10 in response to input of at least one of the one or more pieces of overall information I11 to I13 and one or more pieces of part information I21 to I23, I31 based on one or more parts R, P of the moving object 10. When the resolution of the image of the moving body 10-13 detected from the target image information is less than a threshold (first threshold), the calculation circuit 23 selects a simple model (low resolution model 5) from the multiple trained models, inputs one or more pieces of overall information I11-I13 of the moving body 10-13 detected from the target image information into the simple model (low resolution model 5), and causes the simple model (low resolution model 5) to output an evaluation O1 of the movement direction of the moving body 10 detected from the target image information. When the resolution is equal to or greater than a threshold (first threshold), the arithmetic circuit 23 selects detailed models (medium resolution model 6, high resolution model 7) from the multiple trained models, inputs at least one of one or more pieces of overall information I11-I13 of the moving object 10 detected from the target image information and one or more pieces of part information I21-I23, I31 to the detailed models (medium resolution model 6, high resolution model 7), and causes the detailed models (medium resolution model 6, high resolution model 7) to output evaluations O2, O3 of the movement directions of the moving objects 10-13 detected from the target image information. This configuration enables improvement in the accuracy of the evaluation of the movement directions of the moving objects 10-13.
[0108] The evaluation system 2 described above includes a storage device 22 that stores multiple trained models and an arithmetic circuit 23 that can access the storage device 22. The multiple trained models include a first model (medium-resolution model 6) that has been trained to output an evaluation O2 of the movement direction of the moving objects 10-13 in response to input of one or more pieces of overall information I11-I13 based on the entire moving objects 10-13 (bounding boxes B, B1-B3) in the image information and one or more pieces of first region information I21-I23 based on one or more first regions R of the moving objects 10-13, and a second model (high-resolution model 7) that has been trained to output an evaluation O3 of the movement direction of the moving objects 10-13 in response to input of at least one of the one or more pieces of overall information I11-I13, at least one of the one or more pieces of first region information I21-I23, and one or more pieces of second region information I31 based on one or more second regions P of the moving objects 10-13. The one or more second regions P are smaller than the one or more first regions R. When the resolution of the image of the moving bodies 10-13 detected from the target image information is less than a threshold (second threshold), the arithmetic circuit 23 selects a first model (medium resolution model 6) from the multiple trained models, inputs one or more pieces of overall information I11-I13 and one or more pieces of first region information I21-I23 of the moving bodies 10-13 detected from the target image information to the first model (medium resolution model 6), and causes the first model (medium resolution model 6) to output an evaluation O2 of the movement direction of the moving bodies 10-13 detected from the target image information. When the resolution is equal to or greater than a threshold (second threshold), the arithmetic circuit 23 selects a second model (high-resolution model 7) from the plurality of trained models, inputs at least one of one or more pieces of overall information I11-I13 of the moving objects 10-13 detected from the target image information, at least one of one or more pieces of first region information I21-I23, and one or more pieces of second region information I31 to the second model (high-resolution model 7), and causes the second model (high-resolution model 7) to output an evaluation O3 of the movement direction of the moving objects 10-13 detected from the target image information. This configuration enables improvement in the accuracy of the evaluation of the movement direction of the moving objects 10-13.
[0109] [2. Variable examples] The embodiments of the present disclosure are not limited to the above-described embodiments. The above-described embodiments can be modified in various ways depending on the design, etc., as long as the object of the present disclosure can be achieved. Modifications of the above-described embodiments are listed below. The modifications described below can be applied in appropriate combinations.
[0110] In one variant, the multiple trained models may include two of the low-resolution model 5, the medium-resolution model 6, and the high-resolution model 7, rather than all of them. For example, the arithmetic circuitry 23 may select the low-resolution model 5 when the resolution of the image of the moving body 10 detected from the target image information is less than a threshold, and may select the medium-resolution model 6 or the high-resolution model 7 when the resolution of the image of the moving body 10 detected from the target image information is equal to or greater than the threshold. For example, the arithmetic circuitry 23 may select the medium-resolution model 6 when the resolution of the image of the moving body 10 detected from the target image information is less than a threshold, and may select the high-resolution model 7 when the resolution of the image of the moving body 10 detected from the target image information is equal to or greater than the threshold.
[0111] In one modification, the simple model used as the low-resolution model 5 may be trained to output an evaluation of the moving direction of the moving object 10 in response to input of one or more pieces of overall information I11 to I13 based on the entire moving object 10 in the image information. In the simple model, the number of pieces of overall information is not particularly limited.
[0112] In one variation, the detailed model used as the medium resolution model 6 or the high resolution model 7 may be trained to output an evaluation of the moving direction of the moving object 10 in response to input of at least one of one or more pieces of overall information I11-I13 based on the entire moving object 10 in the image information and one or more pieces of part information I21-I23 based on one or more parts of the moving object 10 in the image information. In the detailed model, the number of pieces of overall information and the number of pieces of part information are not particularly limited. In particular, the detailed model does not necessarily need to use all of the overall information used in the simple model.
[0113] In one variant, the first model used as the medium resolution model 6 may be trained to output an evaluation of the moving direction of the moving object 10 in response to input of one or more pieces of overall information I11-I13 based on the entire moving object 10 in the image information and one or more pieces of first region information I21-I23 based on one or more first regions of the moving object 10 in the image information. In the first model, the number of pieces of overall information and the number of pieces of first region information are not particularly limited.
[0114] In one variation, the second model used as the high-resolution model 7 may be trained to output an evaluation of the movement direction of the moving object 10 in response to input of at least one of one or more pieces of overall information I11-I13 based on the entire moving object 10 in the image information, at least one of one or more pieces of first region information I21-I23 based on one or more first regions of the moving object 10 in the image information, and at least one of one or more pieces of second region information I31 based on one or more second regions of the moving object 10 in the image information. In the second model, the number of pieces of overall information, the number of pieces of first region information, and the number of pieces of second region information are not particularly limited. In particular, the second model does not necessarily need to use all of the overall information and all of the first region information used in the first model.
[0115] In one variant, the interface 21 of the assessment system 2 does not have to comprise both an input / output device and a communication device.
[0116] In one modified example, the imaging system 3 may include a plurality of cameras with different performance (resolution). The evaluation system 2 may evaluate the moving direction of the moving object based on the target image information obtained from each of the plurality of cameras with different performance (resolution).
[0117] In one variation, the evaluation system 2 does not necessarily have to be installed in the mobile device 1. The evaluation system 2 can be used in devices or systems other than the mobile device 1. For example, the evaluation system 2 can be used in an alarm system that issues a warning of a collision with the mobile object 10. The evaluation system 2 can also be implemented by a computer system such as multiple servers. In other words, it is not necessary for multiple functions (components) in the evaluation system 2 to be concentrated in a single housing, and the components of the evaluation system 2 can be distributed across multiple housings. Furthermore, at least some of the functions of the evaluation system 2, for example, some functions of the arithmetic circuit 23, can be implemented by the cloud (cloud computing) or the like.
[0118] [3. Aspects] As is clear from the above-described embodiments and modifications, the present disclosure includes the following aspects. In the following, reference numerals are given in parentheses only to clarify the correspondence with the embodiments. Note that, in consideration of readability of the text, the reference numerals in parentheses may be omitted from the second and subsequent times.
[0119] The first aspect is an evaluation method executed by a calculation circuit (23) that can access a storage device (22) that stores a plurality of trained models. The plurality of trained models include a simple model (low-resolution model 5) that has been trained to output an evaluation of the movement direction of a moving object (10) in response to input of one or more pieces of overall information (I11 to I13) based on the entirety (bounding boxes B, B1 to B3) of the moving object (10 to 13) in image information, and detailed models (medium-resolution model 6, high-resolution model 7) that have been trained to output an evaluation of the movement direction of the moving object (10) in response to input of at least one of the one or more pieces of overall information (I11 to I13) and one or more pieces of part information (I21 to I23, I31) based on one or more parts (R, P) of the moving object (10). The evaluation method selects the simple model (low-resolution model 5) from the plurality of trained models when the resolution of the image of the moving body (10-13) detected from the target image information is less than a threshold (first threshold), inputs the one or more pieces of overall information (I11-I13) of the moving body (10-13) detected from the target image information into the simple model (low-resolution model 5), and causes the simple model (low-resolution model 5) to output an evaluation (O1) of the movement direction of the moving body (10) detected from the target image information. In the evaluation method, when the resolution is equal to or greater than the threshold (first threshold), the detailed model (medium resolution model 6, high resolution model 7) is selected from the plurality of trained models, at least one of the one or more pieces of overall information (I11-I13) of the moving object (10) detected from the target image information and the one or more pieces of part information (I21-I23, I31) are input to the detailed model (medium resolution model 6, high resolution model 7), and the detailed model (medium resolution model 6, high resolution model 7) outputs an evaluation (O2, O3) of the movement direction of the moving object (10-13) detected from the target image information. This aspect enables improvement in the accuracy of the evaluation of the movement direction of the moving object (10-13).
[0120] The second aspect is an evaluation method executed by an arithmetic circuit (23) that can access a storage device (22) that stores a plurality of trained models. The plurality of trained models include a simple model (low-resolution model 5) that has been trained to output an evaluation of the movement direction of a moving object (10) in response to input of one or more pieces of overall information (I11 to I13) based on the entirety (bounding boxes B, B1 to B3) of the moving object (10 to 13) in image information, and detailed models (medium-resolution model 6, high-resolution model 7) that have been trained to output an evaluation of the movement direction of the moving object (10) in response to input of at least one of the one or more pieces of overall information (I11 to I13) and one or more pieces of part information (I21 to I23, I31) based on one or more parts (R, P) of the moving object (10). The evaluation method includes confirming that the resolution of the image of the moving object (10-13) detected from the target image information is less than a threshold (first threshold), selecting the simple model (low-resolution model 5) from the plurality of trained models in response to confirming that the resolution is less than the threshold (first threshold), inputting the one or more pieces of overall information (I11-I13) of the moving object (10) detected from the target image information into the simple model (low-resolution model 5), and causing the simple model (low-resolution model 5) to output an evaluation (O1) of the movement direction of the moving object (10) detected from the target image information. This aspect enables improvement in the accuracy of the evaluation of the movement direction of the moving objects (10-13).
[0121] The third aspect is an evaluation method executed by an arithmetic circuit (23) that can access a storage device (22) that stores a plurality of trained models. The plurality of trained models include a simple model (low-resolution model 5) that has been trained to output an evaluation of the movement direction of a moving object (10) in response to input of one or more pieces of overall information (I11 to I13) based on the entirety (bounding boxes B, B1 to B3) of the moving object (10 to 13) in image information, and detailed models (medium-resolution model 6, high-resolution model 7) that have been trained to output an evaluation of the movement direction of the moving object (10) in response to input of at least one of the one or more pieces of overall information (I11 to I13) and one or more pieces of part information (I21 to I23, I31) based on one or more parts (R, P) of the moving object (10). The evaluation method includes confirming that the resolution of the image of the moving object (10-13) detected from the target image information is equal to or greater than a threshold (first threshold), selecting the detailed model (medium resolution model 6, high resolution model 7) from the plurality of trained models in response to confirming that the resolution is equal to or greater than the threshold (first threshold), inputting at least one of the one or more pieces of overall information (I11-I13) and the one or more pieces of part information (I21-I23, I31) of the moving object (10-13) detected from the target image information into the detailed model (medium resolution model 6, high resolution model 7), and outputting an evaluation (O2, O3) of the movement direction of the moving object (10) detected from the target image information from the detailed model (medium resolution model 6, high resolution model 7). This aspect enables improved accuracy in the evaluation of the movement direction of the moving object (10-13).
[0122] The fourth aspect is an evaluation method based on any one of the first to third aspects. In the fourth aspect, the resolution is determined based on the area of the bounding box (B, B1 to B3) of the moving object (10 to 13) detected from the target image information. This aspect does not require actual measurement of the distances (D1 to D3), thereby reducing the cost of equipment required to implement the evaluation method.
[0123] The fifth aspect is an evaluation method based on any one of the first to fourth aspects. In the fifth aspect, the one or more pieces of overall information (I11 to I13) include at least one of the position of the moving object (10 to 13), the speed of the moving object (10 to 13), and an image of the entire moving object (10 to 13) (bounding box B, B1 to B3). This aspect enables improved accuracy in evaluating the moving direction of the moving object (10 to 13).
[0124] A sixth aspect is an evaluation method based on any one of the first to fifth aspects. In the sixth aspect, the one or more pieces of body part information (I21, I22, I23, I31) include at least one of the following: the position of the one or more parts (R, P) of the moving object (10-13), the orientation of the one or more parts (R, P) of the moving object (10-13), an image of the one or more parts (R, P) of the moving object (10-13), and information based on the correlation between the multiple parts (R, P) of the moving object (10-13). This aspect enables improved accuracy in evaluating the movement direction of the moving object (10-13).
[0125] A seventh aspect is an evaluation method based on the sixth aspect. In the seventh aspect, the positions of the one or more parts (R, P) include the position of the face of the moving object (10 to 13). This aspect enables improvement in the accuracy of evaluation of the moving direction of the moving object (10 to 13).
[0126] The eighth aspect is an evaluation method based on the sixth or seventh aspect. In the eighth aspect, the orientation of the one or more parts (R, P) includes at least one of the orientation of the face of the moving object (10-13) and the line of sight of the moving object (10-13). This aspect enables improvement in the accuracy of evaluation of the moving direction of the moving object (10-13).
[0127] A ninth aspect is an evaluation method based on any one of the sixth to eighth aspects. In the ninth aspect, the image of the one or more regions (R, P) includes an image of the face of the moving object (10 to 13). This aspect enables improvement in the accuracy of evaluation of the moving direction of the moving object (10 to 13).
[0128] A tenth aspect is an evaluation method based on any one of the sixth to ninth aspects. In the tenth aspect, the information based on the correlation between the plurality of parts (R, P) includes information on the posture of the moving object (10 to 13). This aspect enables improvement in the accuracy of evaluation of the moving direction of the moving object (10 to 13).
[0129] An eleventh aspect is an evaluation method based on any one of the first to tenth aspects. In the tenth aspect, the evaluation of the movement direction of the moving body (10-13) includes a probability that the movement direction of the moving body (10-13) is a first movement direction and a probability that the movement direction of the moving body (10-13) is a second movement direction. The first movement direction is a direction in which the moving body (10-13) moves toward an object (mobile device 1). The second movement direction is a direction in which the moving body (10-13) avoids the object (mobile device 1). This aspect enables efficient movement of the object taking into account the movement direction of the moving body (10-13).
[0130] The twelfth aspect is an evaluation method executed by an arithmetic circuit (23) that can access a storage device (22) that stores a plurality of trained models. The plurality of trained models include a first model (medium resolution model 6) trained to output an evaluation (O2) of the movement direction of the moving body (10-13) in response to input of one or more pieces of overall information (I11-I13) based on the entire moving body (10-13) in the image information (bounding boxes B, B1-B3) and one or more pieces of first part information (I21-I23) based on one or more first parts (R) of the moving body (10-13), and a second model (high resolution model 7) trained to output an evaluation (O3) of the movement direction of the moving body (10-13) in response to input of at least one of the one or more pieces of overall information (I11-I13), at least one of the one or more pieces of first part information (I21-I23), and one or more pieces of second part information (I31) based on one or more second parts (P) of the moving body (10-13). The one or more second regions (P) are smaller than the one or more first regions (R). When the resolution of an image of the moving object (10-13) detected from the target image information is less than a threshold (second threshold), the evaluation method selects the first model (medium resolution model 6) from the plurality of trained models, inputs the one or more pieces of overall information (I11-I13) of the moving object (10-13) detected from the target image information and the one or more pieces of first region information (I21-I23) into the first model (medium resolution model 6), and causes the first model (medium resolution model 6) to output an evaluation (O2) of the movement direction of the moving object (10-13) detected from the target image information. The evaluation method includes selecting the second model (high-resolution model 7) from the plurality of trained models when the resolution is equal to or greater than the threshold (second threshold), inputting at least one of the one or more pieces of overall information (I11 to I13) of the moving body (10 to 13) detected from the target image information, at least one of the one or more pieces of first part information (I21 to I23), and the one or more pieces of second part information (I31) into the second model (high-resolution model 7), and outputting an evaluation (O3) of the movement direction of the moving body (10 to 13) detected from the target image information from the second model (high-resolution model 7).This aspect makes it possible to improve the accuracy of evaluation of the moving direction of the moving object (10 to 13).
[0131] The thirteenth aspect is an evaluation method executed by an arithmetic circuit (23) that can access a storage device (22) that stores a plurality of trained models. The plurality of trained models include a first model (medium resolution model 6) trained to output an evaluation (O2) of the movement direction of the moving body (10-13) in response to input of one or more pieces of overall information (I11-I13) based on the entire moving body (10-13) in the image information (bounding boxes B, B1-B3) and one or more pieces of first part information (I21-I23) based on one or more first parts (R) of the moving body (10-13), and a second model (high resolution model 7) trained to output an evaluation (O3) of the movement direction of the moving body (10-13) in response to input of at least one of the one or more pieces of overall information (I11-I13), at least one of the one or more pieces of first part information (I21-I23), and one or more pieces of second part information (I31) based on one or more second parts (P) of the moving body (10-13). The one or more second regions (P) are smaller than the one or more first regions (R). The evaluation method confirms that the resolution of the image of the moving object (10-13) detected from the target image information is less than a threshold (second threshold), and in response to confirming that the resolution is less than the threshold (second threshold), selects the first model (medium resolution model 6) from the multiple trained models, inputs the one or more pieces of overall information (I11-I13) of the moving object (10-13) detected from the target image information and the one or more pieces of first region information (I21-I23) into the first model (medium resolution model 6), and causes the first model (medium resolution model 6) to output an evaluation (O2) of the movement direction of the moving object (10-13) detected from the target image information. This aspect enables improved accuracy in the evaluation of the movement direction of the moving object (10-13).
[0132] The fourteenth aspect is an evaluation method executed by an arithmetic circuit (23) that can access a storage device (22) that stores a plurality of trained models. The plurality of trained models include a first model (medium resolution model 6) trained to output an evaluation (O2) of the movement direction of the moving body (10-13) in response to input of one or more pieces of overall information (I11-I13) based on the entire moving body (10-13) in the image information (bounding boxes B, B1-B3) and one or more pieces of first part information (I21-I23) based on one or more first parts (R) of the moving body (10-13), and a second model (high resolution model 7) trained to output an evaluation (O3) of the movement direction of the moving body (10-13) in response to input of at least one of the one or more pieces of overall information (I11-I13), at least one of the one or more pieces of first part information (I21-I23), and one or more pieces of second part information (I31) based on one or more second parts (P) of the moving body (10-13). The one or more second regions (P) are smaller than the one or more first regions (R). The evaluation method confirms that the resolution of the image of the moving object (10-13) detected from the target image information is equal to or greater than a threshold (second threshold), selects the second model (high-resolution model 7) from the plurality of trained models in response to confirming that the resolution is equal to or greater than the second threshold, inputs at least one of the one or more pieces of overall information (I11-I13) of the moving object (10-13) detected from the target image information, at least one of the one or more pieces of first region information (I21-I23), and the one or more pieces of second region information (I31) into the second model (high-resolution model 7), and causes the second model (high-resolution model 7) to output an evaluation (O3) of the movement direction of the moving object (10-13) detected from the target image information. This aspect enables improved accuracy in the evaluation of the movement direction of the moving object (10-13).
[0133] A fifteenth aspect is a program for causing the arithmetic circuit (23) to execute the evaluation method based on any one of the first to fourteenth aspects. This aspect enables improvement in the accuracy of evaluation of the moving direction of the moving object (10 to 13).
[0134] A sixteenth aspect is an evaluation system (2) comprising: a storage device (22) that stores a plurality of trained models; and an arithmetic circuit (23) that can access the storage device (22). The plurality of trained models include a simple model (low-resolution model 5) that has been trained to output an evaluation of the movement direction of a moving object (10) in response to input of one or more pieces of overall information (I11 to I13) based on the entirety (bounding boxes B, B1 to B3) of the moving object (10 to 13) in image information; and detailed models (medium-resolution model 6, high-resolution model 7) that have been trained to output an evaluation of the movement direction of the moving object (10) in response to input of at least one of the one or more pieces of overall information (I11 to I13) and one or more pieces of part information (I21 to I23, I31) based on one or more parts (R, P) of the moving object (10). When the resolution of the image of the moving body (10 to 13) detected from the target image information is less than a threshold (first threshold), the arithmetic circuit (23) selects the simple model (low-resolution model 5) from the plurality of trained models, inputs the one or more pieces of overall information (I11 to I13) of the moving body (10 to 13) detected from the target image information to the simple model (low-resolution model 5), and causes the simple model (low-resolution model 5) to output an evaluation (O1) of the movement direction of the moving body (10) detected from the target image information. When the resolution is equal to or greater than the threshold (first threshold), the arithmetic circuit (23) selects the detailed model (medium resolution model 6, high resolution model 7) from the plurality of trained models, inputs at least one of the one or more pieces of overall information (I11 to I13) of the moving object (10) detected from the target image information and the one or more pieces of part information (I21 to I23, I31) to the detailed model (medium resolution model 6, high resolution model 7), and causes the detailed model (medium resolution model 6, high resolution model 7) to output an evaluation (O2, O3) of the movement direction of the moving object (10 to 13) detected from the target image information. This aspect enables improvement in the accuracy of the evaluation of the movement direction of the moving object (10 to 13).
[0135] A seventeenth aspect is an evaluation system (2) comprising a storage device (22) that stores a plurality of trained models, and an arithmetic circuit (23) that can access the storage device (22). The plurality of trained models include a first model (medium resolution model 6) trained to output an evaluation (O2) of the movement direction of the moving body (10-13) in response to input of one or more pieces of overall information (I11-I13) based on the entire moving body (10-13) in the image information (bounding boxes B, B1-B3) and one or more pieces of first part information (I21-I23) based on one or more first parts (R) of the moving body (10-13), and a second model (high resolution model 7) trained to output an evaluation (O3) of the movement direction of the moving body (10-13) in response to input of at least one of the one or more pieces of overall information (I11-I13), at least one of the one or more pieces of first part information (I21-I23), and one or more pieces of second part information (I31) based on one or more second parts (P) of the moving body (10-13). The one or more second regions (P) are smaller than the one or more first regions (R). When the resolution of the image of the moving object (10-13) detected from the target image information is less than a threshold (second threshold), the arithmetic circuit (23) selects the first model (medium resolution model 6) from the plurality of trained models, inputs the one or more pieces of overall information (I11-I13) of the moving object (10-13) detected from the target image information and the one or more pieces of first region information (I21-I23) to the first model (medium resolution model 6), and causes the first model (medium resolution model 6) to output an evaluation (O2) of the movement direction of the moving object (10-13) detected from the target image information.When the resolution is equal to or greater than the threshold (second threshold), the arithmetic circuit (23) selects the second model (high-resolution model 7) from the plurality of trained models, inputs at least one of the one or more pieces of overall information (I11-I13) of the moving object (10-13) detected from the target image information, at least one of the one or more pieces of first part information (I21-I23), and the one or more pieces of second part information (I31) to the second model (high-resolution model 7), and causes the second model (high-resolution model 7) to output an evaluation (O3) of the movement direction of the moving object (10-13) detected from the target image information. This aspect enables improvement in the accuracy of the evaluation of the movement direction of the moving object (10-13).
[0136] The second to eleventh aspects can also be modified appropriately and applied to the sixteenth or seventeenth aspect.
[0137] [4. Terminology] In this disclosure, terms related to machine learning are defined and used as follows:
[0138] A "trained model" refers to an "inference program" that incorporates "trained parameters."
[0139] "Trained parameters" refer to parameters (coefficients) obtained as a result of learning using a training dataset. Trained parameters are generated by inputting the training dataset into a training program and mechanically adjusting them for a specific purpose. Although trained parameters are adjusted to suit the purpose of learning, they are simply parameters (numerical information, etc.) on their own, and only function as a trained model when incorporated into an inference program. For example, in the case of deep learning, the main trained parameters are parameters used to weight the links between each node.
[0140] An "inference program" is a program that can output a certain result for an input by applying built-in trained parameters. For example, it is a program that specifies a series of calculation procedures for applying trained parameters acquired as a result of training to an image given as input and outputting a result (authentication or judgment) for that image.
[0141] A "learning dataset," also known as a training dataset, refers to secondary processed data that has been generated to facilitate analysis using the target learning method by converting and processing raw data through preprocessing such as removing missing values and outliers, adding separate data such as label information (ground truth data), or a combination of these. Training datasets may also include data that has been "padded" by applying certain transformations to the raw data.
[0142] "Raw data" refers to data that is primarily acquired by users, vendors, other businesses, research institutions, etc., and that has been converted and processed so that it can be loaded into a database.
[0143] A "learning program" is a program that executes an algorithm to find certain rules from a training dataset and generate a model that expresses those rules. Specifically, this refers to a program that specifies the procedures to be executed by a computer in order to realize learning using the adopted learning method. [Industrial Applicability]
[0144] The present disclosure is applicable to an evaluation method, a program (computer program), and an evaluation system. Specifically, the present disclosure is applicable to an evaluation method, a program (computer program), and an evaluation system related to evaluation of the movement direction of a moving object captured in an image. [Explanation of symbols]
[0145] 1. Mobile device (object) 2. Rating System 22 Storage device 23 Arithmetic circuit 5 Low-resolution model (simple model) 6 Medium resolution model (detailed model, first model) 7 High-resolution model (detailed model, second model) 10,11~13 Mobile B, B1~B3 bounding box R site (1st site) P part (second part) D1~D3 distance I11~I13 General Information I21~I23 Part information (1st part information) I31 Part Information (Second Part Information) O1~O3 rating V1 1st movement direction V2 2nd movement direction
Claims
1. An evaluation method executed by a calculation circuit that can access a storage device that stores a plurality of trained models, The plurality of trained models are a simple model that is trained to output an evaluation of the direction of movement of a moving object in response to input of one or more pieces of overall information based on the entire moving object in image information; a detailed model that is trained to output an evaluation of the moving direction of the moving object in response to input of at least one of the one or more pieces of overall information and one or more pieces of part information based on one or more parts of the moving object; Including, The evaluation method includes: When the resolution of the image of the moving object detected from the target image information is less than a threshold, the simple model is selected from the plurality of trained models, the one or more pieces of overall information of the moving object detected from the target image information are input to the simple model, and an evaluation of the movement direction of the moving object detected from the target image information is output from the simple model; If the resolution is equal to or greater than the threshold, the detailed model is selected from the plurality of trained models, and at least one of the one or more pieces of overall information and the one or more pieces of part information of the moving object detected from the target image information is input to the detailed model, and an evaluation of the moving direction of the moving object detected from the target image information is output from the detailed model. Evaluation method.
2. An evaluation method executed by a calculation circuit that can access a storage device that stores a plurality of trained models, The plurality of trained models are a simple model that is trained to output an evaluation of the direction of movement of a moving object in response to input of one or more pieces of overall information based on the entire moving object in image information; a detailed model that is trained to output an evaluation of the moving direction of the moving object in response to input of at least one of the one or more pieces of overall information and one or more pieces of part information based on one or more parts of the moving object; Including, The evaluation method includes: confirming that the resolution of the image of the moving object detected from the target image information is less than a threshold; selecting the simplified model from the plurality of trained models in response to determining that the resolution is less than the threshold; inputting the one or more pieces of overall information of the moving object detected from the target image information into the simple model, and causing the simple model to output an evaluation of the moving direction of the moving object detected from the target image information; Evaluation method.
3. An evaluation method executed by a calculation circuit that can access a storage device that stores a plurality of trained models, The plurality of trained models are a simple model that is trained to output an evaluation of the direction of movement of a moving object in response to input of one or more pieces of overall information based on the entire moving object in image information; a detailed model that is trained to output an evaluation of the moving direction of the moving object in response to input of at least one of the one or more pieces of overall information and one or more pieces of part information based on one or more parts of the moving object; Including, The evaluation method includes: confirming that the resolution of the image of the moving object detected from the target image information is equal to or greater than a threshold; selecting the detailed model from the plurality of trained models in response to confirming that the resolution is equal to or greater than the threshold; inputting at least one of the one or more pieces of overall information and the one or more pieces of part information of the moving object detected from the target image information into the detailed model, and causing the detailed model to output an evaluation of the moving direction of the moving object detected from the target image information; Evaluation method.
4. the resolution is determined based on the area of a bounding box of the moving object detected from the target image information. The evaluation method according to any one of claims 1 to 3.
5. The one or more pieces of overall information include at least one of a position of the moving object, a speed of the moving object, and an image of the entire moving object. The evaluation method according to any one of claims 1 to 3.
6. the one or more pieces of part information include at least one of: a position of the one or more parts of the moving body; an orientation of the one or more parts of the moving body; an image of the one or more parts of the moving body; and information based on a correlation between the one or more parts of the moving body. The evaluation method according to any one of claims 1 to 3.
7. the positions of the one or more parts include a position of a face of the moving object; The evaluation method according to claim 6.
8. The orientation of the one or more parts includes at least one of the orientation of a face of the moving object and a line of sight of the moving object. The evaluation method according to claim 6.
9. The image of the one or more parts includes an image of a face of the moving object. The evaluation method according to claim 6.
10. the information based on the correlation between the plurality of parts includes information about the posture of the moving body; The evaluation method according to claim 6.
11. the evaluation of the moving direction of the moving object includes a probability that the moving direction of the moving object is a first moving direction and a probability that the moving direction of the moving object is a second moving direction; the first movement direction is a direction in which the moving body moves toward a target object, the second movement direction is a direction in which the moving body avoids an object; The evaluation method according to any one of claims 1 to 3.
12. An evaluation method executed by a calculation circuit that can access a storage device that stores a plurality of trained models, The plurality of trained models are a first model that is trained to output an evaluation of the moving direction of a moving object in response to input of one or more pieces of overall information based on the entire moving object in image information and one or more pieces of first part information based on one or more first parts of the moving object; a second model that is trained to output an evaluation of the moving direction of the moving object in response to input of at least one of the one or more pieces of overall information, at least one of the one or more pieces of first part information, and one or more pieces of second part information based on one or more second parts of the moving object; Including, the one or more second portions are smaller than the one or more first portions; The evaluation method includes: When the resolution of the image of the moving object detected from the target image information is less than a threshold, the first model is selected from the plurality of trained models, the one or more pieces of overall information of the moving object detected from the target image information and the one or more pieces of first part information are input to the first model, and an evaluation of the moving direction of the moving object detected from the target image information is output from the first model; When the resolution is equal to or greater than the threshold, the second model is selected from the plurality of trained models, and at least one of the one or more pieces of overall information of the moving object detected from the target image information, at least one of the one or more pieces of first part information, and the one or more pieces of second part information are input to the second model, and an evaluation of the moving direction of the moving object detected from the target image information is output from the second model. Evaluation method.
13. An evaluation method executed by a calculation circuit that can access a storage device that stores a plurality of trained models, The plurality of trained models are a first model that is trained to output an evaluation of the moving direction of a moving object in response to input of one or more pieces of overall information based on the entire moving object in image information and one or more pieces of first part information based on one or more first parts of the moving object; a second model that is trained to output an evaluation of the moving direction of the moving object in response to input of at least one of the one or more pieces of overall information, at least one of the one or more pieces of first part information, and one or more pieces of second part information based on one or more second parts of the moving object; Including, the one or more second portions are smaller than the one or more first portions; The evaluation method includes: confirming that the resolution of the image of the moving object detected from the target image information is less than a threshold; selecting the first model from the plurality of trained models in response to determining that the resolution is less than the threshold; inputting the one or more pieces of overall information of the moving object detected from the target image information and the one or more pieces of first part information into the first model, and causing the first model to output an evaluation of the moving direction of the moving object detected from the target image information; Evaluation method.
14. An evaluation method executed by a calculation circuit that can access a storage device that stores a plurality of trained models, The plurality of trained models are a first model that is trained to output an evaluation of the moving direction of a moving object in response to input of one or more pieces of overall information based on the entire moving object in image information and one or more pieces of first part information based on one or more first parts of the moving object; a second model that is trained to output an evaluation of the moving direction of the moving object in response to input of at least one of the one or more pieces of overall information, at least one of the one or more pieces of first part information, and one or more pieces of second part information based on one or more second parts of the moving object; Including, the one or more second portions are smaller than the one or more first portions; The evaluation method includes: confirming that the resolution of the image of the moving object detected from the target image information is equal to or greater than a threshold; selecting the second model from the plurality of trained models in response to confirming that the resolution is equal to or greater than the threshold; inputting at least one of the one or more pieces of overall information of the moving object detected from the target image information, at least one of the one or more pieces of first part information, and the one or more pieces of second part information into the second model, and causing the second model to output an evaluation of the moving direction of the moving object detected from the target image information; Evaluation method.
15. In order to cause the arithmetic circuit to execute the evaluation method according to any one of claims 1 to 3 and 12 to 14, program.
16. a storage device that stores a plurality of trained models; an arithmetic circuit that can access the storage device; Equipped with The plurality of trained models are a simple model that is trained to output an evaluation of the direction of movement of a moving object in response to input of one or more pieces of overall information based on the entire moving object in image information; a detailed model that is trained to output an evaluation of the moving direction of the moving object in response to input of at least one of the one or more pieces of overall information and one or more pieces of part information based on one or more parts of the moving object; Including, The arithmetic circuit comprises: When the resolution of the image of the moving object detected from the target image information is less than a threshold, the simple model is selected from the plurality of trained models, the one or more pieces of overall information of the moving object detected from the target image information are input to the simple model, and an evaluation of the movement direction of the moving object detected from the target image information is output from the simple model; If the resolution is equal to or greater than the threshold, the detailed model is selected from the plurality of trained models, and at least one of the one or more pieces of overall information and the one or more pieces of part information of the moving object detected from the target image information is input to the detailed model, and an evaluation of the moving direction of the moving object detected from the target image information is output from the detailed model. Rating system.
17. a storage device that stores a plurality of trained models; an arithmetic circuit that can access the storage device; Equipped with The plurality of trained models are a first model that is trained to output an evaluation of the moving direction of a moving object in response to input of one or more pieces of overall information based on the entire moving object in image information and one or more pieces of first part information based on one or more first parts of the moving object; a second model that is trained to output an evaluation of the moving direction of the moving object in response to input of at least one of the one or more pieces of overall information, at least one of the one or more pieces of first part information, and one or more pieces of second part information based on one or more second parts of the moving object; Including, the one or more second portions are smaller than the one or more first portions; The arithmetic circuit comprises: When the resolution of the image of the moving object detected from the target image information is less than a threshold, the first model is selected from the plurality of trained models, the one or more pieces of overall information of the moving object detected from the target image information and the one or more pieces of first part information are input to the first model, and an evaluation of the moving direction of the moving object detected from the target image information is output from the first model; When the resolution is equal to or greater than the threshold, the second model is selected from the plurality of trained models, and at least one of the one or more pieces of overall information of the moving object detected from the target image information, at least one of the one or more pieces of first part information, and the one or more pieces of second part information are input to the second model, and an evaluation of the moving direction of the moving object detected from the target image information is output from the second model. Rating system.
Citation Information
Patent Citations
Image processing device and image processing method
JP2021077092A
Recognition system, recognition method, program, learning method, trained model, distillation model and training data set generation method
WO2022097371A1