Image processing device, method and program and learning data generator

The image processing apparatus addresses the inefficiencies in generating learning data by evaluating the probability and accuracy of crossing probabilities for objects in images, resulting in efficient and high-quality data for machine learning models.

JP2025082929APending Publication Date: 2025-05-30DENSO TEN LTD
View PDF 1 Cites 0 Cited by

Patent Information

Application Number
JP2023196507
Authority / Receiving Office
JP · JP
Patent Type
Applications
Current Assignee / Owner
Filing Date
2023-11-20
Publication Date
2025-05-30

AI Technical Summary

Technical Problem

Existing methods for generating learning data for machine learning models, such as detecting people in images and inferring crossing probabilities, face inefficiencies due to manual annotation being labor-intensive and prone to errors, while automated methods may introduce inaccuracies.

Method used

An image processing apparatus and method that evaluates the probability and accuracy of crossing probabilities for objects, particularly people, in images relative to roads, using information on object position and orientation, to efficiently generate high-quality learning data.

Benefits of technology

This approach enables the efficient generation of large amounts of high-quality learning data by accurately determining crossing probabilities, reducing manual labor and error rates, and improving the quality of training data for machine learning models.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure 2025082929000001_ABST
    Figure 2025082929000001_ABST
Patent Text Reader

Abstract

To contribute to efficient generation of a large amount of learning data with high quality.SOLUTION: An input image (IN1) in which an object including a person (P [i]) and a road (RD) are reflected is inputted to an image processing device. A controller in the image processing device evaluates a correctness degree of a probability added to the input image, that is, a crossing probability of the object to the road, based on information on a position and a direction of the object relative to the road.SELECTED DRAWING: Figure 10
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention relates to an image processing apparatus, method, and program, and a learning data generation apparatus.

Background Art

[0002] To perform machine learning of AI, a large amount of learning data is required. Various techniques have been proposed for performing annotation in the generation of learning data at low cost and with high accuracy (see, for example, Patent Document 1 below).

Prior Art Documents

Patent Documents

[0003]

Patent Document 1

Summary of the Invention

Problems to be Solved by the Invention

[0004] It is possible to generate a model (learning model) for detecting a person in an image and inferring a crossing probability representing the possibility that the person crosses a road. At this time, by preparing learning data in which the position of the person in the image and the crossing probability of the person to the lane are defined, and performing machine learning using the learning data, a model capable of inferring the crossing probability can be generated. By incorporating the model into a system for performing vehicle driving support or the like, various controls become possible. Here, the road mainly refers to a lane with a crosswalk or the like added, and a person can cross the lane through a crosswalk or the like. There are also lanes without a crosswalk or the like added. Hereinafter, the road where a person can cross is mainly represented by a lane.

[0005] When generating learning data, the method of setting all crossing probabilities manually by human work is inefficient because the amount of human work becomes enormous due to the large amount of learning data required for learning. On the other hand, in the method of setting crossing probabilities by automatic determination by a machine when generating learning data, errors may be mixed in the set values of the crossing probabilities, so it is difficult to ensure the quality of the learning data. In order to efficiently generate a large amount of high-quality learning data, a technique that contributes to an appropriate combination of the former method and the latter method is required.

[0006] An object of the present invention is to provide an image processing apparatus, a method, a program, and a learning data generation apparatus that contribute to efficient generation of a large amount of high-quality learning data.

Means for Solving the Problem

[0007] The image processing apparatus according to the present invention, for an input image in which an object including a person and a road are shown, Based on the information on the position and orientation of the object with respect to the road, the probability added to the input image and the accuracy of the crossing probability of the object to the road are evaluated.

Effect of the Invention

[0008] According to the present invention, the probability added to the input image and the accuracy of the crossing probability of the object to the road are evaluated, and the crossing probability of the learning data can be determined according to the accuracy. This contributes to efficient generation of a large amount of high-quality learning data.

Brief Description of the Drawings

[0009]

Figure 1

Figure 2

Figure 3

Figure 4

Figure 5

Figure 6

Figure 7

Figure 8

Figure 9

Figure 10

Figure 11

Figure 12

Figure 13

Figure 14

Figure 15

Figure 16

Figure 17

Best Mode for Carrying Out the Invention

[0010] Hereinafter, examples of embodiments of the present invention will be specifically described with reference to the drawings. In each of the drawings referred to, the same parts are denoted by the same reference numerals, and redundant descriptions regarding the same parts are omitted in principle. In this specification, for the sake of simplicity of description, by writing symbols or signs referring to information, signals, physical quantities, functional units, circuits, elements, or parts, etc., the names of the information, signals, physical quantities, functional units, circuits, elements, or parts, etc. corresponding to the symbols or signs may be omitted or abbreviated. In this specification, AI is an abbreviation for artificial intelligence.

[0011] The outline of the technology according to this embodiment will be described with reference to FIG. 1. In a system that performs driving assistance for a vehicle, one of the tasks that AI can perform is predicting the behavior of a person. To realize the prediction of a person's behavior in AI, a large amount of training data (teacher data) in which the position of a person on the image and the probability of the person crossing into the lane are defined is required, together with an image including the image of the person (in other words, an image including the image data of the person). As described above, the lane refers to a road that a person can cross.

[0012] The training data generation device 10 generates the above training data based on the input image. The input image supplied to the training data generation device 10 is referred to as the input image IN. The input image IN is a two-dimensional image obtained by photographing the inside of the photographing area where a person and a lane are located. Therefore, the input image IN includes the image of the person and the image of the lane (in other words, the image data of the person and the lane). That is, the input image IN is an image in which a person and a road are reflected.

[0013] Fig. 2 shows a top view of the cooperative vehicle V1 for obtaining the input image IN. The cooperative vehicle V1 is an automobile running on an arbitrary lane. The cooperative vehicle V1 runs forward. A camera C1 is installed on the cooperative vehicle V1. The camera C1 may be a camera of a drive recorder installed on the cooperative vehicle V1. The camera C1 generates a two-dimensional image by photographing within a predetermined photographing area SR1. The two-dimensional image obtained by photographing with the camera C1 is used as the input image IN. However, not all two-dimensional images obtained by photographing with the camera C1 become the input image IN. In Fig. 2, the hatched area corresponds to the photographing area SR1 of the camera C1. The photographing area SR1 includes an area outside the cooperative vehicle V1 and located in front of the cooperative vehicle V1. When the cooperative vehicle V1 runs on the lane, the road surface of that lane is included in the photographing area SR1, and the sidewalk adjacent to that lane may be included in the photographing area SR1.

[0014] Fig. 3 shows an example of the input image IN. The input image IN is a two-dimensional image formed by arranging a plurality of pixels in a matrix in the X-axis and Y-axis directions respectively. The X-axis and Y-axis are orthogonal to each other. The two-dimensional coordinate plane with the X-axis and Y-axis as coordinate axes is called the XY coordinate plane. The input image IN is defined on the two-dimensional coordinate plane. The X-axis is parallel to the horizontal direction of the input image IN, and the Y-axis is parallel to the vertical direction of the input image IN. In the input image IN, the direction toward the positive side of the X-axis is the right direction, and the direction toward the negative side of the X-axis is the left direction. In the input image IN, the direction toward the positive side of the Y-axis is the downward direction, and the direction toward the negative side of the Y-axis is the upward direction. The right and left directions in the input image IN shall be the same as the right and left directions as seen from the driver of the cooperative vehicle V1. Therefore, regarding any target object within the photographing area SR, the position of the target object in the input image IN moves to the right as the target object moves to the right as seen from the driver of the cooperative vehicle V1. Conversely, the position of the target object in the input image IN moves to the left as the target object moves to the left as seen from the driver of the cooperative vehicle V1.

[0015] In the shooting area SR1 of the camera C1, there are included the road RD located in front of the cooperative vehicle V1, the sidewalk SW_L located on the left side of the road RD, and the sidewalk SW_R located on the right side of the road RD. Therefore, the input image IN includes the image of the road RD, the image of the sidewalk SW_L, and the image of the sidewalk SW_R. However, either one of the images of the sidewalk SW_L and the sidewalk SW_R may not be included in the input image IN. Hereinafter, unless otherwise specified, it is assumed that the input image IN includes the image of the road RD, the image of the sidewalk SW_L, and the image of the sidewalk SW_R. Note that including the image of the road RD in the input image IN means, in other words, that the input image IN includes the image data of the road RD. The same applies to the sidewalks SW_L and SW_R and any person described later.

[0016] Also, the crosswalk PX may be included in the shooting area SR1 of the camera C1. In this case, the input image IN also includes the image of the crosswalk PX. The input image IN according to the example of FIG. 3 also includes the image of the crosswalk PX. The crosswalk PX is an area formed within the road RD and is provided for pedestrians and the like to cross the road RD.

[0017] Furthermore, one or more pedestrians and the like are included in the shooting area SR1. Therefore, the input image IN also includes the images of the one or more pedestrians and the like. Pedestrians and the like include not only pedestrians but also persons moving or stationary on a bicycle or a kick scooter.

[0018] The camera C1 is installed on the cooperative vehicle V1 so that the optical axis of the camera C1 faces the front of the cooperative vehicle V1. Therefore, the input image IN has the following characteristics. That is, in the input image IN, the image of the road RD is located in the area including the center of the input image IN. The area where the image of the road RD is located is called the road area. In the input image IN, the image of the sidewalk SW_L is located on the left side of the road area, and the image of the sidewalk SW_R is located on the right side of the road area. Therefore, pedestrians and the like on the sidewalk SW_L are located near the left end of the input image IN, and pedestrians and the like on the sidewalk SW_R are located near the right side of the input image IN.

[0019] The input image IN acquired by the camera C1 is supplied to the learning data generation device 10 via an arbitrary communication line or via an arbitrary recording medium. An interface 20 is connected to the learning data generation device 10 wirelessly or by wire. The interface 20 is a machine interface between the operator OP and the learning data generation device 10. The interface 20 includes a display device 21 and an input device 22. The display content of the display device 21 is visible to the operator OP. The learning data generation device 10 can display an image based on the input image IN on the display device 21 as necessary. The input device 22 is composed of a keyboard, a pointing device, etc., and receives an input of an arbitrary operation from the operator OP. Note that at least one of the display device 21 and the input device 22 may be a device provided in the learning data generation device 10.

[0020] The learning data generation device 10 generates learning data. At this time, the learning data generation device 10 detects a bounding box of a pedestrian or the like in the input image IN based on the input image IN, and determines and sets a crossing probability of a pedestrian or the like with respect to the road RD. The bounding box is hereinafter referred to as BBOX. The learning data generation device 10 performs object detection processing. In the object detection processing, a rectangular area in which a recognition target object is determined to exist in the input image IN is set as the BBOX. The recognition target object is a pedestrian or the like. Therefore, the BBOX of a pedestrian or the like is detected by the object detection processing. When the learning data generation device 10 detects the BBOX of a pedestrian or the like in the object detection processing, it generates BBOX information for specifying the position and shape of the BBOX in the input image IN. From the viewpoint that various image processes can be performed on the input image IN, it can be said that the learning data generation device 10 corresponds to an image processing device or includes an image processing device.

[0021] The training data includes the input image IN and the BBOX information, as well as the crossing probability information indicating the crossing probability of a pedestrian or the like. That the training data includes the input image IN specifically means that the image data of the input image IN is included in the training data. The crossing probability of a pedestrian or the like is the probability that a pedestrian or the like crosses the road RD. Note that the determination of the crossing probability by the training data generation device 10 corresponds to the estimation of the crossing probability. The training data generation device 10 generally makes a determination and setting of the crossing probability without relying on the input information from the operator OP. However, when it is presumed that it is difficult to determine the crossing probability, the operator OP is requested to input the crossing probability. Therefore, in the training data generation device 10, there are cases where the crossing probability is set without relying on the input information from the operator OP and cases where the crossing probability is set based on the input information from the operator OP. The set crossing probability is defined in the crossing probability information.

[0022] The training data is stored in the database 30. The database 30 consists of one or more recording media. By accessing the database 30, the training data generation device 10 stores the training data it has generated in the database 30. However, the training data generated by the training data generation device 10 may be stored in the database 30 via any other computer device.

[0023] The training data is labeled data that has the input image IN as the problem data and the BBOX information and the crossing probability information as the correct labels (which may exist in multiple for one input image IN). Therefore, the training data can also be referred to as a training data set. The training data is supplied to the learning device 40. The learning device 40 has an AI model 41 and uses the training data to train the AI model 41. The learning described in this embodiment refers to supervised machine learning. To distinguish between the image supplied to the training data generation device 10 and the image supplied to the learning device 40, the image supplied to the learning device 40 can also be referred to as a learning image. That is, the learning image is the same as the input image IN, but when the input image IN is used for the training of the AI model 41, the input image IN is referred to as a learning image.

[0024] So far, we have focused on a single input image IN. In reality, however, multiple input images IN are supplied to the learning data generation device 10, and learning data is generated for each input image IN. Therefore, learning data for multiple input images IN is generated, and the AI model 41 is trained based on the learning data for multiple learning images (i.e., the learning data set). The multiple input images IN may be obtained by photographing with the camera C1 of a single cooperative vehicle V1. There may be multiple cooperative vehicles V1. In this case, multiple input images IN may be generated using the multiple cameras C1 in the multiple cooperative vehicles V1.

[0025] The learning device 40 generates a learned AI model 41 by training the AI model 41 based on the learning data. The learned AI model 41 is an algorithm that infers the position, shape, and crossing probability of a pedestrian or the like on the two-dimensional image when a two-dimensional image of the same type as the input image IN is given. By incorporating the learned AI model 41 into a driving support system or the like in an arbitrary vehicle, it is possible to accurately predict the behavior of a person.

[0026] Fig. 4 shows a block diagram of the configuration of the learning data generation device 10. The learning data generation device 10 is composed of one or more server devices connected to a communication network including the Internet. The learning data generation device 10 may be configured using cloud computing. However, the learning data generation device 10 may be installed in the cooperative vehicle V1. The learning data generation device 10 includes a communication unit 11, a storage unit 12, and a controller 13.

[0027] The communication unit 11 transmits and receives arbitrary signals to and from a counterpart device different from the learning data generation device 10. The counterpart device for the communication unit 11 includes an input image supply device (not shown) that supplies the input image IN to the learning data generation device 10, the interface 20, and the database 30. Note that the controller 13 can transmit and receive arbitrary information to and from the counterpart device (the counterpart device for the communication unit 11) using the communication unit 11. However, the description of the communication unit 11 may be omitted below.

[0028] The storage unit 12 is configured to include a non-volatile memory such as a ROM (Read Only Memory) or a flash memory, and a volatile memory such as a RAM (Random Access Memory). In the storage unit 12, in addition to storing each data referred to by the controller 13, various programs to be executed by the controller 13 are stored.

[0029] The controller 13 comprehensively controls the operations of each part in the learning data generation device 10. The controller 13 includes an arithmetic processing unit including a CPU (Central Processing Unit) and a GPU (Graphics Processing Unit) as hardware resources. The controller 13 has function blocks 131 to 139.

[0030] The controller 13 is a program execution device (computer) capable of executing an arbitrary program. By the controller 13 executing the target program, each function of the controller 13 (including the functions of the function blocks 131 to 139) is realized. The target program may be a program stored in the storage unit 12. The target program may be provided to the learning data generation device 10 from an external device (not shown) provided outside the learning data generation device 10 and connected to the learning data generation device 10 wirelessly or by wire. All the operations of the controller 13 described in this embodiment may be operations realized by the controller 13 executing the target program. The target program may be composed of a plurality of programs.

[0031] The function blocks 131 to 139 will be described. The function blocks 131, 132, 133, 134, and 135 are an input image acquisition unit, an object detection unit, an orientation detection unit, a region detection unit, and a distance detection unit, respectively. The function blocks 136, 137, 138, and 139 are a probability determination unit, an evaluation unit, a probability setting unit, and a data output unit, respectively.

[0032] The input image acquisition unit 131 executes an input image acquisition process, and acquires an input image IN from the above-mentioned input image supply device (not shown) in the input image acquisition process. Note that the acquisition, supply, and generation of the input image IN specifically refer to the acquisition, supply, and generation of the image data of the input image IN. The same applies to images other than the input image IN.

[0033] The object detection unit 132 executes an object detection process based on the input image IN. By the object detection process, a recognition target object in the input image IN is detected, and a BBOX is set for each detected recognition target object. As described above, the recognition target object is a pedestrian or the like. When the object detection unit 132 sets a BBOX of a recognition target object in the object detection process, it generates BBOX information for specifying the position and shape of the BBOX in the input image IN. Various object detection AIs for realizing the object detection process have been put into practical use, and the object detection unit 132 can be formed by the object detection AI.

[0034] Fig. 5 shows an input image IN1 which is an example of the input image IN. The input image IN1 includes images of five persons P[1] to P[5]. Also, similar to the input image IN in Fig. 3, the input image IN1 in Fig. 5 includes an image of a road RD, an image of a sidewalk SW_L located on the left side of the road RD, and an image of a sidewalk SW_R located on the right side of the road RD. Also, a crosswalk PX is formed in the road RD in the input image IN1. The object detection unit 132 sets a BBOX for each recognition target object (therefore, for each person) in the object detection process for the input image IN. Therefore, BBOXes are set for each of the persons P[1] to P[5] with respect to the input image IN1. The BBOXes set for the persons P[1] to P[5] are referred to as BBOX[1] to BBOX[5], respectively. BBOX[i] is a rectangular area in which it is determined that the person P[i] exists in the input image IN1 (that is, a rectangular area in which it is determined that the image data of the person P[i] exists). Ignoring the error of this determination, BBOX[i] is a rectangular area in which the person P[i] exists in the input image IN1 (that is, a rectangular area in which the image data of the person P[i] exists). Hereinafter, the error of the determination in the object detection process is ignored unless necessary. i represents an arbitrary integer.

[0035] The BBOX information about BBOX[i] specifies the position and shape of BBOX[i] in the input image IN1. More specifically, for example, the BBOX information about BBOX[i] indicates the origin coordinates, width, and height of BBOX[i] in the input image IN1. The origin coordinates of BBOX[i] refer to the coordinates of the vertex (coordinates on the XY coordinate plane) that is located on the negative side of the X-axis and the negative side of the Y-axis among the four vertices of the rectangle that is the outer shape of BBOX[i]. The width of BBOX[i] is the length of BBOX[i] in the X-axis direction and is expressed by the number of pixels of BBOX[i] in the X-axis direction. The height of BBOX[i] is the length of BBOX[i] in the Y-axis direction and is expressed by the number of pixels of BBOX[i] in the Y-axis direction.

[0036] Basically, the object to be recognized may be only a person (in other words, a pedestrian). The operation examples and the like when the object to be recognized includes an object other than a person will be described later. Hereinafter, unless otherwise specified, the object to be recognized is assumed to be a person (in other words, a pedestrian).

[0037] The orientation detection unit 133 executes orientation detection processing based on the input image IN. In the orientation detection processing, the orientation detection unit 133 detects the orientation of the person in the input image IN for each person. The detected orientation of the person is the orientation of the person's body. Alternatively, the detected orientation of the person is the orientation of the person's face. Hereinafter, for convenience of description, the orientation of the person's body may be referred to as the body orientation, and the orientation of the person's face may be referred to as the face orientation. In the orientation detection processing, both the body orientation and the face orientation of the person may be detected. The body orientation and the face orientation detected in the orientation detection processing are defined on the XY coordinate plane.

[0038] Any person on the input image IN is referred to as the target person, and the body orientation and face orientation of the target person are described. The body orientation of the target person refers to the direction in which the upper body and the lower body of the target person are facing on the input image IN. Assume an axis that is perpendicular to the midline of the target person, is located on the sagittal plane of the target person, and runs from the back to the chest of the target person. Then, it can be understood that the direction parallel to the assumed axis and from the back to the chest of the target person is the body orientation of the target person. The body orientation of the target person may be detected using a known pose estimation AI or the like. The face orientation of the target person refers to the direction in which the face of the target person is facing on the input image IN. It can be understood that the direction from the back of the head of the target person to the glabella of the target person is the face orientation of the target person.

[0039] The orientation detection unit 133 may detect the body orientation and face orientation of the target person with respect to the road lane in the orientation detection process. When the positional relationship between the target person located outside the road lane and the road lane in the input image IN is known, if the body orientation and face orientation of the target person on the XY coordinate plane are detected, the body orientation and face orientation of the target person with respect to the road lane are determined. By detecting the body orientation of the target person with respect to the road lane, it is detected whether the body of the target person is facing the road lane. By detecting the face orientation of the target person with respect to the road lane, it is detected whether the face of the target person is facing the road lane.

[0040] The region detection unit 134 executes a region detection process based on the input image IN. In the region detection process, the region detection unit 134 detects, for each object existing on the input image IN, the region where the image of the object is located (in other words, the region where the image data of the object is located) in pixel units. Therefore, for example, when the input image IN includes images of a road lane, a sidewalk, and a person, in the region detection process, the road lane region where the image of the road lane is located, the sidewalk region where the image of the sidewalk is located, and the person region where the image of the person is located are detected in pixel units. At this time, if the input image IN includes images of a plurality of persons, a person region is detected for each person. The same applies to the sidewalk and the like. The region detection process may be semantic segmentation, and the region detection unit 134 may be configured by a known AI (semantic segmentation AI) that realizes semantic segmentation.

[0041] The distance detection unit 135 executes a distance detection process based on the detection result of the region detection process. In the distance detection process, the distance detection unit 135 detects the distance between any first and second target objects in the input image IN. The distance detection unit 135 detects the distance between the first target region where the image of the first target object is located and the second target region where the image of the second target object is located in the input image IN as the distance between the first and second target objects. The first and second target regions refer to any two regions detected by the region detection process. The detected distance is expressed in units of the size of pixels in the input image IN.

[0042] For example, when the region detection process is executed on the input image IN1 in FIG. 5, the road lane region where the image of the road lane RD is located and the person region where the image of the person P[i] is located are detected. At this time, the distance detection unit 135 can detect the distance between those road lane region and person region as the distance d[i] between the road lane RD and the person P[i] on the input image IN1. When the road lane region for the road lane RD and the person region for the person P[i] are adjacent to each other, the distance d[i] is zero. "d[i]=0" represents that the person P[i] is located on the road lane RD, in other words, represents that the person P[i] is crossing the road lane RD.

[0043] The probability determination unit 136 performs probability determination processing. For any target person in the input image IN, the probability that the target person crosses the roadway is referred to as the crossing probability. In the probability determination processing, the probability determination unit 136 determines the crossing probability of the person in the input image IN based on the detection results of the detection units 132 to 135 without relying on the input information from the operator OP via the input device 22. However, when it is difficult for the probability determination unit 136 to determine the crossing probability, the probability determination processing is not executed. When the input image IN includes images of a plurality of persons, the probability determination processing can be executed for each person. Note that in the probability determination processing, the crossing probability is calculated, and thus the determination of the crossing probability in the probability determination processing is synonymous with the calculation or derivation of the crossing probability. In the controller 13, the crossing probability calculated by the probability determination unit 136 is added to the input image IN (that is, associated with the input image IN). When the input image IN includes images of a plurality of persons, the crossing probabilities calculated for each person are added to the input image IN for each person.

[0044] The evaluation unit 137 performs accuracy evaluation processing. In the accuracy evaluation processing, the evaluation unit 137 evaluates the accuracy of the determination of the crossing probability by the probability determination processing (in other words, evaluates the accuracy of the crossing probability calculated in the probability determination processing) based on the detection results of the detection units 132 to 135. The accuracy evaluated here indicates the degree of accuracy of the crossing probability calculated in the probability determination processing, and in other words, it can also be said to indicate the difficulty (hardness) of determining the crossing probability. Therefore, it can also be said that the accuracy evaluation processing is a difficulty evaluation processing for evaluating the difficulty of determining the crossing probability by the probability determination processing. Hereinafter, for the sake of simplicity of description, the difficulty of determining the crossing probability by the probability determination processing may be expressed as the determination difficulty. Also, the crossing probability calculated in the probability determination processing is hereinafter referred to as the calculated crossing probability.

[0045] The evaluation unit 137 may classify the difficulty of judgment into three or more levels, but here it is assumed that the difficulty of judgment is classified into two levels. That is, the evaluation unit 137 classifies the difficulty of judgment into low difficulty or high difficulty. The low difficulty is lower than the high difficulty. That is, the difficulty of judgment classified as low difficulty is lower than the difficulty of judgment classified as high difficulty. When the evaluation unit 137 evaluates that the difficulty of judgment is relatively low, it classifies the difficulty of judgment as low difficulty, and when it evaluates that the difficulty of judgment is relatively high, it classifies the difficulty of judgment as high difficulty. When the evaluation unit 137 classifies the difficulty of judgment as low difficulty, it classifies the accuracy of the calculated crossing probability as the first accuracy, and when it classifies the difficulty of judgment as high difficulty, it classifies the accuracy of the calculated crossing probability as the second accuracy. The first accuracy is higher than the second accuracy.

[0046] The probability setting unit 138 executes probability setting processing. In the probability setting processing, the probability setting unit 138 sets the crossing probability of the person in the input image IN. When the input image IN includes images of a plurality of persons, the probability setting processing is executed for each person. Although the details will be described later, when the difficulty of judgment for the target person is classified as low difficulty (therefore, when the calculated crossing probability for the target person is classified as the first accuracy), the crossing probability determined in the probability determination processing is set as the crossing probability of the target person. On the other hand, when the difficulty of judgment for the target person is classified as high difficulty (therefore, when the calculated crossing probability for the target person is classified as the second accuracy), the crossing probability of the target person is set based on the input information from the operator OP via the input device 22.

[0047] The data output unit 139 generates a unit data set for each input image IN. Fig. 6(a) shows the structure of the unit data set. The unit data set for a certain input image IN has the image data of the input image IN, the BBOX information generated by the object detection unit 132, and the crossing probability information. The crossing probability information in the unit data set represents the crossing probability for each person set by the probability setting processing. Hereinafter, the crossing probability set by the probability setting processing (probability setting unit 138) is referred to as the set crossing probability.

[0048] Fig. 6(b) conceptually shows the unit data set generated for the input image IN1. The unit data set for the input image IN1 includes the image data of the input image IN1, the BBOX information of each of BBOX[1] to BBOX[5], and the set crossing probabilities of each of the persons P[1] to P[5]. In the unit data set for the input image IN1, the BBOX information of BBOX[i] is associated with the set crossing probability of the person P[i] (i.e., the BBOX information of the person P[i] is associated with the set crossing probability of the person P[i]).

[0049] One unit data set is learning data for one input image IN. When training the AI model 41 (see Fig. 1), a group of unit data sets for a large number of input images IN is prepared as learning data. The data output unit 139 can store the generated unit data set in the storage unit 12. In this case, the unit data set stored in the storage unit 12 is transferred to the database 30 at an arbitrary timing. The storage unit 12 may be the database 30.

[0050] Hereinafter, in a plurality of embodiments, specific operation examples, applied technologies, modified technologies, etc. of each device shown in Fig. 1 will be described. The matters described above in this embodiment are applied to the following respective embodiments unless otherwise specified and without contradiction. In each embodiment, if there are matters conflicting with the above-described matters, the description in each embodiment may be given priority. Also, without contradiction, matters described in any one of the following plurality of embodiments can be applied to any other embodiment (i.e., it is also possible to combine any two or more of the plurality of embodiments).

[0051] <<First Embodiment>> The first embodiment will be described. In the first embodiment, a method for generating learning data by the learning data generation device 10 will be described. FIG. 7 is a flowchart of the generation operation of the unit data set for one input image IN. The operation shown in FIG. 7 is performed for each input image IN. When a program for generating learning data is started by the controller 13, the process proceeds to step S11. Thereafter, when the unit data set is generated in step S19, the generation operation of FIG. 7 ends.

[0052] In step S11, the input image acquisition unit 131 acquires the input image IN from the above-described input image supply device (not shown). The input image IN acquired here is assumed to include an image of the road RD, an image of the sidewalk SW_L, an image of the sidewalk SW_R, and an image of one or more persons. Further, for example, it is assumed that the input image IN is the input image IN1 of FIG. 5. After step S11, the process proceeds to step S12.

[0053] In step S12, the object detection unit 132 executes an object detection process based on the input image IN to detect the recognition target object in the input image IN, and sets a BBOX for each detected recognition target object. It is assumed that the recognition target object is a person (in other words, a pedestrian) as described above. When the object detection unit 132 sets the BBOX of the recognition target object in the object detection process, it generates BBOX information for specifying the position and shape of the BBOX in the input image IN. In this embodiment, it is assumed that the input image IN includes n images of the recognition target object. Then, in step S12, the first to nth BBOXes for the first to nth recognition target objects are set. n represents an arbitrary integer of 2 or more. However, "n = 1" may be possible. After step S12, the process proceeds to step S13. Incidentally, although not particularly considered in FIG. 7, when "n = 0", the operation of FIG. 7 ends without proceeding to step S13.

[0054] In step S13, the controller 13 assigns 1 to the variable i that it manages, and assigns 0 to the flag data FLG[1] to FLG[n] that it manages. Each flag data is binary data having a value of "0" or "1". The flag data FLG[i] is associated with the i-th recognition target object, and thus is associated with the i-th BBOX. After step S13, the process proceeds to step S14.

[0055] In step S14, the controller 13 executes a unit process for setting the crossing probability (hereinafter, may be simply referred to as a unit process). Although the details of the unit process will be described later, in the unit process, "1" flag data is associated with the recognition target object in which the above-described determination difficulty is classified as high difficulty. After step S14, the process proceeds to step S15.

[0056] In step S15, the controller 13 determines whether "i = n" holds. When "i = n" holds (Y in step S15), a transition to step S17 occurs, and when "i = n" does not hold (N in step S15), a transition to step S16 occurs.

[0057] In step S16, the controller 13 adds 1 to the variable i and then causes a transition to step S14. In this case, the unit process for setting the crossing probability in step S14 is executed again. Prior to the description of the processing contents of steps S17 to S19, the unit process for setting the crossing probability in step S14 will be described in detail.

[0058] FIG. 8 is a flowchart of the unit process for setting the crossing probability according to the first embodiment. In the unit process for setting the crossing probability in FIG. 8, first, the process of step S31 is executed. The unit process for setting the crossing probability is executed when the variable i has any integer value from 1 to n. The i-th BBOX that is the target of the unit process for setting the crossing probability is referred to as the target BBOX. Also, the person corresponding to the target BBOX is referred to as the target person. The target BBOX is a rectangular area in which the target person exists in the input image IN (specifically, a rectangular area in which the image of the target person exists).

[0059] In step S31, the orientation detection unit 133 detects the body orientation and face orientation of the target person by executing the above-described orientation detection process based on the input image IN. It is also possible that only one of the body orientation and the face orientation is detected. After step S31, the process proceeds to step S32.

[0060] In step S32, the controller 13 determines whether there is substantially no possibility that the target person crosses the roadway RD based on the detection result in step S31. At this time, the controller 13 (for example, the orientation detection unit 133) detects whether the body or face of the target person faces the roadway RD in the positional relationship between the position of the target person and the roadway RD with respect to the input image IN, and can determine whether there is substantially no possibility that the target person crosses the roadway RD based on the detection result. Determining that there is substantially no possibility that the target person crosses the roadway RD corresponds to the determination result of step S32 being "positive". When the determination result of step S32 is "positive", a transition to step S33 occurs. Determining that there is a possibility that the target person crosses the roadway RD or that the target person is crossing the roadway RD corresponds to the determination result of step S32 being "negative". When the determination result of step S32 is "negative", a transition to step S34 occurs.

[0061] In the determination of step S32, the controller 13 utilizes the finding that the person located on the sidewalk SW_L is located near the left end of the input image IN and the person located on the sidewalk SW_R is located near the right end of the input image IN. Specifically, the controller 13 sets a left region and a right region within the input image IN. A predetermined region located to the left of the vertical line passing through the center of the input image IN and parallel to the Y-axis is set as the left region, and another predetermined region located to the right of the vertical line is set as the right region. The left region is a region that is outside the road lane RD and is presumed to be located on the left side of the road lane RD. The right region is a region that is outside the road lane RD and is presumed to be located on the right side of the road lane RD. Depending on the width of the road lane RD, etc., the left region does not necessarily completely coincide with the region of the sidewalk SW_L, but in the determination in step S32, the left region is regarded as the region of the sidewalk SW_L. The same applies to the right region.

[0062] Therefore, when the target BBOX is located within the left region, the controller 13 regards the target person as being located on the left side of the road lane RD, and when the target BBOX is located within the right region, the controller 13 regards the target person as being located on the right side of the road lane RD. Specifically, for the target person to be located on the left side of the road lane RD means that the target person is outside the road lane RD and is located on the left side of the road lane RD. Specifically, for the target person to be located on the right side of the road lane RD means that the target person is outside the road lane RD and is located on the right side of the road lane RD.

[0063] The controller 13 quantizes the body orientation and face orientation of the target person detected by the orientation detection unit 133 into four levels of orientation respectively. The four levels of orientation are rightward, leftward, upward, and downward. This quantization may also be performed by the orientation detection unit 133.

[0064] As the determination method in step S32, the first to third determination methods are exemplified. The controller 13 can adopt any of the first to third determination methods. In the first determination method, it is determined whether the body of the target person faces the road RD. In the second determination method, it is determined whether the face of the target person faces the road RD. In the third determination method, both the body and face orientations of the target person are considered.

[0065] The first determination method in step S32 will be described. The controller 13 according to the first determination method has the following determination condition CND 1A and CND 1B to determine the success or failure, and based on the determination result, detect the orientation of the target person's body with respect to the road RD. The controller 13 according to the first determination method, if the determination condition CND 1A or CND 1B is satisfied, it is detected that the body of the target person on the sidewalk does not face the road RD. At this time, it is determined that the possibility of the target person crossing the road RD is almost non-existent. The controller 13 according to the first determination method, if neither the determination condition CND 1A nor CND 1B is satisfied, it is determined that the target person may cross the road RD or is in the process of crossing the road RD. The determination condition CND 1A is satisfied only when the target person is located on the left side of the road RD and the body orientation of the target person is leftward, upward, or downward. The determination condition CND 1B is satisfied only when the target person is located on the right side of the road RD and the body orientation of the target person is rightward, upward, or downward. When the target person has the intention to cross the road RD, the body is often turned towards the road RD side. Therefore, when the determination condition CND 1A or CND 1B is satisfied, it is difficult to consider that the target person is moving towards the road RD.

[0066] The controller 13 according to the second determination method has the following determination condition CND 2A and CND 2B to determine the success or failure, and based on the determination result, detect the orientation of the target person's face with respect to the road RD. The controller 13 according to the second determination method, if the determination condition CND 2A or CND2B If it holds, it is detected that the face of the target person on the sidewalk is not facing the road RD. At this time, it is determined that there is almost no possibility that the target person crosses the road RD. The controller 13 according to the second determination method has the determination condition CND 2A and CND 2B If neither holds, it is determined that there is a possibility that the target person crosses the road RD or is crossing the road RD. The determination condition CND 2A holds only when the target person is located on the left side of the road RD and the face orientation of the target person is left, up, or down. The determination condition CND 2B holds only when the target person is located on the right side of the road RD and the face orientation of the target person is right, up, or down. When the target person has the intention to cross the road RD, the face is often turned towards the road RD side. Therefore, when the determination condition CND 2A or CND 2B holds, it is difficult to consider that the target person is moving forward towards the road RD.

[0067] Explain the third determination method in step S32. The controller 13 according to the third determination method determines the success or failure of the following determination conditions CND 3A and CND 3B and detects the body orientation and face orientation of the target person with respect to the road RD based on the determination result. The controller 13 according to the third determination method has the determination condition CND 3A or CND 3B If it holds, it is detected that the body and face of the target person on the sidewalk are not facing the road RD. At this time, it is determined that there is almost no possibility that the target person crosses the road RD. The controller 13 according to the third determination method has the determination condition CND 3A and CND 3B If neither holds, it is determined that there is a possibility that the target person crosses the road RD or is crossing the road RD. The determination condition CND 3A holds only when the target person is located on the left side of the road RD, the body orientation of the target person is left, up, or down, and the face orientation of the target person is left, up, or down. The determination condition CND 3BIt holds only when the subject is located on the right side of the roadway RD, and the orientation of the subject's body is rightward, upward, or downward, and the orientation of the subject's face is rightward, upward, or downward. By referring to both the body orientation and the face orientation, it is possible to more accurately determine whether the subject has the intention to cross the roadway RD.

[0068] Incidentally, although the method of quantizing the body orientation and the face orientation of the subject detected by the orientation detection unit 133 into four levels each has been described above, each of the body orientation and the face orientation of the subject may be quantized more finely. That is, the body orientation and the face orientation of the subject detected by the orientation detection unit 133 may be quantized into five levels or more each, and based on the quantization results, the controller 13 may determine whether the possibility that the subject crosses the roadway RD is substantially zero. When the body orientation of the subject is quantized into m levels, the body orientation of the subject is quantized as being any one of the first to m-th orientations. The same applies to the face orientation of the subject. Each of the first to m-th orientations belongs to any one of rightward, leftward, upward, and downward. Among the first to m-th orientations, there may be an orientation belonging to both rightward and upward, or an orientation belonging to both leftward and upward. Among the first to m-th orientations, there may be an orientation belonging to both rightward and downward, or an orientation belonging to both leftward and downward.

[0069] The determination in step S32 may be made by the evaluation unit 137, and the determination result of step S32 being "positive" corresponds to the above-mentioned determination difficulty being classified as low difficulty. That is, when the determination result of step S32 is "positive", the difficulty of determining the crossing probability by the probability determination process regarding the subject is relatively low, and thus it is classified as low difficulty. That is, based on the information on the orientation of the subject with respect to the roadway RD, the probability added to the input image IN and the accuracy of the crossing probability of the subject are evaluated.

[0070] In step S33, since the determination difficulty is low, the crossing probability of the target person is calculated and determined to be a sufficiently low probability LL through probability determination processing. Also in step S33, the evaluation unit 137 classifies the accuracy of the calculated crossing probability (probability LL) obtained through probability determination processing as the first accuracy. Further in step S33, the probability setting unit 138 sets the calculated crossing probability (probability LL) obtained through probability determination processing as the crossing probability of the target person. When transitioning to step S33, the calculated crossing probability (probability LL) obtained through probability determination processing is added to the input image IN as the crossing probability of the target person. Probability LL may have a predetermined value. Probability LL may have a constant value (for example, 0% or 5%), or may be expressed within a numerical range such as 5% or less. Probability LL may be variable according to the body orientation or face orientation of the target person. For example, when reaching step S33 when the target BBOX is located within the left region, the probability LL when the body orientation of the target person is upward or downward may be made higher than the probability LL when the body orientation of the target person is leftward. When the setting in step S33 is performed, the unit processing in FIG. 8 ends.

[0071] In step S34, the region detection unit 134 performs the above-described region detection processing. As a result, the lane region and the person region in the input image IN are detected. FIG. 9 shows the result of the region detection processing when the input image IN is the input image IN1 (see FIG. 5). In FIG. 9, the dod region R_RD is the lane region where the image of the lane RD is located. In FIG. 9, the hatched regions R_P[1] to R_P[5] are the person regions where the images of the persons P[1] to P[5] are located, respectively. Actually, the sidewalk region and the vehicle region for the sidewalk and the vehicle in the input image IN1 in FIG. 5 are also detected, but in FIG. 9, only the lane region and the person region are shown explicitly to prevent complication of the illustration. Note that the crosswalk PX in FIG. 9 is understood to belong to the lane region R_RD. After step S34, the process proceeds to step S35.

[0072] In step S35, the distance detection unit 135 performs the above-described distance detection process. As a result, the distance between two objects on the input image IN is detected. The detected distance includes the distance between the road lane RD and the target person. Assume that the target person is the person P[i] in the input image IN1 of FIG. 5. Then, the distance detected in step S35 includes the distance d[i] between the road lane RD and the person P[i] on the input image IN1. As the distance d[i], the distance between the road lane area R_RD for the road lane RD and the person area R_P[i] for the person P[i] is detected. When the road lane area R_RD and the person area R_P[i] are adjacent, the distance d[i] is zero. "d[i]=0" indicates that the person P[i] is located on the road lane RD, in other words, the person P[i] is crossing the road lane RD. After step S35, the process proceeds to step S36.

[0073] In step S36, the controller 13 determines whether the target person (therefore the person P[i]) is crossing the road lane RD based on the distance d[i] detected in step S35. The controller 13 determines that the target person is crossing the road lane RD if the distance d[i] is equal to or less than a predetermined threshold distance, and does not determine that the target person is crossing the road lane RD if the distance d[i] is greater than the threshold distance. The threshold distance may be zero. When the threshold distance is zero, the controller 13 determines that the target person is crossing the road lane RD only when the distance d[i] is zero.

[0074] Determining that the target person is crossing the road lane RD corresponds to the determination result of step S36 being "positive". When the determination result of step S36 is "positive", a transition to step S37 occurs. Not determining that the target person is crossing the road lane RD corresponds to the determination result of step S36 being "negative". When the determination result of step S36 is "negative", a transition to step S38 occurs.

[0075] The determination in step S36 may be made by the evaluation unit 137. The determination result of step S36 being "positive" corresponds to the above-described determination difficulty being classified as low difficulty. That is, when the determination result of step S36 is "positive", the difficulty of determining the crossing probability by the probability determination process for the subject is relatively low, and thus it is classified as low difficulty. That is, based on the information on the position of the subject with respect to the lane RD, the probability added to the input image IN and the accuracy of the crossing probability of the subject are evaluated. As a result, in steps S32 and S36, based on the information on the position and orientation of the subject with respect to the lane RD, the probability added to the input image IN and the accuracy of the crossing probability of the subject are evaluated. This evaluation can be performed based on only the position information or only the orientation information among the information on the position and orientation of the subject with respect to the lane RD.

[0076] In step S37, because the determination difficulty is low difficulty, it is calculated and determined by the probability determination process that the crossing probability of the subject is a sufficiently high probability HH. Also in step S37, the evaluation unit 137 classifies the accuracy of the calculated crossing probability (probability HH) by the probability determination process as the first accuracy. Further in step S37, the probability setting unit 138 sets the calculated crossing probability (probability HH) by the probability determination process as the crossing probability of the subject. When shifting to step S37, the calculated crossing probability (probability HH) by the probability determination process is added to the input image IN as the crossing probability of the subject. The probability HH may have a predetermined value. The probability HH may be a constant value (for example, 100% or 95%), or may be expressed in a numerical range such as 95% or more. The probability HH may be variably set according to the distance d[i]. In this case, the smaller the distance d[i], the higher the probability HH. In any case, the probability HH is higher than the above probability LL. When the setting in step S37 is performed, the unit process of FIG. 8 ends.

[0077] The fact that the determination result in step S36 is "negative" corresponds to the above-described high difficulty of determination being classified as high difficulty. That is, when the determination result in step S36 is "negative", the difficulty of determining the crossing probability by probability determination processing for the target person is relatively high, and thus it is classified as high difficulty. When the transition to step S38 occurs, in the probability determination processing executed in step S38, the crossing probability of the target person is calculated according to an algorithm preset based on the orientation of the target person's body and face with respect to the road lane RD and the distance d[i]. At this time, the calculated crossing probability is classified as the second accuracy level in step S38. A required value is set in advance for the accuracy level evaluated by the evaluation unit 137. The first accuracy level satisfies the required value, while the second accuracy level does not satisfy the required value. That the first accuracy level satisfies the required value means that the first accuracy level is greater than or equal to the required value, and that the second accuracy level does not satisfy the required value means that the second accuracy level is lower than the required value. Note that when the transition to step S38 occurs, the calculation of the crossing probability by probability determination processing itself may not be executed. When it cannot be determined that the target person has almost no possibility of crossing the road lane RD (N in step S32), and when it cannot be determined that the target person is crossing the road lane RD (N in step S36), the transition to step S38 occurs. When the transition to step S38 occurs, it is considered that the crossing probability of the target person varies depending on various factors, and it is difficult for the controller 13 to accurately estimate the crossing probability.

[0078] Therefore, in step S38, the controller 13 sets 1 in the flag data FLG[i] in order to have the operator OP specify the crossing probability. When the setting in step S38 is performed, the unit process in FIG. 8 ends.

[0079] The steps S17 to S19 in FIG. 7 will be described. When reaching step S17, the unit process (unit process for setting the crossing probability) of step S14 for the first to nth BBOXes has been executed.

[0080] In step S17, the controller 13 checks whether "1" is set in any of the flag data FLG[1] to FLG[n]. If "1" is set in one or more of the flag data FLG[1] to FLG[n] (Y in step S17), the process proceeds to step S18. If the values of the flag data FLG[1] to FLG[n] are all "0" (N in step S17), the process proceeds to step S19.

[0081] In the case CS_A1 where, in the n unit processes for the first to nth BBOXes, there is no transition to step S38 (see FIG. 8) even once, the process proceeds to step S19 without proceeding to step S18. In case CS_A1, the crossing probabilities are set in step S33 or S37 for each person in the first to nth BBOXes. That is, in case CS_A1, all the crossing probabilities are set by the learning data generation device 10 alone without requiring the input operation of the operator OP.

[0082] In step S19, the unit data set is generated by the data output unit 139. The structure of the unit data set is as described above with reference to FIG. 6(a), and the crossing probability information in the unit data set represents the crossing probabilities of the first to nth persons corresponding to the first to nth BBOXes for each person. In case CS_A1, all the crossing probabilities in the crossing probability information are set in step S33 or S37.

[0083] In the n unit processes for the first to nth BBOXes, the case where there is a transition to step S38 (see FIG. 8) one or more times is referred to as case CS_A2. In case CS_A2, the process proceeds from step S17 to step S18.

[0084] In step S18, the probability setting unit 138 executes an inquiry process described later, and based on the response input information obtained from the operator OP in the inquiry process, sets the crossing probability of the person corresponding to the flag data of "1". That is, the probability setting unit 138 according to step S18 replaces the crossing probability of the person corresponding to the flag data of "1" with the crossing probability based on the response input information from the calculated crossing probability in step S38. In other words, the probability setting unit 138 according to step S18 replaces the crossing probability of the person added to the input image IN (the crossing probability of the person corresponding to the flag data of "1") with the crossing probability based on the response input information from the calculated crossing probability by the probability determination process. When the crossing probability of the person corresponding to the flag data of "1" has not been calculated in the probability determination process, the probability setting unit 138 according to step S18 may simply set the crossing probability of the person corresponding to the flag data of "1" to the crossing probability based on the response input information. For example, consider case CS_A2a belonging to case CS_A2. Fig. 10 shows an example of an image displayed on the display device 21 in the inquiry process according to case CS_A2a.

[0085] The input image IN in case CS_A2a is the input image IN1 in Fig. 5. In the input image IN1, the persons P[1] and P[3] are located on the sidewalk SW_L, and the persons P[2] and P[5] are located on the sidewalk SW_R. However, the person P[5] is located at the boundary between the sidewalk SW_R and the road RD. In the input image IN1, the person P[4] is located on the crosswalk PX within the road RD.

[0086] In case CS_A2a, in the unit process for the person P[1], it is determined that the possibility of the person P[1] crossing the road RD is almost zero, and as a result, the crossing probability of the person P[1] is set to the probability LL (0% in the example of Fig. 10) in step S33. In case CS_A2a, in the unit process for the person P[2], it is determined that the possibility of the person P[2] crossing the road RD is almost zero, and as a result, the crossing probability of the person P[2] is set to the probability LL (0% in the example of Fig. 10) in step S33.

[0087] In case CS_A2a, in the unit process for person P[4], it is determined that person P[4] is crossing the road RD. As a result, at step S37, the crossing probability of person P[4] is set to probability HH (100% in the example of FIG. 10). In case CS_A2a, in the unit process for person P[5], it is determined that person P[5] is crossing the road RD. As a result, at step S37, the crossing probability of person P[5] is set to probability HH (95% in the example of FIG. 10). In the example of FIG. 10, although the distance between person P[5] and the road RD is short but not zero, the crossing probability set for person P[5] is lower than that of person P[4].

[0088] In case CS_A2a, the transition to step S38 occurs only when "i = 3". As a result, among the flag data FLG[1] to FLG[n], only the flag data FLG[3] has "1". In the input image IN1, although person P[3] is facing the road RD, since the distance between person P[3] and the road RD is greater than the threshold distance, the transition to step S38 occurs in the unit process targeting person P[3].

[0089] In step S18 according to case CS_A2a, the probability setting unit 138 causes the input image IN1 to be displayed on the display device 21 together with the BBOX setting information, the probability setting information, and the inquiry information Q in an inquiry process. The following display refers to the display on the display device 21 unless otherwise specified.

[0090] The BBOX setting information represents the BBOX[1] to BBOX[5] set in the object detection process. The outer shapes of the BBOX[1] to BBOX[5] set in the object detection process are superimposed and displayed on the input image IN1.

[0091] The probability setting information represents the crossing probability set in step S33 or S37. The crossing probability set in step S33 or S37 is superimposed on or adjacent to the input image IN1 for display. At this time, the probability setting information is displayed using so-called speech bubbles or the like so that the operator OP can easily recognize that the crossing probability set for the person P[i] and the BBOX[i] are in a corresponding relationship. In the example of FIG. 10, the crossing probabilities of the persons P[1], P[2], P[4], and P[5] (here, 0%, 0%, 100%, 95%) are displayed as the probability setting information.

[0092] The inquiry information Q is information for inquiring the operator OP about what the crossing probability of the person P[3] is. The inquiry information Q is displayed using so-called speech bubbles or the like so that the operator OP can easily recognize that the crossing probability to be inquired about is the crossing probability of the person P[3]. The operator OP who has visually recognized the inquiry information Q determines the crossing probability of the person P[3] after checking the content of the input image IN1, and inputs answer input information specifying the crossing probability of the person P[3] to the input device 22. The answer input information is transmitted to the controller 13 through the input device 22.

[0093] In step S18 according to case CS_A2a, the probability setting unit 138 sets the crossing probability specified by the answer input information as the crossing probability of person P[3]. That is, in step S18 according to case CS_A2a, the probability setting unit 138 replaces the crossing probability of person P[3] added to the input image IN1 with the crossing probability specified by the answer input information from the calculated crossing probability by the probability determination process. When the crossing probability of the person P[3] has not been calculated by the probability determination process, the probability setting unit 138 according to step S18 in case CS_A2a may simply set the crossing probability based on the answer input information as the crossing probability of person P[3] added to the input image IN1. After step S18, the process proceeds to step S19. The crossing probability information in the unit dataset generated in step S19 represents the crossing probabilities of the first to nth persons corresponding to the first to nth BBOXes for each person. In case CS_A2a, among the crossing probability information, the crossing probabilities of persons P[1], P[2], P[4], and P[5] are set in step S33 or S37, while the crossing probability of person P[3] is set based on the answer input information.

[0094] In addition, in the inquiry process, the probability setting unit 138 makes the display modes of the BBOXes (such as BBOX[1]) of the persons whose crossing probabilities are set in step S33 or S37 different from the display mode of the BBOX (BBOX[3]) of the person related to the inquiry. For example, the former BBOX is displayed in blue, and the latter BBOX is displayed in red. Thereby, the operator OP can easily distinguish the former BBOX from the latter BBOX.

[0095] Also, when there are two or more persons corresponding to the flag data of "1", the crossing probability is inquired of the operator OP for each person corresponding to the flag data of "1". In this case, answer input information is obtained for each person corresponding to the flag data of "1" and the crossing probability is set.

[0096] At the stage where the inquiry process is performed, the operator OP may be able to modify the crossing probability set in step S33 or S37. For example, in case CS_A2a, consider a situation where the crossing probability of person P[1] is set to 0% in step S33, but the operator OP determines that the crossing probability of person P[1] is 50%. In this case, the operator OP inputs to the input device 22 a modification operation indicating that the crossing probability of object P[1] is 50%. Then, in response to the modification operation, the probability setting unit 138 changes the crossing probability of person P[1] from 0% to 50%.

[0097] The area detection process only needs to be executed once for one input image IN. Therefore, even if the determination result in step S32 is "negative" multiple times among the n - time unit processes for one input image IN, the area detection process in step S34 only needs to be executed once. The area detection process in step S34 may be executed, for example, between steps S12 and S13.

[0098] The inquiry process may be performed within the unit process. That is, for example, the process of step S18 in FIG. 7 may be performed within step S38 in FIG. 8. However, in this case, when there are multiple persons whose difficulty in determining the crossing probability is classified as high difficulty, the inquiry process will be performed for each person. Therefore, in order to improve the working efficiency of the operator OP, as shown in FIG. 7, it is preferable to perform the inquiry process collectively after all the unit processes for the first to n - th BBOXes (the first to n - th persons) are completed.

[0099] Here, the crossing probability is defined in units of %, but as long as the crossing probability is set in multiple stages, the method of defining the crossing probability is arbitrary. For example, the crossing probability of each person may be set in three stages: "large", "medium", and "small".

[0100] In the flowchart of FIG. 7, when the condition for transitioning from step S17 to step S18 is satisfied for a certain input image IN of interest, a modification may be made to not execute the processes of steps S18 and S19. In this case, the unit data set for the input image IN of interest is not generated, and thus the data related to the input image IN of interest is excluded from the learning data. The above modification can be adopted particularly, for example, when a sufficient amount of unit data sets (a sufficient amount of unit data sets with high accuracy) has been acquired.

[0101] The controller 13 may calculate the accuracy of the crossing probability of the target person based on the orientation of the target person's body with respect to the road lane RD, the orientation of the target person's face with respect to the road lane RD, and the distance between the target person and the road lane RD at step S32 or before step S32. Then, when the calculated accuracy is higher than the threshold value, the determination result of step S32 may be treated as "positive", and when the calculated accuracy is lower than the threshold value, the determination result of step S32 may be treated as "negative". Now, consider the case where the target person is person P[i]. Then, the distance between the target person and the road lane RD is distance d[i]. Let the orientation of person P[i]'s body with respect to the road lane RD be represented by angle θ, and the orientation of person P[i]'s face with respect to the road lane RD be represented by angle φ. Then, for example, the controller 13 may calculate the sum of the first term (k1 / d[i]), the second term (k2×sinθ), and the third term (k3×sinφ) as the accuracy of the crossing probability of the target person. k1, k2, and k3 are constants. Assume that angle θ is 0° when the target person's body is exactly facing the road lane RD, and angle φ is 0° when the target person's face is exactly facing the road lane RD.

[0102] <<Second Embodiment>> The second embodiment will be described. The second embodiment and the third to sixth embodiments described later are embodiments based on the first embodiment. Regarding matters not particularly described in the second to sixth embodiments, the description of the first embodiment is applicable to the second to sixth embodiments as long as there is no contradiction. However, when interpreting the description of the second embodiment, the description of the second embodiment may be given priority for matters that conflict between the first and second embodiments (the same applies to the third to sixth embodiments described later). As long as there is no contradiction, any plurality of the first to sixth embodiments may be combined.

[0103] In the first embodiment, the execution order of each process constituting the operations in FIGS. 7 and 8 can be changed in various ways. For example, in the first embodiment, a modified procedure in which the region detection process and the distance detection process are executed between steps S11 and S13 may be adopted. An explanation will be added to this modified procedure.

[0104] FIG. 11 is a flowchart of the generation operation of a unit data set in which the modified procedure is adopted, and is a flowchart of the generation operation of a unit data set according to the second embodiment. The flowchart of FIG. 11 is a modification of a part of the flowchart of FIG. 7. In the flowchart of FIG. 11, the description of the same parts as the flowchart of FIG. 7 is omitted, and the differences from the flowchart of FIG. 7 will be described.

[0105] First, the processing contents of steps S11 and S12 are as shown in the first embodiment. In the flowchart of FIG. 11, after step S12, the processes of steps S34 and S35 are executed and then the process proceeds to step S13. The processing contents of steps S34 and S35 are as shown in the first embodiment. However, in step S35 according to the second embodiment, it is assumed that the distances between the road lane and each of all the persons on the input image IN are detected. For the sake of specific description, it is assumed that the input image IN is the input image IN1 in FIG. 5. Then, in step S35 according to the second embodiment, the distances between the road lane RD and each of the persons P[1] to P[5] on the input image IN1 are detected. The distance between the road lane RD and the person P[i] on the input image IN1 is denoted as the distance d[i] as described above. As the distance d[i], the distance between the road lane region R_RD for the road lane RD and the person region R_P[i] for the person P[i] is detected.

[0106] In the flowchart of FIG. 11, after step S35, the process proceeds to step S13. The flow of operations after proceeding to step S13 is the same as that in the first embodiment. Therefore, after step S13, after the unit process of step S14 is executed n times, the process reaches step S19 through step S17 or reaches step S19 through steps S17 and S18. However, in the second embodiment, since the processes of steps S34 and S35 are executed before the execution of the unit process of step S14, there is no need to execute the processes of steps S34 and S35 during the unit process of step S14.

[0107] FIG. 12 shows a flowchart of a unit process for setting the crossing probability according to the second embodiment. The unit process of FIG. 12 can be executed in step S14 of FIG. 11. The unit process for setting the crossing probability is executed when the variable i has any integer value from 1 to n. As described in the first embodiment, the i-th BBOX targeted by the unit process for setting the crossing probability is referred to as the target BBOX, and the person corresponding to the target BBOX is referred to as the target person.

[0108] When the unit process of FIG. 12 starts, first, the process of step S36 is executed. The content of the process of step S36 is as shown in the first embodiment. That is, in step S36, the controller 13 determines whether the target person (therefore the person P[i]) is crossing the roadway RD based on the distance d[i] detected in step S35.

[0109] Determining that the target person is crossing the roadway RD corresponds to the determination result of step S36 being "positive". When the determination result of step S36 is "positive", a transition to step S37 occurs. Determining that the target person is not crossing the roadway RD corresponds to the determination result of step S36 being "negative". When the determination result of step S36 is "negative", in the second embodiment, a transition to step S31 occurs.

[0110] The determination in step S36 may be performed by the evaluation unit 137, and the determination result of step S36 being "positive" corresponds to the above-mentioned determination difficulty being classified as low difficulty. That is, when the determination result of step S36 is "positive", the difficulty of determining the crossing probability of the target person by the probability determination process is relatively low, and thus it is classified as low difficulty.

[0111] The content of the process of step S37 executed when the determination result of step S36 is "positive" is as shown in the first embodiment. That is, in step S37, it is determined by the probability determination process that the crossing probability of the target person is a sufficiently high probability HH, and the probability setting unit 138 sets the probability HH for the crossing probability of the target person according to the determination result of the probability determination process. When the setting of step S37 is performed, the unit process of FIG. 12 ends.

[0112] The content of the process of step S31 is as shown in the first embodiment. That is, in step S31, the orientation detection unit 133 detects the body orientation and face orientation of the target person by executing the above-mentioned orientation detection process based on the input image IN. It is also possible that only one of the body orientation and face orientation is detected. After step S31, the process proceeds to step S32.

[0113] In step S32, similar to the first embodiment, based on the detection result in step S31, it is determined whether there is almost no possibility that the target person crosses the roadway RD. Determining that there is almost no possibility that the target person crosses the roadway RD corresponds to the determination result of step S32 being "positive". When the determination result of step S32 is "positive", a transition to step S33 occurs. Determining that there is a possibility that the target person crosses the roadway RD corresponds to the determination result of step S32 being "negative". When the determination result of step S32 is "negative", in the second embodiment, a transition to step S38 occurs.

[0114] As the determination method in step S32, the first, second, or third determination method shown in the first embodiment may be adopted. However, in the second embodiment, the positional relationship between the roadway and each person has already been recognized by the controller 13 through the area detection process in step S34 of FIG. 11. Therefore, by using the result of the area detection process, the positional relationship between the roadway RD and the target person can be accurately grasped.

[0115] Taking the case where the input image IN is the input image IN1 in FIG. 5 as an example, an explanation is added. The controller 13 related to step S32 can specify whether the person P[i] who is the target person is located on the left side or the right side of the roadway RD on the input image IN1 based on the result of the area detection process. If it can be specified whether the person P[i] is located on the left side or the right side of the roadway RD, based on the body orientation or face orientation of the person P[i] (target person), it can be determined whether there is almost no possibility that the target person crosses the roadway RD by the first, second, or third determination method.

[0116] That the person P[i] is located on the left side of the road RD specifically means that the person P[i] is outside the road RD and located on the left side of the road RD. That the person P[i] is located on the right side of the road RD specifically means that the person P[i] is outside the road RD and located on the right side of the road RD. Referring to FIG. 9, on the input image IN1, if the person area R_P[i] does not overlap with the road area R_RD and is located on the left side of the road area R_RD, the person P[i] is located on the left side of the road RD. On the input image IN1, if the person area R_P[i] does not overlap with the road area R_RD and is located on the right side of the road area R_RD, the person P[i] is located on the right side of the road RD. Incidentally, when the person P[i] is located on the road RD, since the determination result of step S36 is "positive", the process does not proceed to step S32.

[0117] The determination in step S32 may be made by the evaluation unit 137, and the determination result of step S32 being "positive" corresponds to the above-mentioned determination difficulty being classified as low difficulty. That is, when the determination result of step S32 is "positive", the difficulty of determining the crossing probability of the subject by the probability determination process is relatively low, and thus it is classified as low difficulty.

[0118] The processing content of step S33 executed when the determination result of step S32 is "positive" is as shown in the first embodiment. That is, in step S33, it is determined that the crossing probability of the subject is a sufficiently low probability LL by the probability determination process, and the probability setting unit 138 sets the probability LL for the crossing probability of the subject according to the determination result of the probability determination process. When the setting of step S33 is performed, the unit process of FIG. 12 ends.

[0119] In the second embodiment, the determination result of step S32 being "negative" corresponds to the above-described determination difficulty being classified as high difficulty. That is, when the determination result of step S32 is "negative", the difficulty of determining the crossing probability by probability determination processing for the subject is relatively high, and thus it is classified as high difficulty. When it cannot be determined that the subject is crossing the road RD (N in step S36), and it also cannot be determined that there is almost no possibility that the subject will cross the road RD (N in step S32), a transition to step S38 occurs. The processing content of step S38 is as shown in the first embodiment. When a transition to step S38 occurs, it is considered that the crossing probability of the subject fluctuates depending on various factors, and it is difficult for the controller 13 to accurately estimate the crossing probability. For this reason, in step S38, the controller 13 sets 1 in the flag data FLG[i] in order to have the operator OP specify the crossing probability. When the setting in step S38 is performed, the unit processing in FIG. 12 ends.

[0120] <<Third Embodiment>> The third embodiment will be described. The recognition target object in the object detection process may include an additional object that is an object other than a person while including a person. The additional object is an object that moves along with the person. There are a first type of additional object and a second type of additional object as the additional objects. The first type of additional object is an object for a person to move on. Examples of the first type of additional object are a bicycle, a kick scooter, and an electric cart. The second type of additional object is an object carried by a person. Examples of the second type of additional object are a trolley, a baby stroller, and a wheelbarrow.

[0121] When the subject is on the first type of additional object, the controller 13 may perform some processing in consideration of the first type of additional object moving along with the subject. When the subject is carrying the second type of additional object, the controller 13 may perform some processing in consideration of the second type of additional object moving along with the subject.

[0122] Figure 13 shows an input image IN2 which is an example of the input image IN. Similar to the input image IN1 in Figure 5, the input image IN2 in Figure 13 includes an image of the road lane RD, an image of the sidewalk SW_L located on the left side of the road lane RD, and an image of the sidewalk SW_R located on the right side of the road lane RD. Also, a crosswalk PX is formed within the road lane RD in the input image IN2. Furthermore, the input image IN2 includes an image of a person P[1] and an image of a bicycle BYC. In the input image IN2, the person P[1] is located on the sidewalk SW_L while riding the bicycle BYC.

[0123] The object detection unit 132 sets a BBOX for each object to be recognized in the object detection process for the input image IN2. Therefore, for the input image IN2, BBOX[1] is set as the BBOX for the person P[1], and BBOX BYC is set as the BBOX for the bicycle BYC. BBOX BYC is a rectangular area where it is determined that the bicycle BYC exists within the input image IN2. In the input image IN2, the BBOX[1] of the person P[1] and the BBOX of the bicycle BYC BYC partially overlap. Based on the fact that the BBOX[1] of the person P[1] and the BBOX of the bicycle BYC BYC partially overlap, the controller 13 determines that the person P[1] is riding the bicycle BYC.

[0124] When it is determined that the person P[1] is riding the bicycle BYC, the controller 13 may perform the process of step S36 (see Figure 8 or Figure 12) considering that the bicycle BYC moves along with the person P[1]. That is, for example, when it is determined that the person P[1] is riding the bicycle BYC, the controller 13, in the input image IN2, BBOX[1] and BBOX BYCThe distance between the synthesis region and the lane region of the road lane RD may be determined as the distance d[i]. Then, in step S36, the controller 13 determines that the subject is crossing the road lane RD if the distance d[i] is less than or equal to a predetermined threshold distance, and does not determine that the subject is crossing the road lane RD if the distance d[i] is greater than the threshold distance. The threshold distance may be zero. The lane region of the road lane RD is the region detected by the region detection process and in which the image of the road lane RD is located, and corresponds to the lane region R_RD in FIG. 9.

[0125] The same applies when the person P[1] is riding on a first type of additional object other than a bicycle. The same also applies when the person P[1] is carrying a second type of additional object.

[0126] It can be said that the controller 13 determines and sets the crossing probability of the object to the road lane based on the input image IN including the image of the object and the image of the road lane. In other words, the input image IN including the image of the object and the image of the road lane is the input image IN in which the object and the road lane (road) are reflected. The object includes at least a person. In the first and second embodiments, the object is the person itself. In contrast, in the third embodiment, it can be said that the object is composed of a person and a first or second type of additional object associated with the person. Even when the object is composed of a person and an additional object (first or second additional object), the determination method in step S32 is the same as that shown in the first or second embodiment, and the presence or absence of the additional object does not affect the determination result in step S32.

[0127] Here, focusing on one object, the operation of the learning data generation device 10 will be summarized. The controller 13 (probability determination unit 136) is capable of executing a probability determination process for calculating the crossing probability of the object on the roadway based on the input image IN having an image of the object including a person and an image of the roadway. The controller 13 (evaluation unit 137) classifies the accuracy of the crossing probability calculated in the probability determination process into a first accuracy or a second accuracy based on the positional relationship between the object and the roadway in the input image IN (see, for example, S32 and S36 in FIG. 8). This corresponds to classifying the difficulty of determining the crossing probability by the probability determination process into low difficulty or high difficulty. Then, when the controller 13 classifies the accuracy of the crossing probability calculated in the probability determination process into the first accuracy, it sets the crossing probability of the object based on the calculation result of the probability determination process (for example, S33 or S37 in FIG. 8). On the other hand, when the controller 13 classifies the accuracy of the crossing probability calculated in the probability determination process into the second accuracy, it sets the crossing probability of the object based on the information input from the operator OP through the input device 22 (for example, S18 in FIG. 7 via S38 in FIG. 8). After that, the controller 13 generates learning data including the input image IN and including the position and shape of the object on the input image IN and the set crossing probability as the correct label. This corresponds to generating a unit data set including the input image IN and including the BBOX information and the probability setting information as the correct label in step S19.

[0128] In the first reference method in which the operator OP designates all the crossing probabilities of the person, the manual work becomes enormous and learning data cannot be efficiently generated. In the second reference method in which all the crossing probabilities of the person are performed by a machine without requiring manual work, a large amount of learning data can be generated in a short time. However, it is often difficult for the machine to correctly judge the crossing probability, and as a result, it is difficult to guarantee the quality of the learning data in the second reference method. On the other hand, according to the learning data generation device 10, the accuracy of the calculated crossing probability is evaluated, and the operator OP is required to input only for those with low accuracy. Therefore, the work load of the operator OP is reduced compared with the first reference method, and the quality of the learning data is easier to guarantee compared with the second reference method. That is, according to the learning data generation device 10, it is possible to efficiently generate a large amount of high-quality learning data.

[0129] The controller 13 can detect at least one of the body orientation of the person with respect to the lane and the face orientation of the person in the input image IN, and classify the accuracy of the calculated crossing probability by the probability judgment process into the first accuracy or the second accuracy based on the detection result (see, for example, S31 and S32). By detecting the orientation of the body or face of the person with respect to the lane, it is possible to estimate the possibility of the person crossing the lane. And if there seems to be no possibility, it can be classified as low difficulty, the crossing probability can be set sufficiently low, and the accuracy can be classified as the first accuracy, or if there seems to be a possibility, it can be classified as high difficulty and the accuracy can be classified as the second accuracy. Through appropriate classification, it is possible to efficiently generate a large amount of high-quality learning data.

[0130] The controller 13 can classify the accuracy of the calculated crossing probability by the probability judgment process into the first accuracy or the second accuracy based on the distance between the object and the lane in the input image IN (see, for example, S36). From the distance between the object and the lane, it is possible to easily estimate whether the object is crossing the lane. And if it is crossing, it can be classified as low difficulty, the crossing probability can be set sufficiently high, and the accuracy can be classified as the first accuracy, or if it is not crossing, it can be classified as high difficulty and the accuracy can be classified as the second accuracy. Through appropriate classification, it is possible to efficiently generate a large amount of high-quality learning data.

[0131] The controller 13 may execute both the first detection process and the second detection process. Here, the first detection process is an orientation detection process for detecting at least one of the body orientation of a person with respect to the lane and the face orientation of the person in the input image IN. The second detection process is a distance detection process for detecting the distance between the object and the lane in the input image IN. Then, the controller 13 can classify the accuracy of the calculated crossing probability by the probability determination process based on the detection results of the first detection process and the second detection process into the first accuracy or the second accuracy (see, for example, S31 to S36 in FIG. 8). By detecting the orientation of the body or face of a person with respect to the lane, it is possible to infer the possibility of the person crossing the lane. And if there seems to be no possibility, it can be classified as low difficulty, the crossing probability can be set sufficiently low, and the accuracy can be classified as the first accuracy, or if there seems to be a possibility, it can be classified as high difficulty and the accuracy can be classified as the second accuracy. Also, it is possible to easily infer whether the object is crossing the lane from the distance between the object and the lane. And if it is crossing, it can be classified as low difficulty, the crossing probability can be set sufficiently high, and the accuracy can be classified as the first accuracy, or if it is not crossing, it can be classified as high difficulty and the accuracy can be classified as the second accuracy. Through appropriate classification, it becomes possible to efficiently generate a large amount of high-quality learning data.

[0132] <<Fourth Embodiment>> The fourth embodiment will be described. In the fourth embodiment, the configuration and operation of the learning device 40 (see FIG. 1) will be described. The learning device 40 may be configured by any one or more server devices. An AI model 41 composed of a neural network is provided in the learning device 40. The learning device 40 includes an arithmetic processing unit including a CPU (Central Processing Unit) and a GPU (Graphics Processing Unit) as hardware resources, and learns the AI model 41 in the arithmetic processing unit.

[0133] The learning device 40 provides the learning images (specifically, the image data of the learning images) in the training data as problem data to the AI model 41, and provides the BBOX information and the crossing probability information in the training data as the correct labels. Since the input image IN is included in the training data as a learning image, each learning image includes an image of a road lane and an image of a person. The AI model 41 can execute the same type of object detection process as the object detection unit 132 (see FIG. 4), detect a person in the learning image, and estimate (infer) the BBOX of the detected person. In addition, the AI model 41 estimates (infers) the crossing probability of each person in the learning image to the road lane.

[0134] The learning device 40 evaluates the error between the BBOX and the crossing probability estimated by the AI model 41 and the BBOX information and the crossing probability information given as the correct labels. Then, the learning device 40 executes the learning of the AI model 41 using a learning algorithm such as the error backpropagation method so that the error is reduced. In this learning, the parameters (weights, etc.) of the AI model 41 are adjusted. When a predetermined learning end condition is satisfied, such as when the above error converges to a sufficiently small value, the learning ends. The AI model 41 after the learning is completed is referred to as a learned model.

[0135] The learned model can be applied to an operation vehicle V2 which is an arbitrary vehicle. FIG. 14 shows a top view of the operation vehicle V2. The operation vehicle V2 may be the same type of vehicle as the cooperative vehicle V1, such as an automobile traveling on a road lane. A camera C2 for photographing the front area of the operation vehicle V2 is installed in the operation vehicle V2. The camera C2 may be a camera of a drive recorder installed in the operation vehicle V2. An AI application device 50 is mounted on the operation vehicle V2. The AI application device 50 may be composed of an ECU (Electronic Control Unit) or the like provided in the operation vehicle V2.

[0136] The learned model is provided in the AI application device 50, and the image data of the captured image of the camera C2 is input to the learned model. Alternatively, the learned model may be installed in an arbitrary operation server device (not shown) provided outside the operation vehicle V2. In this case, the image data of the captured image of the camera C2 is input to the learned model in the operation server device through the AI application device 50. The learned model detects a person in the captured image based on the input image data of the captured image, estimates the BBOX of the detected person, and estimates the probability of the person crossing the road.

[0137] The AI application device 50 can perform various processes using the estimation results of the learned model. For example, the AI application device 50 can detect obstacles around the operation vehicle V2 using a radar or the like installed in the camera C2 or the operation vehicle V2, and detect pedestrians or the like during obstacle detection. The AI application device 50 can give a warning to the driver of the operation vehicle V2, control the running of the operation vehicle V2, or give a warning notification to the surroundings according to the probability of a pedestrian or the like crossing the road. The running control of the operation vehicle V2 includes speed control of the operation vehicle V2. For example, when a pedestrian or the like with a sufficiently high crossing probability is detected, it is possible to reduce the speed of the operation vehicle V2. The warning notification to the surroundings means notifying pedestrians or the like around the operation vehicle V2 that the operation vehicle V2 is approaching.

[0138] The operation vehicle V2 may be a commercial vehicle. Examples of commercial vehicles include taxis and work vehicles in factories. It is also possible to use the estimation results of the learned model for the driving evaluation of commercial vehicles. For example, when there is a pedestrian with a high probability of crossing the road in front of a taxi, the AI application device 50 can detect whether an appropriate speed reduction or braking operation is performed in the taxi, and use the detection result for the driving evaluation of the taxi. Also, for example, based on the crossing probability of pedestrians around the taxi when the taxi applies sudden braking, it is possible to evaluate whether the sudden braking operation was appropriate.

[0139] <<Example 5>> A fifth embodiment will be described. As shown in FIG. 15, the learning data generation device 10 includes a crossing probability determination device 10P. It can be considered that the crossing probability determination device 10P is formed by deleting any partial function of the learning data generation device 10 from all the functions of the learning data generation device 10. For example, the crossing probability determination device 10P has the same hardware configuration as the learning data generation device 10, but does not have the function of performing the above-described inquiry process.

[0140] The crossing probability determination device 10P performs up to the evaluation of the determination difficulty by the probability determination process, associates the evaluation result with the input image IN, and stores it in the storage unit 12 or the database 30 together with the input image IN. However, the controller 13 in the crossing probability determination device 10P may set the crossing probability of the object based on the determination result of the probability determination process when classifying the determination difficulty by the probability determination process as low difficulty (when classifying the accuracy of the calculated crossing probability by the probability determination process as the first accuracy) in the unit process. The controller 13 in the crossing probability determination device 10P may add specific data to the input image IN in association with the object when classifying the determination difficulty by the probability determination process as high difficulty (when classifying the accuracy of the calculated crossing probability by the probability determination process as the second accuracy) in the unit process.

[0141] The controller 13 of the crossing probability determination device 10P can perform an operation in which only the operation of step S18 is excluded from the operations shown in FIGS. 7 and 8. Alternatively, the controller 13 of the crossing probability determination device 10P can perform an operation in which only the operation of step S18 is excluded from the operations shown in FIGS. 11 and 12.

[0142] FIG. 16 shows an example of the operation flowchart of the controller 13 in the crossing probability determination device 10P. The operation flowchart of FIG. 16 is a modification of a part of the operation flowchart of FIG. 7, and the operation flowchart of FIG. 16 can be obtained by replacing step S18 in FIG. 7 with step S18a. FIG. 17 shows another example of the operation flowchart of the controller 13 in the crossing probability determination device 10P. The operation flowchart of FIG. 17 is a modification of a part of the operation flowchart of FIG. 11, and the operation flowchart of FIG. 17 can be obtained by replacing step S18 in FIG. 11 with step S18a. The controller 13 shown below in the fifth embodiment refers to the controller 13 of the crossing probability determination device 10P. Assuming cases CS_E1 and CS_E2, the operation examples of the crossing probability determination device 10P will be described.

[0143] In case CS_E1, in the n unit processes for the first to nth BBOXes, there is no transition to step S38 (see FIG. 8 or FIG. 12) even once. In case CS_E1, the crossing probability is set for each person in the first to nth BBOXes at step S33 or S37. In case CS_E1, since all of the flag data FLG[1] to FLG[n] at the stage when step S17 is reached are "0", the process proceeds to step S19. The operation of step S19 in this case is the same as that in the learning data generation device 10. The controller 13 can store the unit data set generated in step S19 related to case CS_E1 in the storage unit 12 or the database 30.

[0144] In case CS_E2, assume that the input image IN is the input image IN1 in FIG. 5. In case CS_E2, in the unit process for person P[1], it is determined that there is almost no possibility that person P[1] crosses the road RD. As a result, the crossing probability of person P[1] is set to probability LL in step S33. In case CS_E2, in the unit process for person P[2], it is determined that there is almost no possibility that person P[2] crosses the road RD. As a result, the crossing probability of person P[2] is set to probability LL in step S33. In case CS_E2, in the unit process for person P[4], it is determined that person P[4] is crossing the road RD. As a result, the crossing probability of person P[4] is set to probability HH in step S37. In case CS_E2, in the unit process for person P[5], it is determined that person P[5] is crossing the road RD. As a result, the crossing probability of person P[5] is set to probability HH in step S37.

[0145] In case CS_E2, the transition to step S38 occurs only when "i = 3". As a result, among the flag data FLG[1] to FLG[n], only the flag data FLG[3] is set to "1".

[0146] In case CS_E2, at the stage when step S17 is reached, the flag data FLG[3] is set to "1". In the crossing probability determination device 10P, when there is one or more flag data having a value of "1" among the flag data FLG[1] to FLG[n] in step S17, the process proceeds to step S18a.

[0147] In step S18a according to case CS_E2, the controller 13 adds specific data to the input image IN1. At this time, the controller 13 associates the specific data with the person corresponding to the flag data of "1" (i.e., the person P[3] corresponding to the flag data FLG[3]). In step S18a according to case CS_E2, the controller 13 can store the input image IN1 with the specific data added in the storage unit 12 or the database 30 together with the partial correct label. The partial correct label includes the BBOX information generated for the input image IN1, and the method of generating the BBOX information itself is as described above. The partial correct label also includes the crossing probability set in step S33 or S37. Therefore, in case CS_E2, the crossing probabilities of the persons P[1], P[2], P[4], and P[5] are included in the partial correct label, but the crossing probability of the person P[3] is not included in the partial correct label.

[0148] Perform the operation of FIG. 16 or FIG. 17 on the M input images IN. Then, the M input images IN will be divided into a first input image group consisting of M A input images IN and a second input image group consisting of M B input images IN. "M = M A + M B ", and here, both M A and M B are assumed to have integer values of 2 or more. For each input image IN in the first input image group, a transition to step S19 occurs and a unit data set is generated. For each input image IN in the second input image group, a transition to step S18a occurs. Specific data is added only to each input image IN in the second input image group.

[0149] The specific data indicates that there is a BBOX for which the controller 13 cannot set the crossing probability. At an arbitrary timing, the operator OP can display each input image IN with the specific data added on the display device 21 using an arbitrary tool, and manually set the crossing probability of the person for whom the crossing probability is not set.

[0150] Here, focusing on one object, the operation of the crossing probability determination device 10P will be summarized. The controller 13 (probability determination unit 136) can execute a probability determination process for calculating and determining the crossing probability of the object on the road based on the input image IN having the images of the object including a person and the road. The controller 13 (evaluation unit 137) can evaluate the difficulty of the determination by the probability determination process (therefore, the accuracy of the calculated crossing probability) based on the relationship between the object and the road in the input image IN (see, for example, S32 and S36 in FIG. 8).

[0151] In the first reference method in which the operator OP designates all the crossing probabilities of the persons, manual work becomes enormous and learning data cannot be efficiently generated. In the second reference method in which all the crossing probabilities of the persons are mechanically determined without requiring manual work, a large amount of learning data can be generated in a short time. However, in many cases, it is difficult for the machine to correctly determine the crossing probability. As a result, it is difficult to ensure the quality of the learning data in the second reference method. According to the crossing probability determination device 10P, since the difficulty of the determination by the probability determination process is evaluated (since the accuracy of the calculated crossing probability is evaluated), it is possible to classify into those for which the determination by the probability determination process is easy and those for which the determination by the probability determination process is difficult. And it becomes possible to set the crossing probability only for those classified into the latter through the work of the operator OP. For this reason, compared with the first reference method, the work load of the operator OP is reduced, and compared with the second reference method, it is easy to ensure the quality of the learning data. That is, by using the evaluation result of the crossing probability determination device 10P, it becomes possible to efficiently generate a large amount of high-quality learning data.

[0152] The controller 13 (evaluation unit 137) classifies the difficulty of the determination by the probability determination process as low difficulty or high difficulty based on the relationship between the object and the roadway in the input image IN. By this classification, the accuracy of the calculated crossing probability by the probability determination process is classified as the first or second accuracy (see, for example, S32 and S36 in FIG. 8). Then, when the controller 13 classifies the difficulty of the determination by the probability determination process as low difficulty (when classifying the calculated crossing probability as the first accuracy), the controller 13 sets the crossing probability of the object based on the determination result of the probability determination process (for example, S33 or S37 in FIG. 8). On the other hand, when the controller 13 classifies the difficulty of the determination by the probability determination process as high difficulty (when classifying the calculated crossing probability as the second accuracy), the controller 13 associates specific data with the object and adds it to the input image IN (S18a in FIG. 16 or FIG. 17). As a result, it is only necessary to perform the operation of setting the crossing probability by the operator OP on the input image IN to which the specific data is added. Therefore, it is possible to efficiently generate a large amount of high-quality learning data.

[0153] <<Sixth Embodiment>> The sixth embodiment will be described. In the learning data generation step, the above-described setting of the BBOX and calculation of the crossing probability may be performed based on the input image IN using a reference model (not shown) different from the AI model 41 (see FIG. 1). An explanation will be provided for this.

[0154] An AI model Y (not shown) composed of a neural network is provided in a computer device X (not shown) having a hardware configuration equivalent to that of the learning device 40, and machine learning of the AI model Y is executed by the computer device. The learning data used for the machine learning of the AI model Y is, for convenience, referred to as learning data Z. The learning data Z has the same configuration as the learning data to be stored in the database 30, but the amount of data of the learning data Z may be less than the learning data to be stored in the database 30. The learning data Z may be generated by human work. Alternatively, a small amount of learning data generated by the method shown in each of the above embodiments may be used as the learning data Z.

[0155] Computer device X provides the learning images (specifically, the image data of the learning images) in learning data Z to AI model Y as problem data, and provides the BBOX information and crossing probability information in learning data Z as correct labels. AI model Y can execute the same type of object detection process as the object detection unit 132 (see FIG. 4), detect the people in the learning image, and estimate (infer) the BBOX of the detected people. In addition, for each person in the learning image, AI model Y estimates (infers) the crossing probability of the person to the road lane.

[0156] Computer device X evaluates the error between the BBOX and crossing probability estimated by AI model Y and the BBOX information and crossing probability information provided as correct labels. Then, computer device X uses a learning algorithm such as the error backpropagation method to perform the learning of AI model Y so that the error is reduced. In this learning, the parameters (weights, etc.) of AI model Y are adjusted. When a predetermined learning end condition is satisfied, such as the error converging to a sufficiently small value, the learning ends. The AI model Y after the learning ends is the above-mentioned reference model.

[0157] In the sixth embodiment, the reference model is provided in the learning data generation device 10. In the generation process of the learning data to be stored in the database 30, the reference model is responsible for the object detection process and the probability judgment process for the input image IN. That is, the reference model can set the BBOX and calculate the crossing probability for the input image IN. When the reference model sets a BBOX for a certain person in the input image IN, it calculates the crossing probability of the person with respect to the road lane RD of the person, and associates the calculated crossing probability with the person and adds it to the input image IN. When there are multiple people in the input image IN (that is, when the input image IN contains images of multiple people), the reference model sets the BBOX and calculates the crossing probability with respect to the road lane RD for each person, and adds the calculated crossing probability for each person to the input image IN for each person.

[0158] When combining the sixth embodiment with the first embodiment, since the crossing probability is calculated using a reference model separately from the unit process of FIG. 8, the crossing probability is not calculated in the unit process of FIG. 8. For example, the reference model may set the BBOX and calculate the crossing probability at step S12 of FIG. 7 or earlier. Similarly, when combining the sixth embodiment with the second embodiment, since the crossing probability is calculated using a reference model separately from the unit process of FIG. 12, the crossing probability is not calculated in the unit process of FIG. 12. For example, the reference model may set the BBOX and calculate the crossing probability at step S12 of FIG. 11 or earlier.

[0159] Consider case CS_F1 that proceeds to step S33 when the sixth embodiment is combined with the first or second embodiment. In step S33 related to case CS_F1, the evaluation unit 137 classifies the accuracy of the calculated crossing probability by the reference model into the first accuracy, and the probability setting unit 138 sets the calculated crossing probability by the reference model as the crossing probability of the subject as it is. Consider case CS_F2 that proceeds to step S37 when the sixth embodiment is combined with the first or second embodiment. In step S37 related to case CS_F2, the evaluation unit 137 classifies the accuracy of the calculated crossing probability by the reference model into the first accuracy, and the probability setting unit 138 sets the calculated crossing probability by the reference model as the crossing probability of the subject as it is. Therefore, in cases CS_F1 and CS_F2, a unit dataset in a state where the calculated crossing probability by the reference model is added to the input image IN as the crossing probability of the subject will be generated. That is, when a certain input image IN of interest corresponds to case CS_F1 or CS_F2, a unit dataset including the image data of the input image IN of interest, the BBOX information set by the reference model, and the crossing probability information indicating the calculated crossing probability by the reference model is included in the learning data in the database 30.

[0160] Consider case CS_F3 that proceeds to step S38 when combining the sixth embodiment with the first or second embodiment. In step S38 according to case CS_F3, the evaluation unit 137 classifies the accuracy of the calculated crossing probability by the reference model into the second accuracy. Also, in step S38 according to case CS_F3, the controller 13 sets 1 in the flag data FLG[i] (see FIG. 8 or FIG. 12). When 1 is set in the flag data FLG[i], a transition to step S18 occurs as described in the first or second embodiment (see FIG. 7 or FIG. 11).

[0161] In step S18 according to case CS_F3, the probability setting unit 138 executes the above-described inquiry process, and based on the response input information obtained from the operator OP in the inquiry process, sets the crossing probability of the person corresponding to the flag data of "1". That is, in step S18 according to case CS_F3, the probability setting unit 138 replaces the crossing probability of the person corresponding to the flag data of "1" with the crossing probability based on the response input information from the calculated crossing probability by the reference model. In other words, in step S18 according to case CS_F3, the probability setting unit 138 replaces the crossing probability of the person added to the input image IN (the crossing probability of the person corresponding to the flag data of "1") with the crossing probability based on the response input information from the calculated crossing probability by the reference model.

[0162] Alternatively, when case CS_F3 applies to a certain input image IN of interest, a modification may be made to not execute the processes of steps S18 and S19. In this case, the unit data set for the input image IN of interest is not generated, and thus the data related to the input image IN of interest is excluded from the learning data in the database 30.

[0163] <<Supplementary Note>> The road lane RD may be any type of road lane. However, a road where pedestrians or the like may be present in the vicinity is assumed as the road lane RD. Although a public road is mainly assumed as the road lane RD, the road lane RD may be a road provided in a factory (a road on which work vehicles in the factory pass).

[0164] A program for causing a computer device to execute any method described in the embodiments of the present invention, and a non-volatile recording medium storing the program are included within the scope of the embodiments of the present invention. The program for causing a computer device to execute any method described in the embodiments of the present invention may be a subprogram incorporated into an arbitrary main program or called by an arbitrary main program. The learning data generation device 10 and the learning device 40 are a kind of computer device. Any process in the embodiments of the present invention may be realized by hardware such as a semiconductor integrated circuit, software corresponding to the above program, or a combination of hardware and software.

[0165] The embodiments of the present invention can be appropriately modified in various ways within the scope of the technical idea shown in the claims. The above embodiments are merely examples of the embodiments of the present invention, and the meanings of the terms of the present invention or each constituent element are not limited to those described in the above embodiments. The specific numerical values shown in the above description are merely examples, and as a matter of course, they can be changed to various numerical values.

Explanation of Reference Numerals

[0166] 10 Learning data generation device 11 Communication unit 12 Storage unit 13 Controller 131 Input image acquisition unit 132 Object detection unit 133 Orientation detection unit 134 Region detection unit 135 Distance detection unit 136 Probability judgment unit 137 Evaluation unit 138 Probability setting unit 139 Data output unit 20 Interface 21 Display device 22 Input device 30 Database 40 Learning device 41 AI model OP Operator V1 Cooperative vehicle V2 Operational vehicle C1, C2 Camera IN, IN1, IN2 Input image RD Road lane SW_L, SW_R Sidewalk PX Crosswalk P[1]~P[5] Person BBOX[1]~BBOX[5] Bounding box 10P Crosswalk probability judgment device

Claims

1. For an input image showing an object including a person and a road, Based on the information about the position and orientation of the object with respect to the road, A controller that evaluates the probability added to the input image and the accuracy of the probability of the object crossing the road, An image processing apparatus.

2. When the accuracy does not meet the required value, the controller replaces the probability of crossing the road added to the input image. The image processing apparatus according to claim 1.

3. The controller executes a probability determination process of calculating the probability of crossing the road based on the input image and adding it to the input image. In the input image, the controller detects at least one of the orientation of the person's body and the orientation of the person's face with respect to the road, and classifies the accuracy of the probability of crossing the road calculated in the probability determination process based on the detection result into a first accuracy or a second accuracy. The image processing apparatus according to claim 2.

4. The controller executes a probability determination process of calculating the probability of crossing the road based on the input image and adding it to the input image. Based on the distance between the object and the road in the input image, the controller classifies the accuracy of the probability of crossing the road calculated in the probability determination process into a first accuracy or a second accuracy. The image processing apparatus according to claim 2.

5. The controller executes a probability determination process of calculating the probability of crossing the road based on the input image and adding it to the input image. In the input image, the controller executes a first detection process of detecting at least one of the orientation of the person's body and the orientation of the person's face with respect to the road, and a second detection process of detecting the distance between the object and the road in the input image, and classifies the accuracy of the probability of crossing the road calculated in the probability determination process based on the detection results of the first detection process and the second detection process into a first accuracy or a second accuracy. The probability of crossing the road determination apparatus according to claim 2.

6. The controller is configured to be able to execute a probability determination process of calculating the probability of the object crossing the road based on an input image showing an object including a person and a road. Based on the positional relationship between the object and the road in the input image, the controller classifies the accuracy of the probability of crossing the road calculated in the probability determination process into a first accuracy or a second accuracy. The accuracy of the crossing probability calculated in the probability determination process is higher than the second accuracy at the first accuracy, when the controller classifies the accuracy of the crossing probability calculated in the probability determination process as the first accuracy, the controller sets the crossing probability of the object based on the calculation result of the probability determination process; when the controller classifies the accuracy of the crossing probability calculated in the probability determination process as the second accuracy, the controller sets the crossing probability of the object based on the information input from the operator through the input device. The controller generates training data including the input image, the position and shape of the object on the input image, and the set crossing probability as a correct label. A training data generation device.

7. In the input image, the controller detects at least one of the orientation of the person's body with respect to the road and the orientation of the person's face, and classifies the accuracy of the crossing probability calculated in the probability determination process as the first accuracy or the second accuracy based on the detection result. The training data generation device according to claim 6.

8. The controller classifies the accuracy of the crossing probability calculated in the probability determination process as the first accuracy or the second accuracy based on the distance between the object and the road in the input image. The training data generation device according to claim 6.

9. In the input image, the controller executes a first detection process for detecting at least one of the orientation of the person's body with respect to the road and the orientation of the person's face, and a second detection process for detecting the distance between the object and the road in the input image, and classifies the accuracy of the crossing probability calculated in the probability determination process as the first accuracy or the second accuracy based on the detection results of the first detection process and the second detection process. The training data generation device according to claim 6.

10. For an input image showing an object including a person and a road, based on the information on the position and orientation of the person with respect to the road, evaluate the accuracy of the crossing probability of the object to the road, which is the probability added to the input image. An image processing method.

11. For an input image showing an object including a person and a road, based on the information on the position and orientation of the person with respect to the road, cause a computer to perform a method for evaluating the accuracy of the crossing probability of the object to the road, which is the probability added to the input image. , Image processing program.

Citation Information

Patent Citations

  • Annotation device, annotation method, and annotation program

    WO2022130516A1