Crossing prediction system, crossing prediction device, and crossing prediction method
The crossing prediction system addresses the challenge of accurately predicting pedestrian crossings by automatically generating teacher data from learning moving images, reducing manual labor and improving prediction accuracy.
Patent Information
- Application Number
- PCT/JP2024/041679
- Authority / Receiving Office
- WO · WO
- Patent Type
- Applications
- Current Assignee / Owner
- Priority Date
- 2023-11-28
- Filing Date
- 2024-11-25
- Publication Date
- 2025-06-05
AI Technical Summary
Existing systems face challenges in accurately predicting pedestrian crossings at crosswalks due to the time and effort required to manually label large datasets for machine learning, and insufficient data leads to inaccurate predictions.
A crossing prediction system that includes an analysis unit to determine whether a pedestrian has crossed a road from a learning moving image, and a teacher data generation unit to automatically generate teacher data with label information, reducing manual labor and enabling the creation of large datasets efficiently.
The system allows for accurate prediction of pedestrian crossings while reducing labor, enabling the generation of large amounts of high-quality teacher data that improves prediction accuracy.
Smart Images

Figure JP2024041679_05062025_PF_FP_ABST
Abstract
Description
Crossing prediction system, crossing prediction device, and crossing prediction method
[0001] This application claims priority to Japanese Application No. 2023-201091, filed on November 28, 2023, and incorporates by reference all the contents of said Japanese application.
[0002] Patent Document 1 describes a system that analyzes the line of sight of people in an image captured by a camera that captures the waiting area of a crosswalk, detects pedestrians waiting to cross the crosswalk, and, when a waiting pedestrian is detected, controls pedestrian and vehicle traffic lights to allow the pedestrian to cross the crosswalk.
[0003] Japanese Patent Application Laid-Open No. 2017-208141
[0004] One aspect of the crossing prediction system includes an analysis unit that analyzes training video captured from an imaging area adjacent to a road and determines whether a pedestrian shown in the training video has crossed the road, and a training data generation unit that generates training data including the training video and label information indicating the determination result of the analysis unit.
[0005] FIG. 1 is a schematic diagram of a crossing prediction system according to an embodiment. FIG. 2 is a diagram showing an imaging area. FIG. 3 is a block diagram showing the functional configuration of a roadside device and a server device. FIG. 4 is a block diagram showing the functional configuration of an analysis unit. FIG. 5(a) shows an example of a frame included in a learning video. FIG. 5(b) shows an example of a frame including position information of a passerby. FIGS. 6(a) and 6(b) show examples of a movement trajectory of a passerby. FIG. 7 is a block diagram showing an example of the hardware configuration of a roadside device and a server device. FIG. 8 is a sequence diagram showing a crossing prediction method according to an embodiment. FIG. 9(a) shows an example of a first frame in which a passerby enters the imaging area, and FIG. 9(b) shows an example of a second frame immediately before the passerby enters the crosswalk.
[0006] [Problem to be Solved by the Present Disclosure] In order to predict with high accuracy the likelihood of pedestrians crossing a crosswalk, it is conceivable to perform machine learning using images captured by a camera as training data to generate a prediction model that predicts whether pedestrians will cross the crosswalk. However, in order to predict with high accuracy the likelihood of pedestrians crossing a crosswalk, it is necessary to prepare and train a large amount of training data set labeled with true values. Preparing a large amount of manually labeled training data requires a great deal of time and effort. On the other hand, if the amount of training data is insufficient, it becomes difficult to predict with high accuracy the likelihood of pedestrians crossing a crosswalk.
[0007] Effect of the Present Disclosure According to the present disclosure, it is possible to predict with high accuracy the possibility of a pedestrian crossing a road while reducing the amount of work required.
[0008] [Description of Embodiments of the Present Disclosure] First, the contents of the embodiments of the present disclosure will be listed and described.
[0009] [1] A crossing prediction system according to one aspect of the present disclosure includes an analysis unit that analyzes training video captured from an imaging area including an area adjacent to a road and determines whether a pedestrian captured in the training video has crossed the road, and a training data generation unit that generates training data including the training video and label information indicating the determination result of the analysis unit.
[0010] This crossing prediction system determines whether a pedestrian captured in a training video has crossed a road, and generates training data including the training video and label information indicating the determination result. Therefore, the training data can be automatically generated from the training video, reducing the effort required to generate the training data. Furthermore, since it is possible to generate a large amount of training data without much effort, the accuracy of predicting whether a pedestrian will cross a road can be improved.
[0011] [2] The crossing prediction system of [1] may further include a prediction unit that inputs a video captured of the imaging area into a prediction model generated by machine learning training data, and predicts whether a pedestrian captured in the video will cross the road. In this case, it is possible to predict with high accuracy the possibility of a pedestrian crossing the road.
[0012] [3] In the crossing prediction system of [2] above, the prediction unit outputs the probability that a pedestrian shown in the video will cross the road, and the crossing prediction system may further include a control unit that controls a traffic light installed on the road so that the pedestrian shown in the video can cross the road when the probability is higher than a reference value. In this case, the traffic light can be controlled according to the predicted possibility that the pedestrian will cross the road.
[0013] [4] In the crossing prediction system described in any one of [1] to [3] above, the analysis unit may include a detection unit that identifies the position of a passerby in multiple frames included in the learning video, a tracking unit that identifies a movement trajectory of the passerby based on the position of the passerby in the multiple frames, and a determination unit that determines whether the passerby in the learning video has crossed a road based on the movement trajectory. By using the movement trajectory of the passerby, it is possible to determine with high accuracy whether the passerby has crossed a road.
[0014] [5] The crossing prediction system described in any one of [1] to [4] above may include a roadside device installed on or near a road and a server device capable of communicating with the roadside device via a network, wherein the roadside device includes an analysis unit, a training data generation unit, and a prediction unit, and the server device may include a receiving unit that receives training data from the roadside device, a prediction model generation unit that performs machine learning on the training data to generate a prediction model, and a transmitting unit that transmits the generated prediction model to the roadside device. In this case, the server device executes the process of generating the prediction model, which requires a large processing load, thereby reducing the processing load on the roadside device. Therefore, the roadside device can be made smaller.
[0015] [6] In the crossing prediction system described in any one of [1] to [5] above, the training data may further include at least one of information indicating attributes of passersby, information indicating a time period when the learning video was captured, information indicating the weather, and information indicating the road environment. In this case, a prediction model can be generated according to at least one of the attributes of passersby, the time period, the weather, and the road environment, thereby improving the accuracy of predictions of passersby crossing the road.
[0016] [7] In the crossing prediction system described in [5] or [6] above, the training data may further include identification information for identifying the roadside device. In this case, a prediction model can be generated according to the installation environment of the roadside device, thereby improving the accuracy of predicting road crossings by passersby.
[0017] [8] In the crossing prediction system described in [1] to [7] above, when the analysis unit determines that a pedestrian captured in the training video has crossed a road, the training data generation unit may extract from the training video a group of frames in time series from a first frame when the pedestrian enters the image capture area to a second frame immediately before the pedestrian enters the road, and remove from the training video frames all frames other than the group of frames in time series. In this case, training data can be generated that includes information about the movement of the pedestrian from the time when the pedestrian enters the image capture area to the time immediately before the pedestrian enters the road.
[0018] [9] A crossing prediction method according to one aspect includes the steps of: analyzing training videos captured of an imaging area including a surrounding area of a road, and determining whether a pedestrian captured in the training videos has crossed the road; and generating training data including the training videos and label information indicating the determination result of whether the pedestrian captured in the training videos has crossed the road. This crossing prediction method makes it possible to generate a large amount of training data without much effort, thereby improving the accuracy of predicting whether a pedestrian will cross the road.
[0019]
[10] The crossing prediction method described in [9] above may further include a step of inputting a video captured of the imaging area to a prediction model generated by machine learning training data, and predicting whether a passerby captured in the video will cross the road. In this case, it is possible to predict with high accuracy the possibility of a passerby crossing the road.
[0020]
[11] A crossing prediction device according to one aspect of the present disclosure includes an analysis unit that analyzes training videos captured in an imaging area including the surrounding area of a road and determines whether a pedestrian depicted in the training videos has crossed the road, and a training data generation unit that generates training data including the training videos and label information indicating the determination result of the analysis unit.
[0021]
[12] The crossing prediction device described in
[11] above may further include a prediction unit that inputs a video image of the imaging area into a prediction model generated by machine learning the teacher data, and predicts whether a passerby shown in the video image will cross the road.
[0022] [Details of the embodiments of the present disclosure] Specific examples of the embodiments of the present disclosure will be described below with reference to the drawings. The present invention is not limited to these examples, but is defined by the claims, and is intended to include all modifications within the meaning and scope equivalent to the claims. In the description of the drawings, the same elements are given the same reference numerals, and duplicate explanations will be omitted.
[0023] 1 is a schematic diagram of a crossing prediction system 1 according to an embodiment. The crossing prediction system 1 predicts whether a pedestrian will cross a road.
[0024] In the following explanation, a sidewalk 2 and a roadway 3 are provided, and an example will be described in which a prediction is made as to whether a pedestrian P traveling on the sidewalk 2 will cross a crosswalk 4 installed on the roadway 3. The pedestrian P is typically a pedestrian, but the pedestrian P also includes people riding vehicles such as bicycles, kick scooters, and wheelchairs, people using white canes, and people pushing strollers. As shown in FIG. 1 , pedestrian traffic lights S1 are installed on both sides of the crosswalk 4 in the direction of travel. A vehicle traffic light S2 is installed on the roadway 3 for vehicles passing through the crosswalk 4. The sidewalk 2 may be a shoulder or a roadside strip. The roadway 3 is a road that pedestrians P may cross. In the following explanation, a pedestrian P traveling on the sidewalk 2 who intends to cross the crosswalk 4 may be referred to as a "crosser."
[0025] As shown in FIG. 1 , the crossing prediction system 1 includes a camera 10, a roadside device 20, and a server device 30. The camera 10 is an imaging device that is installed, for example, on the top of a support pole 5 installed on a sidewalk 2 and captures moving images of an imaging area R1. Although not limited thereto, the installation height of the camera 10 may be 3 m or more from the ground surface. In one embodiment, as shown in FIG. 2 , the imaging area R1 is an area that includes a waiting area around the roadway 3. The waiting area is located in front of a crosswalk 4, and is an area on the sidewalk 2 where a pedestrian P who intends to cross the crosswalk 4 temporarily waits. As shown in FIG. 2 , the imaging area R1 may include a portion of the crosswalk 4.
[0026] The camera 10 may be disposed at an angle as long as it can capture an image of the image capture area R1. For example, the central axis (optical axis) of the camera 10 may be perpendicular to the ground surface or may be inclined relative to the ground surface. The camera 10 may be installed on the opposite side of the crosswalk 4 from the image capture area R1.
[0027] The camera 10 captures an image of the imaging region R1 at a predetermined frame rate to generate a moving image, which is then transmitted to the roadside unit 20 via wired or wireless communication.
[0028] The roadside device 20 is a computer such as a PLC (Programmable Logic Controller) equipped with a processor, a storage device, a communication device, etc., and is installed on the roadway 3 or in the vicinity of the roadway 3. For example, the roadside device 20 is provided on a support 5. The roadside device 20, for example, loads a program stored in the storage device and executes the loaded program on the processor, thereby realizing various functions described below.
[0029] The roadside device 20 is communicably connected to the camera 10, the server device 30, the pedestrian traffic light S1, and the vehicle traffic light S2. The roadside device 20 generates training data 52 for a prediction model 53 (described later) and has the function of predicting the presence of a pedestrian based on the prediction model 53.
[0030] Fig. 3 is a block diagram showing the functional configuration of the roadside device 20 and the server device 30. As shown in Fig. 3, the roadside device 20 has, as its functional configuration, an acquisition unit 21, an analysis unit 22, a teacher data generation unit 23, a transmission unit 24, a reception unit 25, a prediction unit 26, a control unit 27, and a storage unit 28.
[0031] The acquisition unit 21 acquires, from the camera 10, moving images of the imaging area R1 captured by the camera 10. The acquisition unit 21 outputs, from the moving images captured by the camera 10, moving images in which a passerby P appears to the analysis unit 22 as learning moving images.
[0032] The analysis unit 22 analyzes the learning video output from the acquisition unit 21 and determines whether or not a passerby P appearing in the learning video has crossed the crosswalk 4. Although not limited thereto, the analysis unit 22, for example, identifies the movement trajectory of the passerby P appearing in the learning video and determines whether or not the passerby P has crossed the crosswalk 4 based on the identified movement trajectory.
[0033] FIG. 4 is a block diagram showing an exemplary functional configuration of the analysis unit 22. As shown in FIG. 4, the analysis unit 22 includes a detection unit 41, a tracking unit 42, and a determination unit 43. The detection unit 41 acquires learning videos from the acquisition unit 21, identifies passersby P appearing in each frame (image) of the learning videos through image recognition processing, and outputs position information indicating the position of the identified passersby P. FIG. 5( a) shows an example of one frame included in the learning videos input to the detection unit 41. FIG. 5( b) shows an example of a frame including position information of the passersby P.
[0034] The position information of passerby P can be identified using known image recognition algorithms such as a convolutional neural network (CNN), histogram of oriented gradients (HOG), and you only look once (YOLO). Note that the detection unit 41 may output attribute information indicating the attributes of passerby P in addition to the position information of passerby P. The attribute information of passerby P is information indicating attributes such as the age, gender, and type of passerby P (pedestrian, cyclist, wheelchair user, kickboard user, white cane user, person pushing a stroller, etc.).
[0035] The detection unit 41 identifies the position information and attributes of the passerby P for all frames included in the learning video, and outputs the identified position information and attribute information to the tracking unit 42 .
[0036] The tracking unit 42 identifies the movement trajectory of the passerby P based on the position information of the passerby P in the multiple frames. For example, the tracking unit 42 connects the positions of the passerby P between the multiple frames arranged in time series to generate the movement trajectory of the passerby P. The movement trajectory of the passerby P can be generated using a known tracking algorithm such as a Kalman filter, a particle filter, Optical Flow, or Deep SORT.
[0037] 6( a) and 6(b) show examples of a movement trajectory 50 of a passerby P identified by the tracking unit 42. The tracking unit 42 outputs information indicating the movement trajectory 50 of the identified passerby P to the determination unit 43. Note that if multiple passersby P appear in the learning video, the tracking unit 42 may identify the movement trajectory 50 for each passerby P. In this case, the tracking unit 42 outputs an identifier (ID) that identifies the passerby P to the determination unit 43 in addition to the information indicating the movement trajectory 50 of the passerby P.
[0038] The determination unit 43 determines whether or not a passerby P captured in the learning video has crossed the crosswalk 4 based on the movement trajectory 50 of the passerby P. For example, as shown in FIGS. 6( a) and 6(b), the determination unit 43 sets a region R2 indicating the extent of the crosswalk 4 in each frame of the learning video, and determines whether or not a portion of the movement trajectory 50 is located within the region R2. As shown in FIG. 6(a), if a portion of the movement trajectory 50 is located within the region R2, the determination unit 43 determines that the passerby P has crossed the crosswalk 4. On the other hand, as shown in FIG. 6(b), if the movement trajectory 50 is located outside the region R2, the determination unit 43 determines that the passerby P has not crossed the crosswalk 4. The determination unit 43 outputs information indicating the determination result of whether or not the passerby P has crossed the crosswalk 4.
[0039] The teacher data generation unit 23 generates teacher data 52 including the training video and label information indicating the determination result of the determination unit 43. If it is determined that a passerby P shown in the training video has crossed the crosswalk 4, the teacher data generation unit 23 assigns label information to the training video indicating that a person crossing a crosswalk is present. On the other hand, if it is determined that a passerby P shown in the training video has not crossed the crosswalk 4, the teacher data generation unit 23 assigns label information to the training video indicating that a person crossing a crosswalk is not present. The teacher data generation unit 23 outputs the training video with the assigned label information to the transmission unit 24 as teacher data 52.
[0040] The teacher data 52 generated by the teacher data generator 23 may include attribute information indicating the attributes of the passerby P identified by the detector 41. The teacher data 52 may further include information indicating the time of day (e.g., morning, afternoon, night, etc.) when the learning video was captured, information indicating the weather (rain, snow, fog, etc.) when the learning video was captured, or information indicating the road environment (e.g., the width of the crosswalk 4, the size of the waiting area, the presence or absence of objects that may obstruct passage, such as installations or puddles, on the travel route, etc.). The teacher data 52 may also include identification information for identifying the roadside device 20.
[0041] The transmission unit 24 transmits the teacher data 52 generated by the teacher data generation unit 23 to the server device 30 via the network N. The network N is, for example, a wireless line such as a mobile communication network. The roadside device 20 performs the above-mentioned processing on a plurality of training videos to generate a training dataset including a plurality of pieces of teacher data 52. The training dataset is transmitted to the server device 30 by the transmission unit 24.
[0042] The server device 30 is a stationary or portable computer or workstation equipped with a processor, a storage device, a communication device, etc., and is typically located at a location remote from the roadside device 20. The server device 30, for example, loads a program stored in the storage device and executes the loaded program on the processor, thereby realizing various functions described below.
[0043] The server device 30 is capable of communicating with the roadside devices 20 via the network N. As will be described later, the server device 30 has a function of generating a prediction model 53 that predicts when a passerby P will cross the crosswalk 4 and transmitting the generated prediction model 53 to the roadside device 20. The server device 30 may be connected to multiple roadside devices 20 and generate a prediction model 53 for each of the multiple roadside devices 20. The server device 30 implements various functions, which will be described later, by, for example, loading a program stored in a storage device and executing the loaded program on a processor.
[0044] 3 , the server device 30 has, as its functional configuration, a receiving unit 31, a prediction model generating unit 32, a transmitting unit 33, and a storage unit 34. The receiving unit 31 receives the teacher data 52 transmitted from the transmitting unit 24 of the roadside device 20, and stores the teacher data 52 in the storage unit 34.
[0045] The prediction model generation unit 32 performs machine learning on a training dataset including multiple pieces of training data 52 to generate a prediction model 53. The prediction model 53 is a classifier constructed by machine learning using training video images and label information indicating the presence or absence of a pedestrian crossing. When a video image showing a passerby P is input, the prediction model 53 outputs the probability that the passerby P will cross the crosswalk 4 based on the feature quantities of the input video image. The prediction model 53 is generated by optimizing learning parameters using well-known machine learning algorithms such as a convolutional neural network and a recurrent neural network. The prediction model generation unit 32 stores the generated prediction model 53 in the memory unit 34 and outputs it to the transmission unit 33.
[0046] If the teacher data 52 includes attribute information of passersby P, the prediction model generation unit 32 may generate a prediction model 53 for each attribute of passersby P. Because the behavior of passersby P when crossing the crosswalk 4 differs depending on the attributes of the passersby P, generating a prediction model 53 for each attribute can improve the accuracy of predicting the presence or absence of a pedestrian crossing. Similarly, if the teacher data 52 includes information indicating the time of day when the learning video was captured, information indicating the weather, or information indicating the road environment, a prediction model 53 may be generated for each of these pieces of information.
[0047] Furthermore, when the server device 30 is connected to a plurality of roadside devices 20, the prediction model generation unit 32 may generate a prediction model 53 for each roadside device 20. This allows the generation of a prediction model 53 that is suited to the environment in which the roadside device 20 is installed.
[0048] The transmitter 33 transmits the generated prediction model 53 to the roadside device 20 via the network N. The receiver 25 of the roadside device 20 receives the prediction model 53 from the server device 30 and stores it in the memory 28.
[0049] The prediction unit 26 of the roadside device 20 acquires moving images of the imaging area R1 captured by the camera 10. The moving images acquired by the prediction unit 26 are moving images of the imaging area R1 at the current time (real time), unlike the learning moving images, which are moving images captured in the past. The prediction unit 26 reads out the prediction model 53 from the storage unit 28, inputs the acquired moving images into the prediction model 53, and predicts the presence of a passerby P crossing the crosswalk 4. For example, the prediction unit 26 outputs information indicating the probability that a passerby P captured in the moving images will cross the crosswalk 4 (hereinafter referred to as the "crossing probability") to the control unit 27.
[0050] When the probability of pedestrian P crossing output from the prediction unit 26 is higher than a reference value, the control unit 27 controls the pedestrian traffic light S1 and the vehicular traffic light S2 to allow pedestrian P to cross the crosswalk 4. Specifically, when the probability of pedestrian P crossing is higher than the reference value, the control unit 27 controls the pedestrian traffic light S1 and the vehicular traffic light S2 so that the signal display of the pedestrian traffic light S1 changes to green and the signal display of the vehicular traffic light S2 changes to red. If the pedestrian traffic light S1 is a push-button traffic light that changes the signal display by pressing a button, the control unit 27 may output a control signal to physically or electrically press the button. This allows pedestrian P to cross the crosswalk 4 safely. The reference value is a setting value that is preset by a designer.
[0051] In one embodiment, the roadside device 20 acquires training video images at regular intervals or each time a passerby P is detected, generates training data 52 including training video images and label information, and transmits the training data 52 to the server device 30. The server device 30 performs machine learning each time it receives training data 52 or acquires a training dataset including a certain number of training data sets 52, and updates the prediction model 53. The updated prediction model 53 is transmitted to the roadside device 20, and the prediction model 53 stored in the storage unit 28 of the roadside device 20 is updated. By repeating the above process, the prediction accuracy of the prediction model 53 is improved, making it possible to predict the presence or absence of a pedestrian with high accuracy. Note that the roadside device 20 may generate the training dataset so that the number of training data sets 52 including label information indicating the presence of a pedestrian is equal to the number of training data sets 52 including label information indicating the absence of a pedestrian.
[0052] Next, the hardware configuration of the roadside device 20 and the server device 30 will be described. FIG. 7 is a block diagram showing an example of the hardware configuration of the roadside device 20 and the server device 30. Each of the roadside device 20 and the server device 30 is configured by one or more computers. The computers include a CPU (processor) 101, a main memory unit 102, an auxiliary memory unit 103, a communication control unit 104, an input device 105, and an output device 106. Each of the roadside device 20 and the server device 30 is configured by one or more computers configured by this hardware and software such as programs. Note that the roadside device 20 and the server device 30 may include a GPU (Graphics Processing Unit) as a processor instead of or in addition to the CPU 101.
[0053] When the roadside device 20 or the server device 30 is configured by a plurality of computers, the plurality of computers may be connected locally or via a communication network such as the Internet or an intranet. This connection logically constructs one roadside device 20 or server device 30.
[0054] The CPU 101 executes an operating system, application programs, etc. The main memory unit 102 is composed of a ROM (Read Only Memory) and a RAM (Random Access Memory). The auxiliary memory unit 103 is a storage medium composed of a hard disk, flash memory, etc. The auxiliary memory unit 103 generally stores a larger amount of data than the main memory unit 102. The communication control unit 104 is composed of a network card or a wireless communication module. At least part of the communication function between the camera 10, the roadside unit 20, and the server device 30 is realized by the communication control unit 104. The input device 105 is composed of a keyboard, a mouse, a touch panel, a microphone for voice input, etc. The output device 106 is composed of a display, a printer, etc.
[0055] The auxiliary storage unit 103 stores programs and data necessary for processing. The programs cause the computer to execute each functional element of the roadside device 20 or the server device 30. The programs realize the functions of the roadside device 20 and the server device 30 in the computer. For example, the programs are read by the CPU 101 or the main storage unit 102, and cause at least one of the CPU 101, the main storage unit 102, the auxiliary storage unit 103, the communication control unit 104, the input device 105, and the output device 106 to operate. For example, the programs read and write data from and to the main storage unit 102 and the auxiliary storage unit 103.
[0056] The program may be provided in the form of a tangible storage medium such as a CD-ROM, a DVD-ROM, a semiconductor memory, etc. The program may be provided as a data signal via a communication network.
[0057] The roadside device 20 and the server device 30 do not necessarily have to be configured by a computer that operates by a program, and some or all of the functions of the roadside device 20 and the server device 30 may be implemented in an ASIC (Application Specific Integrated Circuit) that integrates logic circuits.
[0058] Next, a transverse prediction method according to an embodiment will be described. Fig. 8 is a sequence diagram showing the transverse prediction method according to an embodiment. This transverse prediction method is executed by the transverse prediction system 1 described above.
[0059] 8 , first, the acquisition unit 21 of the roadside device 20 acquires a moving image of the imaging region R1 from the camera 10 (step ST1). The moving image acquired by the acquisition unit 21 is a learning moving image among the moving images captured by the camera 10, in which a passerby P is captured.
[0060] Next, the analysis unit 22 analyzes the video (step ST2). For example, the analysis unit 22 identifies the position of a passerby P captured in multiple frames included in the learning video, and identifies a movement trajectory 50 of the passerby P by connecting the positions of the passerby P across the multiple frames (see FIGS. 6( a) and 6(b)). Then, the analysis unit 22 determines whether or not a person crossing the street is present based on the movement trajectory 50 of the passerby P (step ST3).
[0061] If it is determined that a pedestrian P is present, the training data generation unit 23 extracts a time-series group of frames from the first frame (see FIG. 9A) in which the passerby P enters the image capture area R1 to the second frame (step ST4), which is temporally subsequent to the first frame. The second frame is the frame immediately before (e.g., one second before) the passerby P enters the crosswalk 4 (area R2) (see FIG. 9B). The training data generation unit 23 then uses the video containing the extracted group of frames as the training video (hereinafter referred to as the "edited training video"). That is, the edited training video is a video obtained by removing from the pre-edited training video all frames except for the extracted group of frames (frames after the passerby P enters the crosswalk 4). The training data generation unit 23 then assigns label information indicating the presence of a pedestrian P crossing the image to the edited training video (step ST5), thereby generating training data 52.
[0062] On the other hand, if it is determined that no pedestrian is present, the training data generation unit 23 assigns label information indicating that no pedestrian is present to the training video (step ST6). In this case, the training video is a video including a time-series group of frames from the frame in which the passerby P enters the imaging area R1 to the frame in which the passerby P exits the imaging area R1. The transmission unit 24 transmits the training video with the assigned label information to the server device 30 as training data 52 (step ST7).
[0063] The receiving unit 31 of the server device 30 receives the teacher data 52 from the roadside unit 20 (step ST8). The prediction model generating unit 32 performs machine learning on the received teacher data 52 to generate a prediction model 53 (step ST9). If a prediction model 53 generated in the past is stored in the storage unit 34, the prediction model generating unit 32 updates the prediction model 53 stored in the storage unit 34 to the generated prediction model 53.
[0064] Next, the transmitter 33 transmits the generated teacher data 52 to the roadside device 20 (step ST10). The receiver 25 of the roadside device 20 receives the prediction model 53 from the server device 30 (step ST11). If a prediction model 53 generated in the past is stored in the memory 28, the receiver 25 updates the prediction model 53 stored in the memory 28 to the received prediction model 53.
[0065] Next, the prediction unit 26 acquires a moving image of the imaging region R1 from the camera 10 (step ST12). The moving image acquired by the prediction unit 26 is a moving image of the imaging region R1 at the current time captured in real time. The prediction unit 26 inputs the acquired moving image into the prediction model 53 and outputs the probability that the passerby P shown in the moving image will cross the crosswalk 4 (step ST13).
[0066] Next, the control unit 27 determines whether the crossing probability output from the prediction unit 26 is greater than a reference value (step ST14). If the crossing probability is greater than the reference value, the control unit 27 controls the pedestrian traffic light S1 and the vehicular traffic light S2 so that the passerby P can cross the crosswalk 4 (step ST15). On the other hand, if the crossing probability is equal to or less than the reference value, the control unit 27 ends the series of processes without controlling the pedestrian traffic light S1 and the vehicular traffic light S2.
[0067] As described above, the crossing prediction system 1 determines whether a passerby P shown in a learning video has crossed the crosswalk 4, and generates training data 52 including the learning video and label information indicating the determination result. This allows the training data 52 to be generated automatically from the learning video, thereby reducing the effort required to generate the training data 52. Furthermore, since it is possible to generate a large amount of training data 52 without much effort, the prediction accuracy of the prediction model 53 can be improved.
[0068] Furthermore, in the crossing prediction system 1, the server device 30 performs the process of generating the prediction model 53, which requires a large processing load, and therefore the processing load on the roadside device 20 can be reduced. Therefore, the roadside device 20 can be made smaller.
[0069] The above has described the cross-sectional prediction system 1 and the cross-sectional prediction method according to various embodiments, but the present invention is not limited to the above-described embodiments and various modifications can be made without departing from the spirit of the invention.
[0070] For example, in the above embodiment, the teacher data 52 is generated by the roadside device 20, and the prediction model 53 is generated by the server device 30. However, the functional configurations of the roadside device 20 and the server device 30 may be assigned to either the roadside device 20 or the server device 30. That is, the server device 30 may include part or all of the functional configuration of the roadside device 20, or the roadside device 20 may include part or all of the functional configuration of the server device 30. For example, the roadside device 20 may include the prediction model generation unit 32 in addition to the acquisition unit 21, the analysis unit 22, the teacher data generation unit 23, the prediction unit 26, and the control unit 27, generate the teacher data 52, and generate the prediction model 53 by machine learning the teacher data 52. Conversely, the server device 30 may include the acquisition unit 21, the analysis unit 22, the teacher data generation unit 23, the prediction unit 26, and the control unit 27 in addition to the prediction model generation unit 32, generate the teacher data 52, and generate the prediction model 53 by machine learning the teacher data 52. In this case, the server device 30 can also output the probability of the passerby P crossing based on the prediction model 53.
[0071] Furthermore, in the above embodiment, whether or not a passerby P has crossed the crosswalk 4 is determined based on the movement trajectory 50 of the passerby P in the learning video, but any method can be used to determine whether or not the passerby P has crossed the crosswalk 4. For example, if the prediction model 53 has sufficient prediction accuracy, the learning video may be input to the prediction model 53, and if the output probability of the passerby P crossing the crosswalk is equal to or greater than a predetermined threshold, it may be determined that the passerby P has crossed the crosswalk 4.
[0072] Furthermore, in the above embodiment, the prediction model 53 is used to predict whether or not a pedestrian P will cross the crosswalk 4, but the crossing prediction system 1 can also be applied to roads where a crosswalk 4 is not installed. In this case, the crossing prediction system 1 predicts whether or not a pedestrian P shown in a video image will cross the road. Furthermore, the crossing prediction system 1 only needs to predict the presence of a pedestrian, and does not necessarily have to control the pedestrian traffic light S1 and the vehicle traffic light S2.
[0073] REFERENCE SIGNS LIST 1...Crossing prediction system 2...Sidewalk 3...Roadway 4...Crosswalk 5...Support pole 10...Camera 20...Roadside device 21...Acquisition unit 22...Analysis unit 23...Teacher data generation unit 24...Transmission unit 25...Reception unit 26...Prediction unit 27...Control unit 28...Memory unit 30...Server device 31...Reception unit 32...Prediction model generation unit 33...Transmission unit 34...Memory unit 41...Detection unit 42...Tracking unit 43...Determination unit 50...Movement trajectory 52...Teacher data 53...Prediction model 101...CPU (processor) 102...Main memory unit 103...Auxiliary memory unit 104...Communication control unit 105...Input device 106...Output device N...Network P...Passerby R1...Image capture area R2...Area S1...Pedestrian traffic light S2...Vehicle traffic light
Claims
1. A crossing prediction system comprising: an analysis unit that analyzes training video captured in an imaging area including the area surrounding a road, and determines whether a pedestrian appearing in the training video has crossed the road; and a teacher data generation unit that generates teacher data including the training video and label information indicating the determination result of the analysis unit.
2. The crossing prediction system of claim 1, further comprising a prediction unit that inputs a video captured of the imaging area into a prediction model generated by machine learning the training data, and predicts whether a pedestrian captured in the video will cross the road.
3. The crossing prediction system of claim 2, wherein the prediction unit outputs a probability that a pedestrian shown in the video will cross the road, and the crossing prediction system further includes a control unit that controls a traffic light installed on the road to allow the pedestrian shown in the video to cross the road when the probability is higher than a reference value.
4. A crossing prediction system as described in any one of claims 1 to 3, wherein the analysis unit includes: a detection unit that identifies the position of the passerby in multiple frames included in the training video; a tracking unit that identifies a movement trajectory of the passerby based on the position of the passerby in the multiple frames; and a judgment unit that determines whether the passerby appearing in the training video has crossed the road based on the movement trajectory.
5. A cross-sectional prediction system as described in claim 2 or claim 3, comprising: a roadside device installed on or in the vicinity of the road; and a server device capable of communicating with the roadside device via a network, wherein the roadside device comprises the analysis unit, the teacher data generation unit, and the prediction unit, and the server device comprises a receiving unit that receives the teacher data from the roadside device, a predictive model generation unit that performs machine learning on the teacher data to generate the predictive model, and a transmitting unit that transmits the generated predictive model to the roadside device.
6. A crossing prediction system as described in any one of claims 1 to 5, wherein the teacher data further includes at least one of information indicating the attributes of the passersby, information indicating the time of day when the learning video was captured, information indicating the weather, and information indicating the road environment.
7. The crossing prediction system according to claim 5, wherein the teacher data further includes identification information for identifying the roadside device.
8. A crossing prediction system as described in any one of claims 1 to 7, wherein, when the analysis unit determines that the pedestrian appearing in the training video has crossed the road, the teacher data generation unit extracts from the training video a time-series group of frames from a first frame in which the pedestrian enters the imaging area to a second frame immediately before the pedestrian enters the road, and removes from the training video frames excluding the time-series group of frames.
9. A crossing prediction method comprising: a step of analyzing a training video captured in an imaging area including the surrounding area of a road, and determining whether a pedestrian shown in the training video has crossed the road; and a step of generating teacher data including the training video and label information indicating the determination result of whether the pedestrian shown in the training video has crossed the road.
10. The crossing prediction method according to claim 9, further comprising a step of inputting a video captured of the imaging area into a prediction model generated by machine learning the training data, and predicting whether a pedestrian captured in the video will cross the road.
11. A crossing prediction device comprising: an analysis unit that analyzes training video captured in an imaging area including the area surrounding a road, and determines whether a pedestrian appearing in the training video has crossed the road; and a teacher data generation unit that generates teacher data including the training video and label information indicating the determination result of the analysis unit.
12. The crossing prediction device as described in claim 11, further comprising a prediction unit that inputs a video captured of the imaging area into a prediction model generated by machine learning the training data, and predicts whether a pedestrian captured in the video will cross the road.
Citation Information
Patent Citations
Attribute-based pedestrian prediction
JP2022527072A
Track classification
JP2023525054A
Systems and methods for predicting a pedestrian movement trajectory
US20220171065A1
Potential collision warning system based on road user intent prediction
US20220324441A1
Detection system, detection device, and detection device installation method
WO2023195355A1