Pedestrian activity recognition using limited data and meta-learning

By using Siamese neural networks for unsupervised learning, automatic labeling and prediction of pedestrian activities, the existing technology solves the reliance on manually labeled datasets and achieves efficient and accurate pedestrian activity recognition and avoidance in autonomous or semi-autonomous vehicles.

CN114258560BActive Publication Date: 2025-09-30VOLKSWAGEN AG
View PDF 0 Cites 0 Cited by

Patent Information

Application Number
CN202080058496.1
Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Priority Date
2019-08-20
Filing Date
2020-08-13
Publication Date
2025-09-30
Estimated Expiration
2040-08-13

AI Technical Summary

Technical Problem

Existing technologies require large manually annotated datasets to identify and interpret pedestrian activities around vehicles, which is time-consuming and costly, and the complex manually adjusted human motion models are difficult to generalize to new or unseen conditions.

Method used

A Siamese neural network is trained to create clusters of similar activities in an unsupervised manner using datasets from multiple image capture devices, automatically annotate activities, train a spatial-temporal intention prediction model, recognize and predict pedestrian activities, and perform autonomous maneuvers through a vehicle control component.

Benefits of technology

It reduces the reliance on manually annotated data, improves the efficiency and accuracy of identifying and predicting pedestrian activities, and enables safe navigation and avoidance of pedestrians in autonomous or semi-autonomous vehicles.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN114258560B_ABST
    Figure CN114258560B_ABST
Patent Text Reader

Abstract

Pedestrian activity recognition is embodied in methods, systems, non-transitory computer-readable media, and vehicles. A Siamese neural network is trained to recognize multiple pedestrian activities using recordings of the same pedestrian activity from two or more independent training image capture devices. The Siamese neural network is deployed with continuous data collection from additional image capture devices to create a dataset of clusters of similar activities in an unsupervised manner. A spatio-temporal intent prediction model is then trained that can be deployed to recognize and predict pedestrian activities. Based on the likelihood of a particular pedestrian activity occurring or currently in progress, automated vehicle maneuvers can be performed to navigate the situation.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present disclosure relates to recognizing and interpreting scenes around vehicles. In particular, the present disclosure relates to accurately identifying the activities currently being performed by pedestrians around vehicles and predicting the next activities of pedestrians. Background Art

[0002] Correctly identifying and interpreting the scene around a vehicle is essential for enabling autonomous or semi-autonomous vehicles to safely maneuver around or otherwise avoid obstacles and pedestrians. Properly programming these intelligent vehicles with intelligent functionality typically requires very large, annotated datasets to create and train supervised machine learning models to classify pedestrian activity. This data is often manually labeled, which is time-consuming and expensive. These complex, hand-tuned models based on human motion do not always generalize to new or unseen conditions.

[0003] Therefore, there is a need for an effective method for pedestrian activity recognition using limited data and meta-learning. Summary of the Invention

[0004] Methods, systems, and non-transitory computer-readable media for pedestrian activity recognition are disclosed. Further disclosed is a vehicle incorporating a pedestrian activity recognition system. In an illustrative embodiment, a Siamese neural network is trained to recognize multiple pedestrian activities by training it based on two or more inputs, where the inputs are recordings of the same pedestrian activity from two or more separately trained image capture devices. The Siamese neural network is deployed using continuous data collection from the additional image capture devices to create a dataset of multiple activity clusters of similar activities in an unsupervised manner. The Siamese neural network automatically annotates the activities to create an annotated prediction dataset and an annotated non-prediction dataset. The spatiotemporal data samples from the non-prediction annotated dataset and the spatiotemporal data samples from the prediction annotated dataset are then used as input to train a spatiotemporal intent prediction model. This prediction model can then be deployed to recognize and predict pedestrian activities. Based on the likelihood of a particular pedestrian activity occurring or currently in progress, automatic vehicle maneuvers can be performed to navigate a situation. A method for pedestrian activity recognition according to the present invention includes: training a Siamese neural network to recognize multiple activities by training it based on two or more inputs, wherein the inputs are records of the same pedestrian activity from two or more separate training image capture devices; deploying the Siamese neural network model using continuous data collection from additional activity image capture devices to create a dataset of multiple activity clusters of similar activities in an unsupervised manner; employing the Siamese neural network to annotate activities as predicted activities or non-predicted activities so as to create annotated predicted datasets and annotated non-predicted datasets; using spatial-temporal data samples from the annotated non-predicted dataset and spatial-temporal data samples from the annotated predicted dataset as input to train a spatial-temporal intention prediction model; and deploying the intention prediction model to assign the likelihood of a specific activity. Preferably, deploying the Siamese neural network model to cluster similar activities in an unsupervised manner using continuous data collection from the additional activity image capture device comprises: inputting into the Siamese neural network comprising the plurality of activity clusters: the output from the additional activity image capture device; and the dataset of the plurality of activity clusters; determining, by the Siamese neural network, a measure of similarity between the additional activity image capture device output and a data sample from each of the plurality of activity clusters to determine whether the additional activity matches an existing cluster sample; and detecting an activity if the additional activity image capture device output belongs to an activity cluster in the plurality of activity clusters. The method further comprises: creating a new cluster associated with the current activity output if the measure of similarity exceeds a specified measure range for all activity clusters in the plurality of activity clusters, and adding the new cluster to the plurality of activity clusters.Training the Siamese neural network to recognize the plurality of activities may include: creating a similar dataset from at least an output from a first training image capture device of the two or more training image capture devices and a synchronized output from a second training image capture device of the two or more training image capture devices, wherein the outputs reflect the same pedestrian activity; creating a dissimilar dataset from the output from the first training image capture device and a delayed output from the second training image capture device, wherein the outputs reflect different pedestrian activities; and creating a dataset including a plurality of activity clusters by training the Siamese neural network based on the similar dataset and the dissimilar dataset. The method may further include refining the similar dataset and the dissimilar dataset by applying rule-based heuristics to the delayed output from the second training image capture device and the output from the first training image capture device to evaluate dissimilarity. Employing the Siamese neural network to annotate activities as predicted activities or non-predicted activities to create annotated predicted datasets and annotated non-predicted datasets may include: executing an annotation algorithm employing the Siamese neural network to automatically annotate image capture device output associated with an activity captured during a specified time period prior to a detected activity to label it as predictive of the detected activity to create a dataset annotated as predicted activity; and creating negative samples of pedestrian activities captured during a specified time period prior to a period in which one of the plurality of activity clusters was not detected from the image capture device output in which the activity was not detected, and labeling the negative samples as not predictive of the one of the plurality of activity clusters to create a dataset annotated as not predictive of activity. Deploying the intent prediction model to assign a likelihood of a particular activity includes assigning a "1" when an activity is predicted and a zero when an activity is not predicted. The method may further include performing automated vehicle maneuvers based on the assignment of the likelihood of the particular activity.The present invention also accordingly provides a system for identifying pedestrian activities, comprising: one or more processors; one or more storage devices storing computer code on the one or more storage devices, the computer code comprising a Siamese neural network; wherein executing the computer code causes the one or more processors to perform the following method: training the Siamese neural network based on two or more inputs to identify multiple activities, wherein the inputs are records of the same pedestrian activities from two or more separate training image capture devices; deploying the Siamese neural network model using continuous data collection from additional activity image capture devices to create a dataset of multiple activity clusters of similar activities in an unsupervised manner; employing the Siamese neural network to annotate activities as predicted activities or non-predicted activities to create annotated predicted datasets and annotated non-predicted datasets; using spatial-temporal data samples from the annotated non-predicted dataset and spatial-temporal data samples from the annotated predicted dataset as inputs to train a spatial-temporal intention prediction model; and deploying the intention prediction model to assign the likelihood of a specific activity. In the system, deploying the Siamese neural network model to cluster similar activities in an unsupervised manner using continuous data collection from the additional activity image capture device includes the computer code causing the one or more processors to perform the following steps: inputting into the Siamese neural network including the plurality of activity clusters: the output from the additional activity image capture device; and the dataset of the plurality of activity clusters; determining, by the Siamese neural network, a measure of similarity between the additional activity image capture device output and a data sample from each of the plurality of activity clusters to determine whether the additional activity matches an existing cluster sample; and detecting an activity if the additional activity image capture device output belongs to an activity cluster in the plurality of activity clusters. The system may further include the one or more storage devices having computer code stored thereon that, when executed, causes the one or more processors to perform the following steps: creating a new cluster associated with the current activity output if the measure of similarity exceeds a specified measure range for all activity clusters in the plurality of activity clusters, and adding the new cluster to the plurality of activity clusters.In the system, training the Siamese neural network to recognize the plurality of activities includes the computer code causing the one or more processors to perform the following steps: creating a similar dataset from at least an output from a first training image capture device of the two or more training image capture devices and a synchronized output from a second training image capture device of the two or more training image capture devices, wherein the outputs reflect the same pedestrian activity; creating a dissimilar dataset from the output from the first training image capture device and a delayed output from the second training image capture device, wherein the outputs reflect different pedestrian activities; and creating a dataset comprising a plurality of activity clusters by training the Siamese neural network based on the similar dataset and the dissimilar dataset. The system further includes the one or more storage devices having computer code stored thereon, the computer code, when executed, causing the one or more processors to perform the following steps: refining the similar dataset and the dissimilar dataset by applying rule-based heuristics to the delayed output from the second training image capture device and the output from the first training image capture device to evaluate dissimilarity. In the system, employing the Siamese neural network to annotate activities as predicted activities or non-predicted activities to create annotated predicted datasets and annotated non-predicted datasets includes the computer code causing the one or more processors to perform the following steps: executing an annotation algorithm employing the Siamese neural network to automatically annotate image capture device output associated with activities captured during a specified time period prior to a detected activity to label it as predictive of the detected activity to create a dataset annotated as predicted activities; and creating negative samples of pedestrian activities captured during a specified time period prior to a period in which one of the plurality of activity clusters was not detected from the image capture device output in which the activity was not detected, and labeling the negative samples as not predictive of the one of the plurality of activity clusters to create a dataset annotated as not predictive of activities. In the system, deploying the intent prediction model to assign the likelihood of a particular activity includes the computer code causing the one or more processors to perform the steps of assigning a "1" when an activity is predicted and a zero when an activity is not predicted. The system further comprises the one or more storage devices having computer code stored thereon, the computer code, when executed, causing automatic vehicle maneuvering based on the assignment of the likelihood of the particular activity. The present invention also provides an autonomous or semi-autonomous controlled vehicle, wherein the vehicle comprises the system for identifying pedestrian activity.The vehicle further comprises: a vehicle control assembly; and an actuator electronically connected to the vehicle control assembly; wherein the one or more storage devices have computer code stored thereon, which, when executed, causes the actuator to initiate the vehicle maneuver through the vehicle control assembly. The present invention also provides a non-transitory computer-readable medium having computer code stored thereon, which, when executed on one or more processors, causes a computer system to perform the above-described method of the present invention. BRIEF DESCRIPTION OF THE DRAWINGS

[0005] The detailed description refers to the drawings, all of which depict illustrative embodiments.

[0006] Figure 1 Depicts illustrative types of pedestrian activities that a neural network can be trained to recognize.

[0007] Figure 2 An illustrative timeline depicting a pedestrian entering a highway from a curb.

[0008] Figure 3 is a block diagram of an illustrative Siamese neural network that generates an output indicating that two images are similar.

[0009] Figure 4 is a block diagram of an illustrative Siamese neural network that generates an output indicating that two images are dissimilar.

[0010] Figure 5 is a block diagram depicting an illustrative method of obtaining images and creating a dataset to train a Siamese neural network to recognize pedestrian activities.

[0011] Figure 6 is a flowchart of an illustrative approach to training and deploying a Siamese neural network to recognize and predict pedestrian activities.

[0012] Figure 7 is a block diagram of an illustrative autonomous or semi-autonomous vehicle and an associated drive system that recognizes pedestrian activity and causes the vehicle to perform actions based on the recognized activity.

[0013] Figure 8 As Figure 5 A block diagram of an illustrative embodiment of a computing device that drives components of a system. DETAILED DESCRIPTION

[0014] The drawings and descriptions provided herein may have been simplified to illustrate aspects that are relevant to understanding the devices, systems, and methods, while eliminating other aspects that may be found in typical devices, systems, and methods for the purpose of clarity. Those skilled in the art will appreciate that other elements or operations may be desirable or necessary for implementing the devices, systems, and methods described herein. Because such elements and operations are well known in the art, and because they do not contribute to a better understanding of the present disclosure, discussion of such elements and operations may not be provided herein. However, the present disclosure is considered to inherently include all such elements, modifications to the aspects that may be implemented by those skilled in the art, and variations. All embodiments illustrate the broader scope of the concepts described herein.

[0015] The disclosed embodiments provide methods, systems, vehicles, and non-transitory computer-readable media for identifying pedestrian activities. The illustrative embodiments can identify and interpret the scene surrounding a vehicle, including identifying different activities in which pedestrians may engage, and predicting the next activity in which a pedestrian may be involved. Based on the detected or predicted pedestrian activities, the vehicle can automatically perform safe maneuvers, such as controlling the vehicle's speed or steering the vehicle in a new direction. Additionally or alternatively, a warning signal can be generated when a pedestrian activity is detected or predicted.

[0016] A neural network is a set of algorithms designed to recognize patterns. Illustrative embodiments of a pedestrian recognition method and system use a neural network to group data based on similarities found in videos of pedestrian activity. The neural network is trained based on labeled datasets. The datasets can be labeled automatically (i.e., in an unsupervised or semi-supervised manner). The illustrative embodiments include steps for further training the neural network and deploying the trained neural network to recognize pedestrian activity. Further training allows the neural network to predict pedestrian activity.

[0017] Figure 1Depicts illustrative types of pedestrian activities that a neural network can be trained to recognize. Activity types include walking, jogging, running, boxing, hand waving, and hand clapping. Four different examples are provided for each activity type. Using machine learning, a computer can be trained to identify similarities between different images of the same type of activity. For example, a jogger typically takes a shorter stride than a runner. There may be additional identifying features that can be identified in the video, such as arm movements, timing of changes in the pedestrian's posture and position within the environment. For pedestrians, activity annotations can include the activity type and may also include more specific labels, such as walking on the sidewalk, walking towards a paved road, crossing the road, walking in a crosswalk, running, sitting on the curb, waving to another pedestrian, waving to a camera, looking in the direction of a camera or vehicle, walking hurriedly, and walking carefully. In traditional supervised learning, images are manually annotated with activity type labels. The illustrative embodiments described herein automatically annotate or label images.

[0018] Figure 2 An illustrative timeline depicting a pedestrian leaving the curb and entering the roadway. Figure 2 The diagram further depicts how a driver might react to encountering pedestrian activity. Throughout the depicted time period, a pedestrian crosses the road. This is indicated by the top horizontal line labeled "Crossing." During a portion of this time period, the pedestrian, for example, looks away from the direction of travel or intended travel. This is shown in the second image and presented in the third row of the pedestrian activity list. This activity can be identified, for example, by a change in the position of the pedestrian's head or the direction of their gaze. In the third image from the left, the pedestrian uses a hand gesture during a small portion of the entire time period. The diagram also shows two non-contiguous portions of the depicted time period in which the pedestrian is moving quickly (e.g., running), indicated on the fourth row of the pedestrian activity list. Between the two portions of the time period (during which the pedestrian was running), the pedestrian has slowed down. The images are time-stamped and annotated with activity names (e.g., Crossing, Hand Gesture, Gaze, Moving Quickly, and Slowing Down). The use of timestamps eliminates the need for a supervised learning model, simplifying the process and reducing the required time and computing power. In conventional supervised models, annotations are associated with images through manual input. In illustrative embodiments of the disclosed methods and systems, tagging is performed automatically.

[0019] exist Figure 2, below the Pedestrian Activity label, illustrative driver actions are shown that represent how a driver might react to a pedestrian activity. These can also be considered vehicle actions, particularly for autonomous or semi-autonomous vehicles. During an initial approximately 0.03 second period, the vehicle is moving slowly. The vehicle then decelerates further from just before the 0.03 second point to approximately the 0.08 second point. The vehicle then begins to stop and remains stopped until just after the 0.09 second point. The vehicle then begins to move slowly again. The duration, action type, and start time of each action with respect to pedestrian activity are illustrative and may vary depending on the activity and vehicle capabilities.

[0020] In an illustrative embodiment, a Siamese neural network is trained to recognize multiple pedestrian activities. Data is obtained as input to the Siamese neural network, comprising two or more recordings of the same pedestrian activity from two or more separate training image capture devices. The pedestrian activity can be, for example, running, gazing, walking, jogging, waving, or any other pedestrian activity that can be distinguished and classified from other pedestrian activities.

[0021] The term "Siamese neural network" is used herein to refer to a model trained on two or more different inputs that enables it to predict whether the inputs are 'similar' or 'dissimilar'. The term "Siamese neural network" will be used herein for any model with these capabilities. In an illustrative embodiment, a spatiotemporal variant of a Siamese neural network is used that is capable of indicating whether two or more image sequences or videos are similar or dissimilar. In an illustrative embodiment, data from different timestamps separated in time by one minute, for example, will be dissimilar approximately 90% of the time. The greater the difference in the timestamps, the more dissimilar the predicted activity will be.

[0022] Figure 3 is a block diagram of an illustrative Siamese neural network 300 that generates an output indicating that two images are similar. A first image 302 of a person running is compared to a second image 304 also of a person running. The Siamese neural network is trained using weights to maximize the accuracy of the neural network. Training can be accomplished by making incremental adjustments to the weights. The weighting encourages two very similar images to be mapped to the same location. Here, the Siamese neural network is fully trained for use as an analysis tool for new image inputs (i.e., first image 302 and second image 304).

[0023] The twin networks 306 and 308 are joined by a function 310. The input images, i.e., the first image 302 and the second image 304, are filtered through one or more hidden or intermediate layers in a feature extraction process. Each filter picks up a different signal or feature to determine how much overlap exists between the new image and images or reference images for various types of pedestrian activities. As the image passes through the various filters, it is described mathematically. Various types of feature engineering can be used to map image features. A loss or cost function is calculated based on feature similarity. In the illustrative embodiment, a ternary cost function is used. Thresholds or ranges are provided as a basis for representing images as "similar" or "dissimilar."

[0024] The illustrative embodiments focus on detecting similarities and dissimilarities between multiple different activities. This approach, known as "metric learning," is a field of meta-learning within the broader context of machine learning. This approach can have significant advantages: significantly less annotated data, even orders of magnitude less, is required to learn new activities or categories.

[0025] Figure 4 FIG4 is a block diagram of an illustrative Siamese neural network 400 that generates an output indicating that two images are dissimilar. A first image 402 depicts a person running. A second image 404 depicts a person standing with their arms raised above their head. Siamese networks 406 and 408 are trained to process the images and provide a similarity measure in block 410 that does not meet a threshold for being designated as “similar.” Accordingly, the system outputs a conclusion that the images are dissimilar.

[0026] against Figure 3 and Figure 4 The described functionality can be used to predict whether two different pedestrian activities are the same or different, predict which known activities a new activity matches, and predict new activities that are not previously known in observational data. Siamese neural networks can also be used to identify new categories of pedestrian activities. Furthermore, Siamese neural networks can be configured to predict what subsequent activities an observed pedestrian may engage in by categorizing activity records based on what they will subsequently do, rather than what they are currently doing.

[0027] Figure 5 is a block diagram depicting an illustrative method for obtaining images and creating a dataset to train a Siamese neural network to recognize pedestrian activities. In the illustrative embodiment, data is collected to train a Siamese neural network to recognize multiple activities by training it based on two or more inputs, where the inputs are recordings of the same pedestrian activity from two or more separate training image capture devices 502, 504. Similar datasets 506 and dissimilar datasets 508 are created and stored in a memory 512 of a type suitable for and compatible with the intended pedestrian recognition method and system.

[0028] Figure 6 FIG6 is a flow chart of an illustrative method for training and deploying a Siamese neural network to recognize and predict pedestrian activities. In step 602, a similar dataset 506 is created from the output from a first training image capture device 502 and the synchronized output from a second training image capture device 504, where the outputs reflect the same pedestrian activities. Note that multiple training image capture devices may be used, with the first training image capture device and the second training image capture device being part of the plurality of training image capture devices.

[0029] It can be assumed that synchronized image capture devices record the same pedestrian activity because they are capturing images at the same time and are positioned to capture pedestrian activity in the same spatial region. Accordingly, the data collected from each camera can be automatically annotated as "similar" for use in training the Siamese neural network.

[0030] In step 604, a distinct dataset 508 is created from the output from the first training image capture device 502 and the delayed output 510 from the second training image capture device 504, where the output reflects different pedestrian activities. The delay is predefined and can be, for example, 30 seconds, or, by way of further example, can be in the range of approximately 10 seconds to approximately 30 seconds. It can be assumed that the delay produces images of different pedestrian activities, and therefore the images can be automatically annotated as distinct. The length of the delay can be selected, for example, based on the location of the image capture device, where the location can induce specific types of pedestrian activities. Pedestrian activities can vary depending on the environment, which can affect how pedestrians react or what activities they engage in and the order of such activities.

[0031] Optionally, in step 606 , similar and dissimilar datasets 506 , 508 are refined by applying rule-based heuristics 514 to the delayed output 510 from the second training image capture device 504 and to the output of the first training image capture device 502 to assess dissimilarity.

[0032] In step 608, the illustrative embodiment further includes creating a dataset comprising a plurality of pedestrian activity clusters. The activity cluster dataset is created by training a Siamese neural network based on the similar dataset and the dissimilar dataset to create a dataset comprising a plurality of activity clusters. Thus, the activity cluster dataset is created in an unsupervised manner.

[0033] More specifically, in an illustrative embodiment, a Siamese neural network may be trained in step 608 using the similar dataset created in step 602 and the dissimilar dataset created in step 604. The neural network is trained to determine whether two datasets input into the Siamese neural network belong to the same activity, and therefore should exist in the same activity cluster, or whether they belong to different activities.

[0034] In step 610, a Siamese neural network can be deployed using the continuous data collection input from the additional image capture device and the input from the database containing the multiple activity clusters formed in step 608. Each activity can be stored by its name or a numerical ID number corresponding to other samples classified into the ID of the cluster created in step 608. This can be achieved by inputting the output from the additional activity image capture device into the Siamese neural network including the multiple activity clusters and also inputting the database of the multiple activity clusters into the Siamese neural network. The additional image capture device can be one of the training image capture devices 502, 504, but for the sake of clarity, it will be referred to as the additional image capture device when used for continuous data collection.

[0035] In step 612, the Siamese neural network determines a measure of similarity between the additional activity image capture device output and data samples from each of the plurality of activity clusters to determine whether the additional activity matches an existing cluster sample. In step 614, if the additional activity image capture device output belongs to one of the plurality of activity clusters, the Siamese neural network detects the activity. More specifically, in the illustrative embodiment, after obtaining similarity scores for all pairs of samples that are found to be similar to each other, the samples are drawn close to each other to form clusters. A clustering algorithm (e.g., k-means, Gaussian mixture, model, or other cluster analysis algorithm compatible with this method) is run to give each cluster an identity, i.e., a sequence number or label. After executing the clustering algorithm, a human can optionally manually review the activities contained within each cluster and adjust them, for example, by merging two clusters into one cluster or dividing a cluster into two clusters.

[0036] In step 616, if the similarity measure exceeds the specified measure range for all activity clusters in the plurality of existing activity clusters, a new cluster associated with the current activity output is created. The new activity cluster is then added to the existing plurality of activity clusters. If the additional activity image capture device output belongs to the new activity cluster, the Siamese neural network can detect the activity. In this way, the system uses a spatiotemporal Siamese neural network model to cluster similar activities in a semi-supervised manner without requiring access to a large amount of annotations manually labeled by humans.

[0037] In steps 618 and 620, a Siamese neural network can be used to automatically annotate activities as predicted or non-predicted, creating an annotated predicted dataset and an annotated non-predicted dataset. To create the dataset annotated as predicted activities, in step 618, the Siamese neural network is employed by the annotation algorithm to annotate activity samples according to a defined time period-based logic. In other words, the algorithm annotates image capture device output associated with pedestrian activities captured during a specified time period prior to the detected activity by labeling them as predicted for the detected activity.

[0038] The following is an illustrative embodiment of a method for predicting pedestrian activities before they occur ("pedestrian intention prediction"). The collected dataset has a complete chronological record of each event that occurred sequentially. The trained Siamese neural network model is used to detect the time of, for example, a pedestrian action (such as waving or making a gesture). In this illustrative example, the pedestrian action is detected at time t=25 seconds. A fixed offset from the time of the pedestrian action is provided, for example dt=10 seconds. The annotation time period is specified as, for example, 5 seconds. Based on the fixed offset of dt=10 seconds and the annotation period of 5 seconds, the dataset from t=25-dt-5 to t=25-dt seconds can be annotated as containing predictive information for the target pedestrian activity. The imagery so annotated will be from t=10 to t=15 seconds.

[0039] Return Reference Figure 2 During the time period represented by the bar in the second row from the top, the Siamese neural network model detects the gesture identified in the third image from the left. When the Siamese neural network model makes a positive detection of a gesture, it can annotate that segment as a predicted segment / activity for a fixed time before the start of the detected gesture. Therefore, the "deceleration" segment in the fifth row from the top can be considered a prediction of the gesture because it occurs in the time period before the gesture.

[0040] In step 620, a negative sample of pedestrian activity may be created from the image capture device output in which no pedestrian activity was detected. The negative sample of pedestrian activity is a pedestrian activity captured during a specified time period prior to the period in which one of the plurality of activity clusters was not detected. In this manner, the pedestrian activity is automatically labeled as not predicting one of the plurality of activity clusters. A dataset of pedestrian activities annotated as not predicting an activity is created.

[0041] In step 622, a spatio-temporal intent prediction model is trained using the spatio-temporal data samples from the non-predicted annotated dataset and the spatio-temporal data samples from the predicted annotated dataset as input. The intent prediction model can then be deployed to assign the likelihood of a particular activity based on the predicted and non-predicted information.

[0042] Video of pedestrian activity is input to an intention prediction model. The intention prediction model compares the video with non-predicted and predicted activities to determine whether the pedestrian activity is likely to occur. In an illustrative embodiment, the intention prediction model may assign the likelihood of a particular activity by assigning a "1" when the activity is predicted and a zero when the activity is not predicted. Other methods of assigning the likelihood of a particular activity may also be used in the intention prediction model. When used in an autonomous or semi-autonomous vehicle control unit, automatic vehicle maneuvers can be performed based on the assignment of the likelihood of a particular activity. Additionally or alternatively, an audio or visual warning may be generated to alert the driver about nearby pedestrians. The visual warning may be presented on a display unit within the vehicle to which the driver may be aware.

[0043] Vehicle maneuvers in response to pedestrian activity may include, for example, slowing or stopping the vehicle by applying brakes or redirecting the vehicle by changing the steering angle. Coordination with the navigation system may also be implemented to further automatically guide the vehicle so that safety precautions are taken.

[0044] Figure 7 is a block diagram of an illustrative autonomous or semi-autonomous vehicle and an associated drive system that recognizes pedestrian activity and causes the vehicle to perform actions based on the recognized activity. The vehicle 700 includes a computing device 702 having a neural network 704 as implemented in the steps of an illustrative embodiment of a pedestrian recognition method. Data from a sensor 706 (e.g., an image capture device included on the vehicle 700) is input to the computing device 702. The computing device 702 includes a pedestrian recognition prediction model that acts on the data from the sensor 706 in accordance with the methods disclosed herein. The data from the sensor 706 may include video data that is processed in multiple frames or a single frame. Additional details of the computing device 702 are provided in Figure 6 Shown in.

[0045] Vehicle 700 has various components 708, such as a braking system and a steering system. Each system may have its own electronic control unit. An electronic control unit may also be designed to control more than one vehicle system. One or more actuators 710 are associated with each vehicle component 708. Computing device 702 generates signals based on input from sensors 706. Signals from computing device 702 are input to actuators 710, which provide electronic instructions to act on vehicle components 708. The actuators 710 are associated with the vehicle components 708. For example, actuator 710 may receive a signal from computing device 702 to stop vehicle 700. Actuator 710 will then activate the vehicle's braking system to execute the instruction from computing device 702.

[0046] Figure 7 The illustrative system depicted in FIG. 7 relies on input to a computing device 702 from sensors 706 located on a vehicle 700. In further embodiments, the computing device 702 may receive signals from external sensors, such as cameras attached to infrastructure components. The signals from the external sensors may be processed by the computing device 702 to enable the vehicle 700 to perform a safe maneuver. The onboard navigation system may coordinate with the computing device 702 to implement the maneuver.

[0047] Figure 8 As Figure 7FIG2 is a block diagram of an illustrative embodiment of a computing device 702 as a component of a vehicle system. Computing device 702 includes a memory device 802, which can be a single memory device or multiple devices, for storing executable code to implement any portion of the pedestrian recognition method disclosed herein, including, for example, algorithms for implementing Siamese neural network training and deployment. Further included in memory device 802 may be stored data, such as data representing characteristics of each pedestrian's activity. One or more processors 804 are coupled to memory device 802 via a data interface 806. Processor 804 may be any device(s) configured to execute one or more applications and analyze and process data in accordance with embodiments of the pedestrian recognition method. Processor 804 may be multiple processors acting individually or in concert, or a single processor. Processor 804 may be, for example, a microprocessor, a special-purpose processor, or other device capable of processing and transforming electronic data. Processor 804 executes instructions stored on memory device 802. Memory device 802 may be integrated with processor 804 or a separate device. Illustrative types and characteristics of memory device 802 include volatile and / or non-volatile memory. Various types of memory may be used, as long as the type(s) are compatible with the system and its functionality. Illustrative examples of memory types include, but are not limited to, various types of random access memory, static random access memory, read-only memory, magnetic disk storage devices, optical storage media, and flash memory devices. This description of memory also applies, to some extent, to memory 512, which stores similar database 506 and distinct database 508.

[0048] Input / output devices 808 are coupled to the data interface 806. This may include, for example, image capture devices and actuators. A network interface 810 is also shown coupled to the data interface 806, which may couple the computing device components to a private or public network 812.

[0049] Ground truth data collection can be accomplished using image capture devices positioned to obtain video or images to provide paired data. In an illustrative embodiment, two or more image capture devices (e.g., cameras) are mounted alongside a highway, each with a different viewing angle but with overlapping focal regions. The overlapping regions of the resulting images can be used to train a Siamese neural network based on image similarities across a pair or a collection of images from the image capture devices. A pair or group of images captured at different times and not covering the same activity can be used for training based on a dissimilarity metric.

[0050] The image capture devices may be mounted, for example, on infrastructure poles or other support structures at traffic intersections.The image capture devices may also be mounted on a single vehicle at different mounting locations but with overlapping areas of interest.

[0051] Siamese neural networks can be trained based on recognizing various types of activities (e.g., gestures) by obtaining image data from locations other than vehicles or roadways. For example, image capture devices can be installed at locations where social or commercial gatherings occur, such as restaurants, cafeterias, banks, or other establishments where people gather. Image capture devices can be mounted on a wall, for example, with multiple possible pairs of cameras, where they capture images of people communicating with each other using gestures or other movements.

[0052] Already available or collected annotations (even in small quantities) can be used by unsupervised models in a more effective manner than traditional supervised models. The existing annotations allow the model to be trained on most or all possible combinations and permutations of input data pairs or groupings extracted from this existing small annotated dataset, while still learning the ability to predict whether the input pairs or groupings are similar or dissimilar.

[0053] The illustrative embodiments include a non-transitory computer-readable medium on which computer code is stored, which, when executed on one or more processors, causes a computer system to perform the method for pedestrian activity recognition as described herein. The term "computer-readable medium" may be, for example, a machine-readable medium that can store data in a format readable by a mechanical device. Examples of computer-readable media include, for example, semiconductor memories (e.g., flash memories, solid-state drives, SRAM, DRAM, EPROM, or EEPROM), magnetic media (e.g., magnetic disks, optical disks), or other forms of computer-readable media that can be functionally implemented to store code and data for the execution of embodiments of the pedestrian activity recognition method described herein.

[0054] Advantageously, the illustrative embodiments of pedestrian activity recognition and prediction do not rely on extensive annotations, as conventional methods do. As those skilled in the art will appreciate, pedestrian activity recognition and prediction are particularly challenging problems due to the amount of computational power and accuracy required. Importantly, the classification or categorization of the present embodiments focuses on detecting similarities and dissimilarities between multiple different activities. This metric learning approach can require orders of magnitude less annotated data to learn new activities or categories than conventional methods. Clustering activities in this unsupervised manner using two or more image capture devices to obtain paired data can reduce processing time. Furthermore, the disclosed illustrative embodiments of pedestrian recognition allow for the detection of unknown new activities as well as the recognition of known activities. This allows the system to operate with little or no manually labeled data. Using multiple image capture devices to concurrently record pedestrian activities from different viewpoints also offers advantages over conventional methods by enabling the system to automatically capture both grouped or paired data streams and unpaired data to generate similar and dissimilar datasets, both used to train the pedestrian recognition system.

[0055] Various types of neural networks can be used in the illustrative embodiments, as long as they can be trained and deployed to recognize pedestrian activity. In the illustrative embodiments, each neural network can be a convolutional neural network with shared parameters. During feature extraction, similarity values ​​derived from the extraction of comparable hidden layer features (feature extraction) are calculated and output. Convolutional neural networks can cover the temporal portion of spatiotemporal neural networks. Convolutional neural networks are particularly suitable for detecting features in images. Convolutional neural networks move filters across an image and use convolution operations to calculate values ​​associated with the filters. Filters can be associated with any feature found in an image and can represent aspects of an activity or person that the system wishes to identify. For example, a filter can be associated with whether a person is identified as running, for example, based on the position of their legs or the tilt of their body. Filters can be assigned specific values, which are then updated automatically during neural network training operations. Once the filters are passed over an image, a feature map is generated. Multiple filter layers can be used, which will generate additional feature maps. The filters provide translation invariance and parameter sharing. Pooling layers can be included to identify inputs for use in subsequent layers. The process can include multiple convolutional layers, each followed by a pooling layer. A fully connected layer may also be included before the classification output of the convolutional neural network. Nonlinear layers (such as rectified nonlinear units) may be implemented between convolutional layers to improve the robustness of the neural network. In summary, the input is fed into a convolutional layer, which may be followed by a nonlinear layer, one or more additional convolutional layers and nonlinear layers may follow, and this sequence can continue until a fully connected layer is reached before a pooling layer is provided.

[0056] Neural network training according to the illustrative embodiments can be end-to-end for learning of spatiotemporal features.Alternatively or additionally, feature extraction can be used in conjunction with separate classification.

[0057] In illustrative embodiments of the pedestrian recognition method and system, a single neural network can capture both spatial and temporal information, or a network can capture combined spatial-temporal information.

[0058] Various embodiments of the present invention have been described, each having different combinations of elements. The present invention is not limited to the specific embodiments disclosed, but may include different combinations of disclosed elements, omissions of some elements, or replacement of elements by structural equivalents of such elements.

[0059] It is further noted that while the embodiments are described primarily with respect to pedestrian activity near motor vehicles, the methods and systems are applicable to human or animal activity in other contexts. In general, the disclosed embodiments are applicable to recurring activities that can be recognized by a trained neural network and used to generate motions for a vehicle or other device.

Claims

1. A method for pedestrian activity recognition, comprising: training a Siamese neural network to recognize multiple activities by training it based on two or more inputs, wherein the inputs are recordings of the same pedestrian activity from two or more separate training image capture devices; deploying the Siamese neural network model with continuous data collection from an additional activity image capture device to create a dataset of multiple activity clusters of similar activities in an unsupervised manner; employing the Siamese neural network to annotate activities as predicted activities or non-predicted activities to create an annotated prediction dataset and an annotated non-predictive dataset; training a spatio-temporal intent prediction model using the spatio-temporal data samples from the annotated non-prediction dataset and the spatio-temporal data samples from the annotated prediction dataset as input; and The intent prediction model is deployed to assign the likelihood of a particular activity.

2. The method for pedestrian activity recognition according to claim 1, wherein: Deploying the Siamese neural network model to cluster similar activities in an unsupervised manner using continuous data collection from the additional activity image capture device includes: Inputting into the Siamese neural network comprising the plurality of activity clusters: output from said additional motion image capture device; and said dataset of said plurality of activity clusters; determining, by the Siamese neural network, a measure of similarity between the additional activity image capture device output and data samples of each of the plurality of activity clusters to determine whether the additional activity matches an existing cluster sample; and If the additional motion image capturing device outputs an activity belonging to one of the plurality of activity clusters, then activity is detected.

3. The method for pedestrian activity recognition according to claim 2, further comprising: If the measure of similarity exceeds a specified measure range for all activity clusters in the plurality of activity clusters, a new cluster associated with the current activity output is created and added to the plurality of activity clusters.

4. The method for pedestrian activity recognition according to claim 1, wherein: Training the Siamese neural network to recognize the plurality of activities comprises: creating a similar dataset from at least an output from a first training image capture device of the two or more training image capture devices and a synchronized output from a second training image capture device of the two or more training image capture devices, wherein the outputs reflect the same pedestrian activity; creating distinct data sets from the output from the first training image capture device and the delayed output from the second training image capture device, wherein the output reflects different pedestrian activities; and A dataset including a plurality of activity clusters is created by training a Siamese neural network based on the similar dataset and the dissimilar dataset.

5. The method for pedestrian activity recognition as claimed in claim 4 further comprises refining the similar dataset and the dissimilar dataset by applying rule-based heuristics to the delayed output from the second training image capture device and the output of the first training image capture device to evaluate dissimilarity.

6. The method for pedestrian activity recognition according to claim 1, wherein: Employing the Siamese neural network to annotate activities as predicted activities or non-predicted activities to create annotated prediction datasets and annotated non-predictive datasets includes: executing an annotation algorithm employing the Siamese neural network to automatically annotate image capture device output associated with activity captured during a specified time period prior to a detected activity to label it as predictive of the detected activity to create a dataset annotated as predicted activity; and Negative samples of pedestrian activities captured during a specified time period preceding a period in which one of the plurality of activity clusters was not detected are created from image capture device output in which no activity was detected, and the negative samples are labeled as not predicting the one of the plurality of activity clusters to create a dataset annotated as not predicting activity.

7. The method for pedestrian activity recognition according to claim 1, wherein: Deploying the intent prediction model to assign the likelihood of a particular activity includes assigning a "1" when an activity is predicted and assigning a zero when an activity is not predicted.

8. The method of pedestrian activity recognition of claim 1, further comprising performing automatic vehicle maneuvers based on the assignment of the likelihood of the particular activity.

9. A system for identifying pedestrian activities, comprising: one or more processors; one or more storage devices having computer code stored thereon, the computer code comprising a Siamese neural network; Execution of the computer code causes the one or more processors to perform the following method: training the Siamese neural network to recognize a plurality of activities based on two or more inputs, wherein the inputs are recordings of the same pedestrian activity from two or more separate training image capture devices; deploying the Siamese neural network model with continuous data collection from an additional activity image capture device to create a dataset of multiple activity clusters of similar activities in an unsupervised manner; employing the Siamese neural network to annotate activities as predicted activities or non-predicted activities to create an annotated prediction dataset and an annotated non-predictive dataset; training a spatio-temporal intent prediction model using the spatio-temporal data samples from the annotated non-prediction dataset and the spatio-temporal data samples from the annotated prediction dataset as input; and The intent prediction model is deployed to assign the likelihood of a particular activity.

10. The system of claim 9, wherein: Deploying the Siamese neural network model to cluster similar activities in an unsupervised manner using continuous data collection from the additional activity image capture device includes the computer code causing the one or more processors to perform the steps of: Inputting into the Siamese neural network comprising the plurality of activity clusters: output from said additional motion image capture device; and said dataset of said plurality of activity clusters; determining, by the Siamese neural network, a measure of similarity between the additional activity image capture device output and data samples of each of the plurality of activity clusters to determine whether the additional activity matches an existing cluster sample; as well as If the additional motion image capturing device outputs an activity belonging to one of the plurality of activity clusters, then activity is detected.

11. The system of claim 10 , further comprising the one or more storage devices having computer code stored thereon, the computer code, when executed, causing the one or more processors to perform the following steps: if the measure of similarity exceeds a specified measure range for all activity clusters in the plurality of activity clusters, creating a new cluster associated with the current activity output and adding the new cluster to the plurality of activity clusters.

12. The system of claim 9, wherein: Training the Siamese neural network to recognize the plurality of activities includes the computer code causing the one or more processors to perform the following steps: creating a similar dataset from at least an output from a first training image capture device of the two or more training image capture devices and a synchronized output from a second training image capture device of the two or more training image capture devices, wherein the outputs reflect the same pedestrian activity; creating distinct datasets from the output from the first training image capture device and the delayed output from the second training image capture device, wherein the output reflects different pedestrian activities; as well as A dataset including a plurality of activity clusters is created by training a Siamese neural network based on the similar dataset and the dissimilar dataset.

13. The system of claim 12, further comprising the one or more storage devices having computer code stored thereon, the computer code, when executed, causing the one or more processors to perform the following steps: refining the similar data set and the dissimilar data set by applying rule-based heuristics to the delayed output from the second training image capture device and the output of the first training image capture device to evaluate dissimilarity.

14. The system of claim 9, wherein: Employing the Siamese neural network to annotate activities as predicted activities or non-predicted activities to create annotated prediction datasets and annotated non-predictive datasets includes the computer code causing the one or more processors to perform the following steps: executing an annotation algorithm employing the Siamese neural network to automatically annotate image capture device output associated with activity captured during a specified time period prior to a detected activity to label it as predictive of the detected activity to create a dataset annotated as predicted activity; as well as Negative samples of pedestrian activities captured during a specified time period preceding a period in which one of the plurality of activity clusters was not detected are created from image capture device output in which no activity was detected, and the negative samples are labeled as not predicting the one of the plurality of activity clusters to create a dataset annotated as not predicting activity.

15. The system of claim 9, wherein: Deploying the intent prediction model to assign a likelihood of a particular activity includes the computer code causing the one or more processors to perform the steps of assigning a "1" when an activity is predicted and assigning a zero when an activity is not predicted.

16. The system of claim 9, further comprising the one or more storage devices having computer code stored thereon that, when executed, causes automatic vehicle maneuvering based on the assignment of the likelihood of the particular activity.

17. An autonomous or semi-autonomous controlled vehicle, wherein the vehicle comprises the system for identifying pedestrian activities according to claim 9.

18. The vehicle of claim 17, further comprising: vehicle control components; as well as an actuator electronically connected to the vehicle control assembly; wherein the one or more storage devices have stored thereon computer code that, when executed, causes the actuator to initiate the vehicle maneuver via the vehicle control assembly.

19. A non-transitory computer readable medium having computer code stored thereon, the computer code, when executed on one or more processors, causing a computer system to perform the following method: training a Siamese neural network to recognize multiple activities by training it based on two or more inputs, wherein the inputs are recordings of the same pedestrian activity from two or more separate training image capture devices; deploying the Siamese neural network model with continuous data collection from an additional activity image capture device to create a dataset of multiple activity clusters of similar activities in an unsupervised manner; employing the Siamese neural network to annotate activities as predicted activities or non-predicted activities to create an annotated prediction dataset and an annotated non-predictive dataset; training a spatio-temporal intent prediction model using the spatio-temporal data samples from the annotated non-prediction dataset and the spatio-temporal data samples from the annotated prediction dataset as input; and The intent prediction model is deployed to assign the likelihood of a particular activity.

20. The non-transitory computer readable medium of claim 19, wherein: Deploying the Siamese neural network model to cluster similar activities in an unsupervised manner using continuous data collection from the additional activity image capture device includes: Inputting into the Siamese neural network comprising the plurality of activity clusters: output from said additional motion image capture device; and said dataset of said plurality of activity clusters; determining, by the Siamese neural network, a measure of similarity between the additional activity image capture device output and data samples of each of the plurality of activity clusters to determine whether the additional activity matches an existing cluster sample; and If the additional motion image capturing device outputs an activity belonging to one of the plurality of activity clusters, then activity is detected.

21. The non-transitory computer-readable medium of claim 20, further comprising: If the measure of similarity exceeds a specified measure range for all activity clusters in the plurality of activity clusters, a new cluster associated with the current activity output is created and added to the plurality of activity clusters.

22. The non-transitory computer readable medium of claim 19, wherein: Training the Siamese neural network to recognize the plurality of activities comprises: creating a similar dataset from at least an output from a first training image capture device of the two or more training image capture devices and a synchronized output from a second training image capture device of the two or more training image capture devices, wherein the outputs reflect the same pedestrian activity; creating distinct data sets from the output from the first training image capture device and the delayed output from the second training image capture device, wherein the output reflects different pedestrian activities; and A dataset including a plurality of activity clusters is created by training a Siamese neural network based on the similar dataset and the dissimilar dataset.

23. The non-transitory computer-readable medium of claim 22, further comprising refining the similar dataset and the dissimilar dataset by applying rule-based heuristics to the delayed output from the second training image capture device and the output of the first training image capture device to evaluate dissimilarity.

24. The non-transitory computer readable medium of claim 19, wherein: Employing the Siamese neural network to annotate activities as predicted activities or non-predicted activities to create annotated prediction datasets and annotated non-predictive datasets includes: executing an annotation algorithm employing the Siamese neural network to automatically annotate image capture device output associated with activity captured during a specified time period prior to a detected activity to label it as predictive of the detected activity to create a dataset annotated as predicted activity; and Negative samples of pedestrian activities captured during a specified time period preceding a period in which one of the plurality of activity clusters was not detected are created from image capture device output in which no activity was detected, and the negative samples are labeled as not predicting the one of the plurality of activity clusters to create a dataset annotated as not predicting activity.

25. The non-transitory computer readable medium of claim 19, wherein: Deploying the intent prediction model to assign the likelihood of a particular activity includes assigning a "1" when an activity is predicted and assigning a zero when an activity is not predicted.

26. The non-transitory computer-readable medium of claim 19, further comprising performing an automated vehicle maneuver based on the assignment of the likelihood of the particular activity.