Systems and methods for predicting pedestrian intent
By employing computer vision to analyze human images and map detected poses to intentions, the system addresses the limitations of conventional pedestrian intention prediction methods, providing more accurate and culturally aware predictions for autonomous vehicles.
Patent Information
- Application Number
- JP2020552164
- Authority / Receiving Office
- JP · JP
- Patent Type
- Patents
- Current Assignee / Owner
- Priority Date
- 2017-12-13
- Filing Date
- 2018-12-13
- Publication Date
- 2025-05-19
- Estimated Expiration
- 2038-12-13
AI Technical Summary
Conventional systems for determining pedestrian intentions in autonomous vehicles rely on crude methods that fail to consider human body language and cultural gestures, resulting in inaccurate intention predictions.
The system uses computer vision to analyze a series of human images, detect keypoints to determine human poses, and map these poses to predicted intentions using a database that accounts for cultural variations and geographical differences.
This approach enables more accurate and culturally sensitive prediction of pedestrian intentions, improving the safety and effectiveness of autonomous vehicle operations.
Smart Images

Figure 0007679197000001 
Figure 0007679197000002 
Figure 0007679197000003
Abstract
Description
Technical Field
[0001] This disclosure generally relates to the field of automation, or computer vision for guiding autonomous vehicles, and more specifically to applying computer vision to predict the intentions of pedestrians.
Background Art
[0002] This application claims the benefit of U.S. Provisional Application No. 62 / 598,359, filed Dec. 13, 2017, which is hereby incorporated by reference in its entirety.
[0003] Related conventional systems for determining pedestrian intentions use crude and primitive means for performing intention determination, such as analyzing the direction or speed in which a pedestrian is moving, and analyzing the distance of the pedestrian from the edge of the road. Conventional systems take these variables into a static model and roughly determine the most likely intentions of the user. This approach is overly broad and a uniform approach that fails to access other things such as human body language in the form of explicit or implicit gestures, and human perception of the surroundings of the individual. From a technical perspective, conventional systems lack the technical sophistication to understand human poses based on a series of human images, and thus cannot convert gestures or recognition information into more accurate intention predictions.
Brief Description of the Drawings
[0004] The disclosed embodiments have other advantages and features that will become more readily apparent from the detailed description, the appended claims, and the accompanying drawings (or figures). A brief introduction to the figures is shown below.
[0005]
Figure 1
Figure 2
Figure 3
Figure 4
Figure 5
Figure 6
Figure 7
Figure 8
Figure 9
Figure 10
Figure 11
[0006] The drawings (figures), and the following description, relate to preferred embodiments by way of example only. It should be noted that from the following discussion, alternative embodiments of the structures and methods disclosed herein will be readily recognized as viable alternatives that may be used without departing from the principles claimed.
[0007] Next, some embodiments will now be referred to in detail, examples of which are illustrated in the accompanying drawings. It should be noted that whenever possible, the same or similar reference numbers may be used in the figures and may indicate the same or similar functions. The figures are for the purpose of illustration only and depict embodiments of the disclosed system (or method). Those skilled in the art will readily recognize from the following description that alternative embodiments of the structures and methods illustrated herein may be used without departing from the principles of the invention described herein.
[0008] (Overview of the Configuration) One embodiment of the disclosed systems, methods, and computer-readable storage media includes determining a human intention by converting a series of human images (e.g., from a video) into an accumulation of keypoints and determining the human pose. The determined pose is then mapped to a human intention based on a mapping from known poses to intentions. These mappings from poses to intentions may vary depending on various gestures that teach poses that may represent different intentions depending on human culture, and the geography of the location where the video is captured, which may vary depending on habits. For this and other purposes, in some embodiments of the present disclosure, the processor obtains a plurality of series of images from feeds provided by vision and / or depth sensors such as cameras (e.g., cameras providing video feeds), far-infrared sensors, and LIDAR. In situations where the camera provides a video feed, the camera can be mounted on or near the dashboard of an autonomous vehicle, or anywhere on or in the vehicle (e.g., incorporated into the vehicle body structure). The processor determines, for each image of the plurality of series of images, the respective keypoints corresponding to a human, such as points corresponding to the human's head, arms, and legs, and their relative positions to each other, as the images change from image to image in the series of images.
[0009] The processor aggregates respective keypoints for respective images into a human pose. For example, based on the movement of a human leg across a series of images as determined from keypoint analysis, the processor aggregates the keypoints into a vector indicating the movement of the keypoints across the images. Next, the processor transmits a query to a database to determine whether it is known to map to a given intention, such as an intention to cross a street, and compares it with a plurality of template poses that convert candidate poses of the intention into an intention. The processor receives a response message from the database indicating either the human intention or the inability to identify the location of the matching template. In response to the response message indicating the human intention, the processor outputs a command corresponding to the intention (e.g., to stop the vehicle or to warn the driver of the vehicle).
[0010] (Overview of the system) FIG. 1 illustrates an embodiment of a system for determining human intention using an intention determination service according to some embodiments of the present disclosure. The system 100 includes a vehicle 110. While the vehicle 110 is depicted as an automobile, the vehicle 110 may be any electric device configured to move automatically or semi-automatically near a human. As used herein, the term "automatic" and variations thereof refer not to fully automated operations or other functions (e.g., "semi-automatic"), but to semi-automatic operations that rely on human input to operate some functions. For example, the vehicle 110 may be a drone or a bipedal robot configured to fly or walk near a human without being commanded (or while being commanded to perform at least some functions) to navigate in any direction or at any speed. The vehicle 110 is configured to operate safely around a human, determine (or infer) the intention of a nearby human such as the human 114, and take a safely coordinated action considering the intention of the human 114.
[0011] Camera 112 is operably coupled to vehicle 110 and is used to determine the intention of human 114. Other sensors 113, such as a microphone, can be operably coupled to vehicle 110 to supplement the intention determination of human 114 (for example, when human 114 says "I intend to cross the street now", vehicle 110 can activate a safety mode to decelerate or stop the vehicle based on the microphone input so that vehicle 110 can avoid a collision). As used herein, the term "operably coupled" refers to direct attachment (for example, wiring to or incorporating into the same circuit board), indirect attachment (for example, two components connected through a wire or a wireless communication protocol), and the like. A series of images are captured by camera 112 and provided to intention determination service 130 via network 120. The technical aspects of network 120 will be described in more detail with respect to FIG. 2. In the following disclosure, intention determination service 130 determines the intention of human 114, as will be described in more detail with respect to other figures such as FIG. 3. While network 120 and intention determination service 130 are depicted as being remote from vehicle 110, network 120 and / or intention determination service 130 can be implemented wholly or partially within vehicle 110.
[0012] (Configuration of Computer Equipment) FIG. 2 is a block diagram illustrating exemplary components of a machine capable of reading instructions from a machine-readable medium and executing them in a processor (or controller). Specifically, FIG. 2 illustrates a schematic representation of a machine in an exemplary form of a computer system 200 that has program code (e.g., software) therein for causing a machine to execute any one or more of the methods described herein. The program code may be comprised of instructions 224 executable by one or more processors 202. In alternative embodiments, the machine may operate as a stand-alone device or may be connected (e.g., networked) to other machines. In a networked deployment, the machine may operate in the capacity of a server machine, a client machine, in a server-client network environment, or as a peer machine in a peer-to-peer (or distributed) network environment.
[0013] The machine may be a server computer, a client computer, a personal computer (PC), a tablet PC, a set-top box (STB), a personal digital assistant (PDA), a cellular telephone, a smartphone, a web appliance, a network router, switch or bridge, an in-vehicle computer (e.g., a computer embedded in vehicle 110 for operating vehicle 110), an infrastructure computer, or any machine capable of executing (a series of or other) instructions 224 to specify actions to be taken by that machine. Further, while only a single machine is illustrated, the term "machine" shall also be taken to include any collection of machines that individually or jointly execute instructions 124 to perform any one or more of the methods described herein.
[0014] The exemplary computer system 200 includes a processor 202 (e.g., a central processing unit (CPU), a graphics processing unit (GPU), a digital signal processor (DSP), one or more application specific integrated circuits (ASICs), one or more radio frequency integrated circuits (RFICs), or any combination thereof), a main memory 204, and a static memory 206, and is configured to communicate with each other via a bus 208. The computer system 200 may further include a visual display interface 210. The visual interface may include a software driver that enables the user interface to be displayed on a screen (or display). The visual interface can display the user interface directly (e.g., on the screen) or indirectly (e.g., via a visual projection unit) on a surface or window. For ease of explanation, the visual interface may be described as a screen. The visual interface 210 may include a touch screen or may interface with a touch screen. Also, the computer system 200 includes an alphanumeric input device 212 (e.g., a keyboard or a touch screen keyboard), a cursor control device 214 (e.g., a mouse, a trackball, a joystick, a human sensor, or other pointing device), a storage unit 216, a signal generation device 218 (e.g., a speaker), and a network interface device 220, and is configured to communicate via the bus 208. The signal generation device 218 can output to the visual interface 210 or output signals to other devices such as a vibration generator and a robotic arm, and warn the human driver of the vehicle 110 of danger to pedestrians (e.g., attract the driver's attention by vibrating the driver's seat, display a warning, and when the processor 202 determines that the driver has fallen asleep, gently tap the driver using the robotic arm, etc.).
[0015] The memory unit 216 includes a machine-readable medium 222 that stores instructions 224 (e.g., software) that implement any one or more of the methods or functions described herein. Also, the instructions 224 (e.g., software) may be present in whole or at least partially within the main memory 204 or within the processor 202 (e.g., within the processor's cache memory) during its execution by the computer system 200, and the main memory 204 and the processor 202 also constitute a machine-readable medium. The instructions 224 (e.g., software) can be transmitted or received over the network 226 via the network interface device 220.
[0016] While the machine-readable medium 222 is shown as a single medium in the exemplary embodiment, the term "machine-readable medium" is to be understood to include a single medium or multiple media (e.g., a centralized or distributed database, or associated caches and servers) that can store instructions (e.g., instructions 224). Also, the term "machine-readable medium" can include any medium that is capable of storing instructions (e.g., instructions 224) for execution by a machine, and it is to be understood that it causes the machine to execute any one or more of the methods disclosed herein. The term "machine-readable medium" includes, but is not limited to, data repositories in the form of solid-state memory, optical media, and magnetic media.
[0017] (Configuration of the Intent Judgment Service) As described above, the systems and methods described herein generally are directed to determining human intent (e.g., human 114) based on human poses as determined from images of cameras mounted or integrated in an autonomous vehicle (e.g., camera 112 of vehicle 110). The images are captured by a processor (e.g., processor 202) of the intent determination service 130. FIG. 3 illustrates an embodiment of an intent determination service that includes modules and databases that support the service, according to some embodiments of the present disclosure. The intent determination service 330 is a detailed view of the intent determination service 130 and has the same functionality as described herein. The intent determination service 330 may be a single server that includes the various modules and databases depicted therein, or may be distributed across multiple servers and databases. As described above, the intent determination service 330 may be implemented, in whole or in part, within the vehicle 110. The distributed servers and databases may be first party or third party services and databases accessible by the network 120.
[0018] The intention determination service 330 includes a keypoint determination module 337. After capturing an image from the camera 112, the processor 202 of the intention determination service 330 executes a human detection module 333. The human detection module 333 determines whether a human is in the image by using computer vision and / or object recognition techniques in analyzing the composition of the image, and determines whether edges or other features in the image match known human forms. If it is determined that there is a human in the image, the processor 202 of the intention determination service 330 executes the keypoint determination module 337 to determine the human keypoints. This determination will be described with respect to FIG. 4. FIG. 4 depicts keypoints for the detection of poses associated with various human activities according to some embodiments of the present disclosure. The human 410 is standing in a neutral position without making a gesture. The keypoints are indicated by circles superimposed on the human 410. Lines are used to illustrate pose features and will be described in more detail below.
[0019] The key point determination module 337 determines where the key points are on the human 410. The key point determination module 337, for example, first identifies the outline of the human body, determines where the key points are, and then matches the predetermined points of the human body as key points based on the outline or other methods of key point / pose determination. For example, the key point determination module 337 may refer to a template of a human image that indicates the key points applied to the ankles, knees, thighs, center of the torso, wrists, elbows, shoulders, neck, and forehead of each human. The key point determination module 337 can identify each of these points by comparing the human image with the template of the human image, and the template indicates the features corresponding to each region where the key points are present. Alternatively, the key point determination module 337 can be configured to distinguish and identify each body part to which the key points are applied using computer vision, identify the positions of these body parts, and apply the key points. As depicted by humans 420, 430, 440, 450, and 460, the key point determination module 337 is configured to identify the key points of a human regardless of the different poses of the human. The behavior of the human pose is considered and will be described below with reference to other modules in FIG. 3.
[0020] After determining which keypoints are used for each image in which a human (e.g., human 114) exists, the processor 202 of the intention determination service 330 executes the keypoint aggregation module 331. The keypoint aggregation module 331 can aggregate keypoints for a single image or a series of consecutive images. Aggregating keypoints for a single image creates a representation of the pose at a specific point in time. Aggregating keypoints for a series of consecutive images creates a vector of keypoint movement over time that can be mapped to known movements or gestures. When the keypoint aggregation module 331 aggregates keypoints for a single image, referring back to FIG. 4, the keypoint aggregation module 331 can aggregate the keypoints by mapping the keypoints using straight connectors.
[0021] The keypoint aggregation module 331 determines which keypoints to connect in such a way based on a predefined schema. For example, as described above, the keypoint module 331 knows that each keypoint corresponds to a specific body part (e.g., ankle, knee, thigh, etc.). The keypoint aggregation module 331 accesses the predefined schema and determines that the ankle keypoint is connected to the knee keypoint and connected to the thigh keypoint, etc. These straight connectors of the keypoints are represented for each of the humans 410, 420, 430, 440, 450, and 460 depicted in FIG. 4. Instead of, or in addition to, remembering or generating the straight connectors, the keypoint aggregation module 331 can remember the distances between each keypoint of the predefined schema. As will be explained later, the relative positions of these lines, or the distances between the keypoints, are used to determine the user's pose. For example, the distance between the hand and the head of human 450 is much shorter than the distance between the hand and the head of human 410, which will be used to determine that it is highly likely that human 450 is holding a phone up to their head with respect to human 450.
[0022] When the keypoint aggregation module 331 aggregates the keypoints of a series of images, the keypoint aggregation module 331 aggregates the keypoints for each of the individual images and then calculates the motion vectors for each of the keypoints from each of the respective images to each of the next series of respective images. For example, in a situation where the first image includes a human 410 and the next image includes a human 410, which is the same human but in a different pose, the keypoint aggregation module 331 calculates the motion vector and depicts the transformation of the movement of the left foot and knee of the human between the two images. This motion vector is illustrated by the lower line of the human 430, which indicates the distance between where the keypoint of the left foot was previously and where it has moved.
[0023] Also, the keypoint aggregation module 331 can indicate the direction in which the user is moving or facing based on the movement of the keypoints between the images, with reference to other objects in the images, and based on the positions of the keypoints. FIG. 5 depicts keypoints for detecting the same pose regardless of the direction in which a human is facing the camera according to some embodiments of the present disclosure. When looking at the images in a two-dimensional space, the keypoint aggregation module 331 can determine where the direction in which the human is facing is based on the proximity of the keypoints. For example, a human 114 is depicted in FIG. 5 and is in various rotational positions such as position 510, position 520, position 530, position 540, and position 550. In image 540, compared to image 510, the keypoints of the shoulders and neck are closer to each other. These relative distances between the keypoints and the changes in these relative distances between the keypoints between the images are analyzed by the keypoint aggregation module 331 to determine the rotation of the human 114 in three-dimensional space. The keypoint aggregation module 331 can detect the direction in which the human 114 is rotating in determining the rotation, and in combination with understanding other objects (e.g., roads, curbs) in the image, based on this information and the aggregation of the keypoints based on the rotation, send it to the intention determination module 332 to be able to determine the intention of the user.
[0024] After aggregating the key points, the processor 202 of the intent determination service 330 executes the intent determination module 332 and optionally feeds the aggregated key points along with additional information. The intent determination module 332 uses the aggregated key points along with potential additional information to determine the intent of the human 114. In some embodiments, the intent determination module 332 classifies the aggregated key points as poses. As used herein, a pose may be a stationary human at a given point in time, or a human movement such as a gesture. Thus, the classified pose may be an accumulation of stationary poses or, as described above, a vector that describes the movement between key points from image to image.
[0025] The intent determination module 332 queries the pose template database 334 to determine the intent of the human 114. The pose template database 334 includes entries for aggregations of various key points and motion vectors between the aggregations. These entries are mapped to corresponding poses. For example, referring back to FIG. 4, the template database 334 may include a record of an aggregation of key points that matches the aggregation of key points for the human 450 and may refer to a pose corresponding to "on the phone". Any number of corresponding poses can be mapped to aggregations of key points and motion vectors, such as standing, sitting, riding a bicycle, running, on the phone, holding hands with another human, etc. Further, an aggregation of key points may match more than one record. For example, an aggregation of key points may match both a template corresponding to running and a template corresponding to being on the phone.
[0026] The pose template database 334 maps the aggregation of key points to the movements of the human 114, or the characteristics of the human 114 that transcend the pose, and can also include the state of mind. FIG. 6 depicts key points for detecting whether the user is attentive and detecting the type of distraction according to some embodiments of the present disclosure. The pose template database 334 maps, for example, the aggregation of key points for the face of the human 114 based on whether the user is distracted or paying attention. If the human is distracted, the pose template database 334 can map the aggregation of key points by adding them for distraction. For example, the pose template database 334 may include a record 610 that maps the aggregation of key points to an attentive human 114. The pose template 334 also includes a record 620, a record 630, and a record 640, each of which can be mapped to a human who is distracted, looking sideways, on the phone, or looking downward and thus distracted. When receiving a matching template from the pose template 334 in response to a query, the intention determination module 332 can determine that the pose of the human 114 matches one or more poses and the state of mind.
[0027] Depending on where a pose is created, there are situations where the same pose can mean two different things. For example, if human 114 raises his hand in Europe, the normal cultural response is to interpret this as a "stop" signal. In a country where this gesture usually means "continue", if human 114 raises his hand, the intention judgment module 332, based on the records in the pose template database 334 that show the correspondence between the key points and such meanings, will conclude that it is incorrect to judge this as a stop gesture. Thus, in some embodiments, the pose template database 334 simply details what the pose is, for example, that a human hand is being waved, and another database, the cultural pose intention database 336, converts that pose into the human's desire or intention depending on the human's geographical location. This conversion will be described with respect to FIG. 7, and some embodiments of the present disclosure depict a geographical map based on gesture differences.
[0028] Using the above system and method, the intention determination module 332 determines, based on the entries in the pose template database 334, that human 710 is raising his hand and that human 720 is holding his hand horizontally. After making this determination, the intention determination module 332 refers to the poses of the determined humans 710 and 720 for the entries in the cultural pose intention database 336. The cultural pose intention database 336 maps the pose to an intention based on the pose being determined and the location from the acquired image. For example, the image of human 710 was acquired in Europe, while the image of human 720 was acquired in Japan. Each entry in the cultural pose intention database 336 maps the pose to the corresponding cultural meaning at the location where the image was acquired. Thus, the meaning of the pose of human 710 will be determined based on the entry corresponding to the location in Europe, whereas the meaning of the pose of human 720 will be determined based on the entry corresponding to the location in Japan. In some embodiments, the cultural pose intention database 336 is distributed among several databases, each database corresponding to a different location. The appropriate database can be referenced based on the geographical location where the image was acquired. In a situation where the intention determination service 330 is partially or fully installed within the vehicle 110, the processor 202 can automatically download data from the cultural pose intention database 336 in response to detecting that the vehicle 110 has entered a new geographical location (e.g., when the vehicle 110 crosses a border and enters a different city, state, country, principality, etc.). This improves the efficiency of the system by avoiding latency in requesting cultural pose intention data via an external network (e.g., network 120).
[0029] After determining the pose of human 114, arbitrarily determine other information such as the cultural meaning of the pose and whether human 114 is distracted in the above-described method, and the intention determination module 332 can predict the intention of human 114. FIG. 8 depicts a flowchart for mapping some exemplary poses to predicted intentions according to some embodiments of the present disclosure. In practice, the data flow shown in FIG. 8 does not need to start every time a human is detected by the human detection module 333. For example, if a human is detected by the human detection module 333 but the human is determined to be far from the vehicle 110 or far along the path the vehicle 110 is traveling, the intention of the human is not important for the progress of the vehicle 110 and, as a result, does not need to be determined. Thus, in some embodiments, the processor 202 of the intention determination service 330 can determine an initial prediction as to whether the activity of human 114 will result in the vehicle 110 needing to change its progress along the path or can affect the safety of human 114 when the human detection module 333 detects human 114.
[0030] When the intention determination service 330 determines an initial prediction that the activity of human 114 may be important for the progress of the vehicle 110, as described above, the intention determination module 332 determines the pose of human 114 and, in some embodiments, its corresponding meaning. Thus, as shown in data flow 810, the intention determination module 332 determines that human 114 has raised the left arm. The intention determination module 332 then determines the predicted intention of human 114 therefrom. In some embodiments, the pose template database 334 shows the intention that matches the pose as described above, and thus the intention determination module 332 determines the intention of human 114, such as waiting for the human to cross, as shown in data flow 810.
[0031] The pose intention mapping in the pose template database 334 may be supplemented by information from additional sources. For example, the intention determination module 332 can use deep learning to convert a pose (and potential additional information such as the distance from the human 114 to the road) into the intention of the human 114. For example, given a dataset of classified pedestrian intentions (such as crossing and not crossing) for training a deep learning framework, the intention determination module 332 can predict data from the next image in a series of real-time acquired images as to whether the human 114 intends to cross or not cross and with what confidence.
[0032] Furthermore, the intention determination module 332 performs statistical analysis on the movements, body language, interactions with the social infrastructure, activities, and actions of the human 114, and provides different inputs such as the previous actions of the human 114, the classification of the human 114 by type (e.g., a single pedestrian, a group, a cyclist, a disabled person, a child, etc.), the speed of the pedestrian, the context of the interaction (e.g., crosswalk, middle of the street, city, countryside, etc.) to determine the future actions of the human 114. Additionally, the intention determination module 332 can apply a psychological behavior model to improve intention prediction. The psychological behavior model is related to the ability of a bottom-up, data-driven approach such as the above-described embodiments that are complemented or extended using a psychological model for novel actions of pedestrian intentions from pedestrian movements, body language, and computer vision of actions. Moreover, the intention determination module 332 can improve its prediction using sensors other than the camera 112, such as a microphone that analyzes the conversation of the human 114 (e.g., sensor 113). Combining these improvements, data flow 820, data flow 830, and data flow 840 show further examples of the conversion from an intention pose to an intention as determined by the intention determination module 332.
[0033] Approaching the danger zone of vehicle 110 is not the only factor in triggering the judgment of intention towards human 114. Figure 9 depicts, according to some embodiments of the present disclosure, key points used to determine the direction of a human by referring to a landmark by referring to that landmark. In image 910, human 114 has its back to the curb as judged using the above-described system and method. Based on this, the intention judgment module 332 can judge that it is less likely that human 114 will approach the road, and thus vehicle 110 does not need to change its route based on the intention of human 114. However, in image 920, human 114 is facing towards the curb, and in image 930, human 114 is facing the curb. Based on these activities, the image judgment module 332 can be made to judge that the intention of human 114 must be judged to ensure that if the intention of human 114 is to enter the road, vehicle 110 is instructed not to cause danger to human 114.
[0034] While the above description analyzes key points at a macro level, the key point judgment module 337 can judge and process key points granularly to provide more robust information to the intention judgment module 332 when judging the intention of human 114. Figure 10 depicts, according to some embodiments of the present disclosure, a key point cluster for identifying various parts of a human body. Image 1010 illustrates exemplary key points of a human face. By using a larger number of key points, the intention judgment module can judge the emotional state of human 114 (e.g., frowning, happy, etc.), which provides more information for the intention judgment module 332 to judge the intention of human 114. Similarly, the intention judgment module 332 can judge more accurate gestures and the like by judging more key points in the arms as illustrated in image 1020, the legs as illustrated in image 1030, and the body as illustrated in image 1040.
[0035] Figure 11 depicts a flow diagram for converting an image received from a video feed into an intent determination, according to some embodiments of the present disclosure. Process 1100 begins with the processor 202 of the intent determination service 330 obtaining a plurality of series of images from the video feed (e.g., from camera 112 via network 120, as described with reference to FIG. 1 above). The images can be obtained for a predetermined length of time, or for a predetermined number of frames relative to the current time. Images in which no humans are detected (e.g., by the human detection module 333) may be discarded. Next, the processor 202 of the intent determination service 330 determines each key point corresponding to a human in each of the plurality of series of images (e.g., by executing the key point determination module 337, as described above with reference to FIG. 3).
[0036] Next, the processor 202 of the intention determination service 330 aggregates 1106 the respective keypoints for each image into a human pose (e.g., by executing the keypoint aggregation module 331 as described above for FIG. 3). When there are multiple humans in the image, the keypoints for each human (or, as described above, if the intention of any one of the humans may affect the operation of the vehicle 110) are aggregated, and a pose is determined for each. Next, the processor 202 of the intention determination service 330 compares 1108 the pose with an intention template, for example, by causing the intention determination module 332 to execute, transmitting a query to the pose template database 334, and comparing the pose with a plurality of template poses that convert the pose from candidate poses to intentions. Next, the intention determination module 332 determines 1110 whether there is a matching template. For example, the intention determination module 332 receives from the database a response message indicating either the human intention or the inability to identify a matching template. If the response message indicates the human intention, the intention determination module 332 determines that there was a matching template and outputs 1112 a command corresponding to the determined intention (e.g., turn to avoid the movement of human 114, decelerate to enable human 114 to cross, etc.).
[0037] If it is indicated that a template matching the response message cannot be identified, the intent processor 202 of the intent determination service 330 commands the vehicle 110 to enter the safety mode 1114. For example, the intent determination service 330 can include a database 335 of safety mode operations. As described above, the intent determination service 330 can determine the position of a human, as well as other obstacles and features within the image, from a series of images. The safety mode operation database 335 indicates the safety operations to be taken when there is no knowledge of the intent of the human 114, based on the detected obstacles and features, and the relative distance of those obstacles and features, and the human from the vehicle 110. For example, when entering the safety operation mode, it includes commanding the vehicle to perform at least one of turning, accelerating, stopping the movement, providing control to the operator, sounding the horn, or transmitting a message through sound. The safety operation mode can be entered until the intent of the human 114 is determined, or the human 114 leaves the field of view of the camera 112, or leaves the dangerous area, as determined based on the objects detected within the field of view of the camera 112.
[0038] (Considering additional configurations) Throughout this specification, multiple instances can implement a component, operation, or structure that is described as a single instance. Although the individual operations of one or more methods are illustrated and described as separate operations, one or more of the individual operations can be performed simultaneously, and it is not necessary to perform the operations in the order illustrated. Similarly, the structure and functions presented as a single component can be implemented as separate components. These, and other variations, modifications, additions, and improvements are within the scope of the subject matter of this specification.
[0039] In this specification, certain embodiments are described as including logic or several components, or modules or mechanisms. A module can be constituted by either a software module (e.g., code implemented on a machine-readable medium or within a transmission signal) or a hardware module. A hardware module is a tangible unit capable of performing a specific operation and can be configured or arranged in a specific manner. In an exemplary embodiment, one or more computer systems (e.g., stand-alone, client, or server computer systems), or one or more hardware modules of a computer system (e.g., a processor or a group of processors) may be configured by software (e.g., an application or a part of an application) like a hardware module that operates to perform the specific operations described herein.
[0040] In various embodiments, the hardware module may be implemented mechanically or electronically. For example, the hardware module may comprise a dedicated circuit or logic that is permanently configured (e.g., a specific-purpose processor such as a field-programmable gate array (FPGA) or an application-specific integrated circuit (ASIC)) to perform a specific operation. Also, the hardware module can include programmable logic or a circuit (e.g., one included within a general-purpose processor or other programmable processor) that is temporarily configured by software to perform a specific operation. It will be recognized that the decision to implement a hardware module mechanically in a dedicated permanently configured circuit or a temporarily configured circuit (e.g., configured by software) can be driven by cost and time considerations.
[0041] Accordingly, the term "hardware module" should be understood to encompass a tangible entity, which is an entity that is physically constructed and permanently configured (e.g., hardwired), or temporarily configured (e.g., programmed) to operate in a particular manner or to perform particular operations as described herein. As used herein, "hardware-implemented module" refers to a hardware module. Considering embodiments in which a hardware module is temporarily configured (e.g., programmed), each hardware module need not be configured or instantiated at any one instance over time. For example, if a hardware module includes a general-purpose processor configured using software, the general-purpose processor may be configured at different times as different hardware modules. Thus, software can configure the processor, for example, to configure a particular hardware module at one moment and a different hardware module at another moment.
[0042] A hardware module can provide information to, or receive information from, other hardware modules. Thus, the hardware modules described can be considered to be communicatively coupled. When multiple such hardware modules are present simultaneously, communication can be achieved through signal transmission that connects the hardware modules (e.g., across appropriate circuitry and buses). In embodiments where multiple hardware modules are configured or instantiated at different times, communication between such hardware modules can be achieved, for example, through the storage and retrieval of information in a memory structure accessed by the multiple hardware modules. For example, one hardware module can perform an operation and store the output of that operation in a memory device to which it is communicatively coupled. A further hardware module can then later access the memory device, retrieve, and process the stored output. Also, a hardware module can initiate communication with an input device or an output device and can operate on resources (e.g., collect information).
[0043] The various operations of the exemplary methods described herein can be performed, at least in part, by one or more processors that are temporarily (e.g., by software) or permanently configured to perform the relevant operations. Such processors, whether temporarily or permanently configured, can configure processor-implemented modules that operate to perform one or more operations or functions. The modules referred to herein can, in some exemplary embodiments, include processor-implemented modules.
[0044] Similarly, the methods described herein may be implemented, at least in part, on a processor. For example, at least some of the operations of the method can be performed by one or more processors or processor-implemented hardware modules. Certain operations may be distributed among one or more processors and exist not only within a single machine but also across several machines. In some exemplary embodiments, a single processor, or multiple processors, can be located in a single location (e.g., within a home environment, within an office environment, or as a server farm), while in other embodiments, the processors can be distributed across several locations.
[0045] Also, one or more processors can operate to support the operations associated with a "cloud computing" environment or "software as a service" (SaaS). For example, at least some of the operations may be performed by a group of computers (as examples of machines that include processors), and these operations can be accessed via a network (e.g., the Internet) and one or more appropriate interfaces (e.g., application programming interfaces (APIs)).
[0046] Certain operations may be distributed among one or more processors and exist not only within a single machine but also across several machines. In some exemplary embodiments, one or more processors, or processor-implemented modules, can be located in a single geographical location (e.g., within a home environment, within an office environment, or within a server farm). In other exemplary embodiments, one or more processors, or processor-implemented modules, can be distributed across several geographical locations.
[0047] Part of this specification is presented from the perspective of algorithms of operations in bits in a machine memory (such as a computer memory), or data stored as binary digital signals, or symbolic representations. These algorithms, or symbolic representations, are examples of techniques used by those skilled in the data processing field to convey the content of their work to other skilled persons. As used herein, an "algorithm" is a consistent series of operations leading to a desired result, or a similar process. In this context, algorithms, and operations, involve physical manipulations of physical quantities. Usually, but not necessarily so, such quantities can take the form of electrical, magnetic, or optical signals that can be stored, accessed, transformed, combined, compared, or otherwise manipulated by a machine. For mainly reasons of common usage, it may be convenient to refer to such signals using words such as "data", "content", "bit", "value", "element", "symbol", "character", "term", "number", or "digit". However, these words are merely convenient labels and are associated with appropriate physical quantities.
[0048] Unless otherwise specified, the descriptions in this specification using words such as "processing", "computing", "calculating", "judging", "presenting", or "displaying" can refer to the actions, or processes, of a machine (such as a computer) that manipulates, or transforms, data represented as physical (electronic, magnetic, or optical) quantities in one or more memories (such as volatile memory, non-volatile memory, or a combination thereof), in registers, or in other machine components that receive, store, transmit, or display information.
[0049] As used herein, a reference to "one embodiment", or "an embodiment", means that a particular element, feature, structure, or characteristic described in connection with that embodiment is included in at least one embodiment. The appearances of the phrase "in one embodiment" in various places in this specification are not necessarily all referring to the same embodiment.
[0050] Some embodiments can be described using the terms "coupled" and "connected" along with their derivatives. It should be understood that these terms are not intended to be synonyms of each other. For example, some embodiments can be described using the term "connected" to indicate that two or more elements are in direct physical or electrical contact with each other. In another example, some embodiments can be described using the term "coupled" to indicate that two or more elements are in direct physical or electrical contact. However, the term "coupled" can also mean that two or more elements are not in direct contact with each other but still cooperate or interact with each other. Embodiments are not limited in this context.
[0051] As used herein, the terms "comprise," "comprising," "include," "including," "have," "having," or any other variations thereof are intended to include non-exclusive inclusion. For example, a process, method, article, or apparatus that comprises a list of elements is not necessarily limited to only those elements, and may include other elements not expressly listed or inherent to such process, method, article, or apparatus. Further, unless expressly stated to the contrary, "or" refers to "inclusive or" and not "exclusive or." For example, the condition A or B is satisfied by any one of the following: A is true (or present) and B is false (or absent), A is false (or absent) and B is true (or present), and both A and B are true (or present).
[0052] Furthermore, the use of "a" or "an" is employed to describe elements and components of embodiments herein. This is merely for convenience and to impart a general sense of the invention. This description includes one or at least one and, unless it is clearly meant to be otherwise, the singular form should be read as including the plural form.
[0053] By reading this disclosure, those skilled in the art will come to understand additional alternative structural and functional designs for systems and processes for determining human intent through the principles disclosed herein. Thus, while specific embodiments and applications have been illustrated and described, it should be understood that the disclosed embodiments are not limited to the exact structures and components disclosed herein. Various modifications, changes, and variations will become apparent to those skilled in the art, and without departing from the spirit and scope defined in the appended claims, the methods, and the arrangement, operation, and details of the devices disclosed herein can be made.
Claims
1. 1. A method implemented by a computer system having one or more processors, comprising: the one or more processors acquiring a plurality of sequential images from a video feed; determining, by the one or more processors, respective key points corresponding to humans in each image in the plurality of sequential images; the one or more processors identifying regions within a body contour of the human; comparing, by the one or more processors, each region in the body contour to a template in which predefined points indicate each of the various parts included in each region in the contour relative to the human body; applying, to each region within the body contour, the key points corresponding to the predefined points located at a macro level within the various parts of the template, by the one or more processors; applying, to at least one of the regions within the body contour, each of the key points corresponding to the predefined points located at a granular level along the detailed features within the various features included in the template; and aggregating, by the one or more processors, each of the key-points for each image into a pose of the human, determining, by the one or more processors, a plurality of sets of keypoints, each set of keypoints corresponding to a distinct body part of the human; the one or more processors respectively determine a relationship in which an end keypoint in each set of keypoints is sequentially connected to an end keypoint in each other set of keypoints, and store relative positions and relative distances between the connected end keypoints in each of the two sets of keypoints in the connected relationship; the one or more processors determine a vector of relative keypoint movement based on how the end keypoint in each of the sets of keypoints moves relative to the end keypoint in the other set of keypoints as a change in relative position and distance between the end keypoints in each of the two sets of keypoints in the plurality of sequential images over time; the one or more processors mapping the vectors of relative keypoint movement to the pose; and sending a query to a database, by the one or more processors, to compare the pose to a number of template poses for candidate pose-to-intent translation; receiving, by the one or more processors, a response message from the database indicating either the human's intent or an inability to identify a matching template; outputting, by the one or more processors, an instruction corresponding to the intent of the human in response to the response message indicating the intent of the human; A method for providing the above.
2. in response to the response message indicating that the matching template cannot be identified; outputting a command to stop normal operation of the vehicle capturing the video feed and enter a safe operating mode; the one or more processors monitoring the video feed to determine when the human is not present within an image of the video; in response to determining that the human is not present within the video image, the one or more processors output a command to deactivate the safe mode of operation and resume normal operation of the vehicle; Further comprising: The method of claim 1.
3. the safe operating mode may be any of a plurality of modes, and the command includes an indication of a particular one of the plurality of modes, the particular one of the plurality of modes being selected based on a position of the person and other obstacles relative to the vehicle. The method of claim 2.
4. the safe operating mode, when entered, includes commanding the vehicle to at least one of turn, accelerate, stop moving, provide control to an operator, output a visual, audio, or multimedia message, and sound a horn; The method of claim 2.
5. 1. A method implemented by a computer system having one or more processors, comprising: the one or more processors acquiring a plurality of sequential images from a video feed; determining, by the one or more processors, respective key points corresponding to humans in each image in the plurality of sequential images; aggregating, by the one or more processors, each of the key-points for each image into a pose of the human, determining, by the one or more processors, a plurality of sets of keypoints, each set of keypoints corresponding to a distinct body part of the human; determining, by the one or more processors, a vector of relative keypoint movement based on how an end keypoint in each set of keypoints moves relative to an end connected keypoint in each other set of keypoints; the one or more processors mapping the vectors of relative keypoint movement to the pose; and sending a query to a cultural pose intention database, the one or more processors to compare the poses to a plurality of template poses that convert a candidate pose into an intention indicating a respective future action; the one or more processors determining a geographic location of a vehicle capturing the video feed; accessing, by the one or more processors, an index of a plurality of cultural pose intention databases, the index including entries that correspond to each cultural pose intention database of the plurality of cultural pose intention databases, each database location and each address; the one or more processors comparing the geographic location against the entries to determine which cultural pose intent database to query, wherein a matching cultural pose intent database in the plurality of cultural pose intent databases comprises: selecting, by the one or more processors, from the plurality of cultural pose intent databases, a cultural pose intent database to be used in connection with converting the pose into an intent indicative of a respective future action, the selection being performed based on a location at which the video feed was taken, where the pose is indicated to be converted into a first intent by a first database of the plurality of cultural pose intent databases corresponding to a first location, and the pose is indicated to be converted into a second intent, different from the first intent, by a second database of the plurality of cultural pose intent databases corresponding to a second location, different from the first location. based on the respective database locations of the matching cultural pose intent database that match the geographic location by the one or more processors assigning the matching cultural pose intent database as the database to which the query is directed; and receiving, by the one or more processors, a response message from the selected cultural pose intent database indicating either the human's intent indicating a future action or an inability to identify a matching template; outputting, by the one or more processors, in response to the response message indicating the intention of the human indicating the future action, an instruction corresponding to the intention; A method for providing the above.
6. the one or more processors determining whether the plurality of sequential images includes a human; in response to determining that the plurality of sequential images includes the human, the one or more processors perform the step of determining respective key points corresponding to the human in each image of the plurality of sequential images; in response to determining that the plurality of sequential images does not include the human, the one or more processors discarding the plurality of sequential images and acquiring a next plurality of sequential images from the video feed; Further comprising:
2. The method of claim 1.
7. The step of acquiring the plurality of sequential images from the video feed comprises: determining, by the one or more processors, a quantity representing either an amount of time prior to a current time or an amount of consecutive images prior to a currently captured image; the one or more processors obtaining from the video feed a quantity of the successive images relative to either a current time or a currently captured image; Further comprising: The method of claim 1.
8. the intent is determined using at least one of machine learning, statistical analysis, and applying a psychological behavioral model to each image of the plurality of sequential images. The method of claim 1.
9. 1. A computer-readable storage medium comprising encoded instructions that, when executed by a processor of a client device, cause the processor to: acquiring a plurality of successive images from a video feed; determining respective key points corresponding to humans in each image in the plurality of sequential images, identifying regions within a body contour of the human; comparing each region in the body contour to a template having predetermined points indicating each of the various locations included in each region in the contour with respect to the human body; applying, to each of the regions of the body contour, the key points corresponding to the predetermined points located at a macro level within the various regions of the template; applying, to at least one of the various features included in each region of the body contour, each of the key points corresponding to the predetermined points arranged at a granular level along the detailed features within the various features included in the template; and aggregating each of the keypoints for each image into a pose of the human, determining a plurality of sets of key points, each set of key points corresponding to a distinct body part of the human; determining a relationship in which a key point at one end of each set of key points is sequentially connected to a key point at one end of each set of other key points, and storing the relative positions and relative distances between the connected key points at one end of each of the two sets of key points in the connected relationship; determining a vector of relative keypoint movement based on how the end keypoint in each set of keypoints moves relative to the end keypoint in the other set of keypoints as the relative positions and distances between the end keypoints in each of the two sets of keypoints in the plurality of consecutive images change over time; mapping the vectors of relative keypoint movement to the pose; and and sending a query to a database to identify templates that match the pose by comparing the pose to a number of template poses that convert the candidate pose to an intent, each template corresponding to an associated intent; receiving a response message from the database indicating either the human's intent based on a matching template or an inability to identify the matching template; outputting the intent of the human in response to the response message indicating the intent of the human; A computer-readable storage medium for causing a computer to perform the above steps.
10. The instructions cause the processor to, in response to the response message indicating a failure to identify the matching template: outputting a command to stop normal operation of the vehicle capturing the video feed and enter a safe operating mode; monitoring the video feed to determine when the human is not present within an image of the video; in response to determining that the human is not present within the video image, outputting a command to deactivate the safe mode of operation and resume normal operation of the vehicle; Further, 10. The computer-readable storage medium of claim 9.
11. the safe operating mode may be any of a plurality of modes, and the command includes an indication of a particular one of the plurality of modes, the particular one of the plurality of modes being selected based on a position of the person and other obstacles relative to the vehicle. The computer-readable storage medium of claim 10.
12. the safe operating mode, when entered, includes commanding the vehicle to perform at least one of turning, accelerating, stopping movement, providing control to an operator, and sounding a horn; The computer-readable storage medium of claim 10.
13. The instructions cause the processor to, when sending the query to the database: determining a geographic location of a vehicle capturing the video feed; accessing an index of a plurality of candidate databases, the index including entries associating each candidate database of the plurality of candidate databases with each database location and each address; comparing the geographic location against the entries to determine candidate databases to query, wherein a matching candidate database among the plurality of candidate databases is selected based on the respective database locations of the matching candidate databases that match the geographic location; assigning the candidate match database as the database to which the query is to be sent; Further, 10. The computer-readable storage medium of claim 9.
14. The instructions cause the processor to: determining whether the plurality of sequential images includes a human; performing said determining of each key point corresponding to said human in each image of said plurality of sequential images in response to determining that said plurality of sequential images includes said human; in response to determining that the plurality of sequential images does not include the human, discarding the plurality of sequential images and obtaining a next plurality of sequential images from the video feed; Further, 10. The computer-readable storage medium of claim 9.
15. The instructions cause the processor to, when acquiring the plurality of sequential images from the video feed: determining a quantity representing either the length of time prior to the current time or the amount of consecutive images prior to the currently captured image; obtaining from the video feed a quantity of said successive images relative to either a current time or a currently captured image; Further, 15. The computer-readable storage medium of claim 14.
16. The intent is determined using at least one of machine learning, statistical analysis, and applying a psychological behavioral model to each image of the plurality of sequential images.
10. The computer-readable storage medium of claim 9.
Citation Information
Patent Citations
Operation recognition device
JP2011186576A
Information processing device and method, and program
JP2012066026A
Three-dimensional posture estimation device, three-dimensional posture estimation method and program
JP2013020578A
Image processing device, image processing method, and image processing program
WO2015186436A1
Gesture recognition device, gesture recognition method, and information processing device
WO2016167331A1