Managing traffic light detection
By adopting a cross-checking mechanism in autonomous vehicles and cross-verification of the side view traffic light detection system and the front view system, the accuracy problem of the vehicle when detecting traffic lights is solved and safety is improved.
Patent Information
- Application Number
- CN202380068496.3
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Priority Date
- 2022-07-28
- Filing Date
- 2023-07-28
- Publication Date
- 2025-05-06
AI Technical Summary
Autonomous vehicle can be difficult to accurately analyze when detecting traffic lights, especially when the front-view traffic light detection system is malfunctioning or unavailable, which can cause the vehicle to operate in an undesirable manner, such as running a red light.
The cross-checking mechanism is adopted to collect cross-traffic information through an independent side-view traffic light detection system and cross-verified with the results of the front-view traffic light detection system to ensure the accuracy of the traffic light status at the intersection.
Improves the accuracy of traffic light detection of the vehicle at the intersection, prevents undesired operations due to system failures, and enhances driving safety.
Smart Images

Figure CN119948541A_ABST
Abstract
Description
Background Art
[0001] Autonomous vehicles use sensors to generate sensor data associated with their environment and use the sensor data to perceive and operate within that environment. However, certain objects, such as traffic lights, may be difficult to analyze. BRIEF DESCRIPTION OF THE DRAWINGS
[0002] Figure 1 is an example environment in which a vehicle including one or more components of an autonomous system may be implemented;
[0003] Figure 2 is a diagram of one or more systems of a vehicle including an autonomous system;
[0004] Figure 3 yes Figure 1 and Figure 2 a diagram of one or more devices and / or components of one or more systems;
[0005] Figure 4A is a diagram of some components of an autonomous system;
[0006] Figure 4B is a graph of the implementation of a neural network;
[0007] Figure 4C and Figure 4D is a diagram illustrating an example operation of a CNN;
[0008] Figure 5 A block diagram showing an architecture for managing traffic light detection;
[0009] Figure 6 An example of traffic light detection at an intersection is shown;
[0010] Figure 7 is a flow chart of a process for managing traffic light detection; and
[0011] Figure 8 is a flow chart of another process for managing traffic light detection. DETAILED DESCRIPTION
[0012] In the following description, for the purpose of explanation, many specific details are set forth in order to provide a thorough understanding of the present disclosure. However, it will be apparent that the embodiments described in the present disclosure can be implemented without these specific details. In some instances, well-known configurations and devices are illustrated in block diagram form to avoid unnecessarily obscuring aspects of the present disclosure.
[0013] In the accompanying drawings, for ease of description, the specific arrangement or order of schematic elements (such as those representing systems, devices, modules, instruction blocks and / or data elements, etc.) is illustrated. However, those skilled in the art will understand that, unless explicitly described, the specific order or arrangement of schematic elements in the accompanying drawings is not intended to mean that a specific processing order or sequence, or separation of processing is required. In addition, unless explicitly described, the inclusion of schematic elements in the accompanying drawings is not intended to mean that such elements are required in all embodiments, nor is it intended to mean that the features represented by such elements cannot be included in some embodiments or cannot be combined with other elements in some embodiments.
[0014] In addition, in the accompanying drawings, connecting elements (such as solid or dotted lines or arrows, etc.) are used to illustrate the connection, relationship or association between or among two or more other schematic elements, and the absence of any such connecting elements is not intended to mean that there can be no connection, relationship or association. In other words, some connections, relationships or associations between elements are not illustrated in the accompanying drawings so as not to obscure the present disclosure. In addition, for ease of illustration, a single connecting element can be used to represent multiple connections, relationships or associations between elements. For example, if the connecting element represents the communication of a signal, data or instruction (e.g., "software instruction"), it will be understood by those skilled in the art that such an element can represent one or more than one signal path (e.g., bus) that may be needed to affect the communication.
[0015] Although the terms "first", "second" and / or "third", etc. are used to describe various elements, these elements should not be limited by these terms. The terms "first", "second" and / or "third" are only used to distinguish one element from another. For example, a first contact may be referred to as a second contact, and similarly, a second contact may be referred to as a first contact without departing from the scope of the described embodiments. Both the first contact and the second contact are contacts, but they are not the same contacts.
[0016] The terms used in the description of the various embodiments described herein are included only for the purpose of describing a particular embodiment and are not intended to be limiting. As used in the description of the various embodiments described and in the appended claims, the singular forms "a", "an" and "the" are also intended to include plural forms and can be used interchangeably with "one or more than one" or "at least one", unless the context clearly indicates otherwise. It will also be understood that the term "and / or" as used herein refers to and includes any and all possible combinations of one or more than one of the associated listed items. It will also be understood that when the terms "include", "comprise", "have" and / or "have" are used in this specification, it is specifically stated that there are stated features, integers, steps, operations, elements and / or components, but it does not exclude the presence or addition of one or more than one other features, integers, steps, operations, elements, components and / or groups thereof. As used herein, "satisfy" refers to meeting a predetermined condition or requirement, for example, not greater than a predetermined threshold or not less than a certain value.
[0017] As used herein, the terms "communication" and "communicating" refer to at least one of receiving, receiving, transmitting, transmitting and / or providing information (or information represented by, for example, data, signals, messages, instructions and / or commands, etc.). For a unit (e.g., a device, a system, a component of a device or system, and / or a combination thereof) to communicate with another unit, this means that the unit is able to directly or indirectly receive information from the other unit and / or send (e.g., transmit) information to the other unit. This can refer to a direct or indirect connection that is wired and / or wireless in nature. In addition, even if the transmitted information can be modified, processed, relayed and / or routed between the first unit and the second unit, the two units can communicate with each other. For example, even if the first unit passively receives information and does not actively transmit information to the second unit, the first unit can communicate with the second unit. As another example, if at least one intermediary unit (e.g., a third unit located between the first unit and the second unit) processes the information received from the first unit and transmits the processed information to the second unit, the first unit can communicate with the second unit. In some embodiments, a message may refer to a network packet (eg, a data packet, etc.) that includes data.
[0018] As used herein, the term "if" is optionally interpreted to mean "when," "at," "in response to being determined to be," and / or "in response to being detected," etc., depending on the context. Similarly, the phrases "if it is determined" or "if [the stated condition or event] is detected" are optionally interpreted to mean "when determining," "in response to being determined to be" or "when [the stated condition or event] is detected," and / or "in response to being detected," etc., depending on the context. In addition, as used herein, the terms "have," "have," or "possess," etc. are intended to be open-ended terms. Furthermore, unless expressly stated otherwise, the phrase "based on" is intended to mean "based at least in part on."
[0019] Reference will now be made in detail to embodiments, examples of which are illustrated in the accompanying drawings. In the following detailed description, numerous specific details are set forth in order to provide a thorough understanding of the various embodiments described. However, it will be apparent to one of ordinary skill in the art that the various embodiments described may be implemented without these specific details. In other cases, well-known methods, processes, components, circuits, and networks have not yet been described in detail in order not to unnecessarily obscure aspects of the embodiments.
[0020] General Overview
[0021] In some aspects and / or embodiments, the systems, methods, and computer program products described herein include and / or implement techniques for detecting traffic lights. A vehicle (e.g., an autonomous vehicle) is configured to manage traffic light detection at an intersection to make reliable decisions. Specifically, when a vehicle is approaching an intersection, a forward-looking traffic light detection (TLD) system of the vehicle (e.g., using a front camera) classifies the current state (e.g., red, yellow, or green) of a traffic light at the intersection. In order to prevent the vehicle from operating in an undesirable manner (e.g., running a red light when the forward-looking TLD system fails or is unavailable), the vehicle uses a separate and independent side-looking TLD system to cross-check the forward-looking TLD system results, wherein the side-looking TLD system includes sensors (e.g., left / right side cameras, left / right light detection and ranging (Light Detection and Ranging, LiDAR) sensors and / or left / right radio detection and ranging (Radio Detection and Ranging, Radar) sensors) to monitor (e.g., track) cross traffic events and / or behaviors at intersections. The collected side-looking sensor data can be used (e.g., through a sensor tracking algorithm or a machine learning model) to detect different types of cross traffic information at various locations at the intersection (e.g., at road sections perpendicular to the driving direction of the vehicle). Based on the detected crossing traffic information, the vehicle derives the status of the traffic light at the intersection in front of the vehicle by, for example, determining whether the field of view of the crossing traffic is blocked, whether the crossing traffic is stopping or moving, whether the distance between the crossing traffic and the intersection is decreasing or increasing, and / or whether the crossing traffic is decelerating or accelerating.
[0022] When the forward-looking TLD system detects a red or yellow light, the vehicle uses the detected traffic light to make a decision (e.g., slow down or stop completely), and the vehicle can optionally bypass the cross-check using the side sensor. When the forward-looking TLD system detects a green light, the vehicle uses the side-looking TLD results to cross-check the forward-looking TLD results by detecting crossing traffic events / behaviors at the intersection, and derives the corresponding traffic light state. The vehicle compares the green light state detected by the forward-looking TLD system with the traffic light state derived by the side sensor. If the two states are the same, the vehicle determines that the current state of the traffic light is green and travels through the intersection. If the two states are different, the vehicle determines that the current state of the traffic light is red and continues to stop or slow down, which can prevent possible scenarios of operating in an unexpected manner (e.g., running a red light).
[0023] By means of the implementation of the systems, methods and computer program products described herein, a technology for managing traffic light detection is enabled. First, the technology enhances the forward traffic light detection (TLD) system of a vehicle by adding the crossing traffic information at the intersection, which provides comprehensive, consistent, cross-checked and accurate traffic light information at the intersection. Second, the technology uses a side-view sensor independent of the forward TLD system to derive a separate and independent TLD result to cross-check the forward TLD result, which ensures the reliability of the cross-check. Third, the technology uses different types of side-view sensors (e.g., cameras, LiDAR sensors and / or Radar sensors) to detect different types of crossing traffic information at the intersection, which provides multiple and / or cascade checks on the traffic light state at the intersection to ensure the correctness of the TLD results based on the side sensors for reliable cross-checking. Fourth, the technology actively prevents the vehicle from operating in an unexpected manner (e.g., running a red light at an intersection when the forward TLD system fails or is unavailable), which ensures reliable decision-making and improves driving safety. For example, the technology prevents side and near traffic conflicts that occur when a vehicle is operating in an undesirable manner (e.g., running a red light) without a "crumple zone" to the side of the vehicle. A near traffic conflict is an event in which no property is damaged and no personal injury is sustained, but loss or injury could occur without an evasive maneuver. A near traffic conflict is an event that occurs defined by the likelihood of a traffic conflict without a vehicle-initiated maneuver. A traffic conflict is a contact between a vehicle and an actor in the environment.
[0024] Reference now Figure 1, illustrates an example environment 100 in which vehicles including autonomous systems and vehicles not including autonomous systems operate. As illustrated, the environment 100 includes vehicles 102a-102n, objects 104a-104n, routes 106a-106n, areas 108, vehicle-to-infrastructure (V2I) devices 110, a network 112, a remote autonomous vehicle (AV) system 114, a fleet management system 116, and a V2I system 118. The vehicles 102a-102n, the vehicle-to-infrastructure (V2I) devices 110, the network 112, the autonomous vehicle (AV) system 114, the fleet management system 116, and the V2I system 118 are interconnected (e.g., establish connections for communication, etc.) via wired connections, wireless connections, or a combination of wired or wireless connections. In some embodiments, objects 104a-104n are interconnected with at least one of vehicles 102a-102n, vehicle-to-infrastructure (V2I) devices 110, networks 112, autonomous vehicle (AV) systems 114, fleet management systems 116, and V2I systems 118 via wired connections, wireless connections, or a combination of wired or wireless connections.
[0025] Vehicles 102a-102n (individually referred to as vehicles 102 and collectively referred to as vehicles 102) include at least one device configured to transport goods and / or people. In some embodiments, vehicles 102 are configured to communicate with V2I devices 110, remote AV systems 114, fleet management systems 116, and / or V2I systems 118 via network 112. In some embodiments, vehicles 102 include cars, buses, trucks, and / or trains, etc. In some embodiments, vehicles 102 are similar to vehicles 200 described herein (see Figure 2 ). In some embodiments, vehicles 200 in the set of vehicles 200 are associated with an autonomous queue manager. In some embodiments, vehicles 102 travel along respective routes 106a-106n (individually referred to as routes 106 and collectively referred to as routes 106) as described herein. In some embodiments, one or more vehicles 102 include an autonomous system (e.g., an autonomous system that is the same as or similar to autonomous system 202).
[0026] Objects 104a-104n (individually referred to as objects 104 and collectively referred to as objects 104) include, for example, at least one vehicle, at least one pedestrian, at least one cyclist, and / or at least one structure (e.g., a building, a sign, a fire hydrant, etc.), etc. Each object 104 is stationary (e.g., located at a fixed location and over a period of time) or moves (e.g., has a speed and is associated with at least one trajectory). In some embodiments, objects 104 are associated with corresponding locations in area 108.
[0027] Routes 106a-106n (individually referred to as routes 106 and collectively referred to as routes 106) are each associated with (e.g., specify a series of actions (also referred to as trajectories) connecting states along which the AV can navigate. Each route 106 begins at an initial state (e.g., a state corresponding to a first spatiotemporal location and / or speed, etc.) and ends at a final target state (e.g., a state corresponding to a second spatiotemporal location different from the first spatiotemporal location) or a target zone (e.g., a subspace of acceptable states (e.g., terminal states)). In some embodiments, the first state includes a location where one or more individuals will board the AV, and the second state or zone includes one or more locations where one or more individuals boarding the AV will disembark. In some embodiments, routes 106 include multiple acceptable state sequences (e.g., multiple spatiotemporal location sequences) that are associated with (e.g., define multiple trajectories). In an example, routes 106 include only high-level actions or imprecise state locations, such as a series of connecting roads indicating a change of direction at a roadway intersection, etc. Additionally or alternatively, the route 106 may include more precise actions or states, such as, for example, specific target lanes or precise locations within lane regions and target speeds at those locations, etc. In an example, the route 106 includes a plurality of precise state sequences along at least one high-level action with a limited look-ahead horizon to an intermediate target, wherein a combination of consecutive iterations of the limited horizon state sequences cumulatively correspond to a plurality of trajectories that collectively form a high-level route terminating at a final target state or region.
[0028] The area 108 includes a physical area (e.g., a geographic region) in which the vehicle 102 can navigate. In an example, the area 108 includes at least one state (e.g., a country, a province, a separate state of a plurality of states included in a country, etc.), at least a portion of a state, at least one city, at least a portion of a city, etc. In some embodiments, the area 108 includes at least one named thoroughfare (referred to herein as a "road"), such as a highway, an interstate highway, a parkway, a city street, etc. Additionally or alternatively, in some examples, the area 108 includes at least one unnamed road, such as a driveway, a section of a parking lot, a section of an open space and / or undeveloped area, a dirt road, etc. In some embodiments, the road includes at least one lane (e.g., a portion of the road that the vehicle 102 can traverse). In an example, the road includes at least one lane associated with (e.g., identified based on) at least one lane marking line.
[0029] The vehicle-to-infrastructure (V2I) device 110 (sometimes referred to as a vehicle-to-infrastructure or vehicle-to-everything (V2X) device) includes at least one device configured to communicate with the vehicle 102 and / or the V2I system 118. In some embodiments, the V2I device 110 is configured to communicate with the vehicle 102, the remote AV system 114, the queue management system 116, and / or the V2I system 118 via the network 112. In some embodiments, the V2I device 110 includes a radio frequency identification (RFID) device, a sign, a camera (e.g., a two-dimensional (2D) and / or three-dimensional (3D) camera), lane markings, street lights, parking meters, etc. In some embodiments, the V2I device 110 is configured to communicate directly with the vehicle 102. Additionally or alternatively, in some embodiments, the V2I device 110 is configured to communicate with the vehicle 102, the remote AV system 114, and / or the fleet management system 116 via the V2I system 118. In some embodiments, the V2I device 110 is configured to communicate with the V2I system 118 via the network 112.
[0030] The network 112 includes one or more wired and / or wireless networks. In an example, the network 112 includes a cellular network (e.g., a long-term evolution (LTE) network, a third generation (3G) network, a fourth generation (4G) network, a fifth generation (5G) network, a code division multiple access (CDMA) network, etc.), a public land mobile network (PLMN), a local area network (LAN), a wide area network (WAN), a metropolitan area network (MAN), a telephone network (e.g., a public switched telephone network (PSTN)), a private network, an ad hoc network, an intranet, the Internet, a fiber-based network, a cloud computing network, etc., and / or a combination of some or all of these networks, etc.
[0031] The remote AV system 114 includes at least one device configured to communicate with the vehicle 102, the V2I device 110, the network 112, the fleet management system 116, and / or the V2I system 118 via the network 112. In an example, the remote AV system 114 includes a server, a server group, and / or other similar devices. In some embodiments, the remote AV system 114 is co-located with the fleet management system 116. In some embodiments, the remote AV system 114 participates in the installation of some or all of the components of the vehicle (including autonomous systems, autonomous vehicle computing, and / or software implemented by autonomous vehicle computing, etc.). In some embodiments, the remote AV system 114 maintains (e.g., updates and / or replaces) these components and / or software during the life of the vehicle.
[0032] The queue management system 116 includes at least one device configured to communicate with the vehicles 102, the V2I devices 110, the remote AV system 114, and / or the V2I system 118. In an example, the queue management system 116 includes a server, a server group, and / or other similar devices. In some embodiments, the queue management system 116 is associated with a ridesharing company (e.g., an organization for controlling the operation of multiple vehicles (e.g., vehicles including autonomous systems and / or vehicles not including autonomous systems)).
[0033] In some embodiments, the V2I system 118 includes at least one device configured to communicate with the vehicle 102, the V2I device 110, the remote AV system 114, and / or the fleet management system 116 via the network 112. In some examples, the V2I system 118 is configured to communicate with the V2I device 110 via a connection other than the network 112. In some embodiments, the V2I system 118 includes a server, a server group, and / or other similar devices. In some embodiments, the V2I system 118 is associated with a municipality or a private agency (e.g., a private agency for maintaining the V2I device 110, etc.).
[0034] supply Figure 1The number and arrangement of elements illustrated are examples. Figure 1 There may be additional elements, fewer elements, different elements, and / or differently arranged elements than those illustrated. Additionally or alternatively, at least one element of environment 100 may be described as being Figure 1 Additionally or alternatively, at least one set of elements of environment 100 may perform one or more functions described as being performed by at least one different set of elements of environment 100.
[0035] Reference now Figure 2 , vehicle 200 (which can be Figure 1 102) includes, or is associated with, autonomous system 202, powertrain control system 204, steering control system 206, and braking system 208. In some embodiments, vehicle 200 is similar to vehicle 102 (see Figure 1). In some embodiments, the autonomous system 202 is configured to give the vehicle 200 autonomous driving capabilities (e.g., implementing at least one driving automation or maneuver-based function, feature, and / or device, etc., which enables the vehicle 200 to partially or completely operate without human intervention, including but not limited to fully autonomous vehicles (e.g., vehicles that abandon reliance on human intervention, such as Level 5 ADS-operated vehicles, etc.), highly autonomous vehicles (e.g., vehicles that abandon reliance on human intervention in certain situations, such as Level 4 ADS-operated vehicles, etc.), and / or conditionally autonomous vehicles (e.g., vehicles that abandon reliance on human intervention in limited situations). Tools, such as Level 3 ADS operated vehicles, etc.). In one embodiment, the autonomous system 202 includes the operational or tactical functions required to enable the vehicle 200 to operate in traffic on the road and continuously perform part or all of the dynamic driving tasks (DDT). In another embodiment, the autonomous system 202 includes an advanced driver assistance system (ADAS) including driver support features. The autonomous system 202 supports various levels of driving automation ranging from no driving automation (e.g., Level 0) to full driving automation (e.g., Level 5). For a detailed description of fully autonomous vehicles and highly autonomous vehicles, reference may be made to SAE International's standard J3016: Taxonomy and Definitions for Terms Related to On-Road Motor Vehicle Automated Driving Systems, the entire contents of which are incorporated by reference. In some embodiments, the vehicle 200 is associated with an autonomous queue manager and / or a ride-sharing company.
[0036] Autonomous system 202 includes a sensor suite that includes one or more devices such as camera 202a, LiDAR sensor 202b, Radar sensor 202c, and microphone 202d. In some embodiments, autonomous system 202 may include more or fewer devices and / or different devices (e.g., ultrasonic sensors, inertial sensors, GPS receivers (discussed below), and / or odometer sensors for generating data associated with an indication of the distance that vehicle 200 has traveled, etc.). In some embodiments, autonomous system 202 uses one or more devices included in autonomous system 202 to generate data associated with environment 100 described herein. Data generated by one or more devices of autonomous system 202 can be used by one or more systems described herein to observe the environment (e.g., environment 100) in which vehicle 200 is located. In some embodiments, autonomous system 202 includes communication device 202e, autonomous vehicle computing 202f, drive-by-wire (DBW) system 202h, and safety controller 202g.
[0037] The camera 202a includes a communication device 202e, an autonomous vehicle computer 202f, and / or a safety controller 202g configured to communicate with the communication device 202e via a bus (e.g., Figure 3 The camera 202a includes at least one device for communicating with the autonomous vehicle computing 202f (e.g., an image processing unit 202a, a bus ... Figure 1In some embodiments, the autonomous vehicle computing 202f determines a depth to one or more objects in a field of view of at least two of the plurality of cameras based on image data from the at least two cameras. In some embodiments, the camera 202a is configured to capture images of objects within a distance relative to the camera 202a (e.g., up to 100 meters and / or up to 1 kilometer, etc.). Thus, the camera 202a includes features such as sensors and lenses that are optimized for sensing objects at one or more distances relative to the camera 202a.
[0038] In an embodiment, the camera 202a includes at least one camera configured to capture one or more images associated with one or more traffic lights, street signs, and / or other physical objects that provide visual navigation information. In some embodiments, the camera 202a generates traffic light data associated with the one or more images. In some examples, the camera 202a generates traffic light detection (TLD) data associated with one or more images including a format (e.g., RAW, JPEG, and / or PNG, etc.). In some embodiments, the camera 202a that generates TLD data differs from other systems incorporating cameras described herein in that the camera 202a may include one or more cameras with a wide field of view (e.g., a wide-angle lens, a fisheye lens, and / or a lens with a viewing angle of about 120 degrees or greater, etc.) to generate images related to as many physical objects as possible.
[0039] The light detection and ranging (LiDAR) sensor 202b includes a communication device 202e, an autonomous vehicle computing device 202f, and / or a safety controller 202g configured to communicate with the communication device 202e via a bus (e.g., Figure 3The LiDAR sensor 202b includes at least one device that communicates with a bus (the same or similar bus as the bus 302 of the embodiment of the present invention). The LiDAR sensor 202b includes a system configured to emit light from a light emitter (e.g., a laser emitter). The light emitted by the LiDAR sensor 202b includes light outside the visible spectrum (e.g., infrared light, etc.). In some embodiments, during operation, the light emitted by the LiDAR sensor 202b encounters a physical object (e.g., a vehicle) and is reflected back to the LiDAR sensor 202b. In some embodiments, the light emitted by the LiDAR sensor 202b does not penetrate the physical object encountered by the light. The LiDAR sensor 202b also includes at least one light detector that detects the light emitted from the light emitter after encountering the physical object. In some embodiments, at least one data processing system associated with the LiDAR sensor 202b generates an image (e.g., a point cloud and / or a combined point cloud, etc.) representing objects included in the field of view of the LiDAR sensor 202b. In some examples, at least one data processing system associated with the LiDAR sensor 202b generates an image representing the boundaries of the physical object and / or the surface of the physical object (e.g., the topology of the surface), etc. In such examples, the image is used to determine the boundaries of the physical object in the field of view of the LiDAR sensor 202b.
[0040] The radio detection and ranging (Radar) sensor 202c includes a sensor configured to communicate with the communication device 202e, the autonomous vehicle computing device 202f, and / or the safety controller 202g via a bus (e.g., Figure 3 At least one device for communicating with a bus (same or similar bus as bus 302 of the embodiment of the present invention). Radar sensor 202c includes a system configured to transmit (pulsed or continuous) radio waves. The radio waves transmitted by Radar sensor 202c include radio waves within a predetermined spectrum. In some embodiments, during operation, the radio waves transmitted by Radar sensor 202c encounter physical objects and are reflected back to Radar sensor 202c. In some embodiments, the radio waves transmitted by Radar sensor 202c are not reflected by some objects. In some embodiments, at least one data processing system associated with Radar sensor 202c generates a signal representing an object included in the field of view of Radar sensor 202c. For example, at least one data processing system associated with Radar sensor 202c generates an image representing the boundary of a physical object and / or the surface of a physical object (e.g., the topology of the surface), etc. In some examples, the image is used to determine the boundary of a physical object in the field of view of Radar sensor 202c.
[0041] The microphone 202d includes a microphone configured to communicate with the communication device 202e, the autonomous vehicle computing device 202f, and / or the safety controller 202g via a bus (e.g., Figure 3 At least one device for communicating with the vehicle 200 (the same or similar bus as bus 302 of FIG. 1 ). Microphone 202d includes one or more microphones (e.g., an array microphone and / or an external microphone, etc.) that capture an audio signal and generate data associated with (e.g., representing) the audio signal. In some examples, microphone 202d includes a transducer device and / or the like. In some embodiments, one or more systems described herein can receive the data generated by microphone 202d and determine the position (e.g., distance, etc.) of an object relative to vehicle 200 based on an audio signal associated with the data.
[0042] The communication device 202e includes at least one device configured to communicate with the camera 202a, the LiDAR sensor 202b, the Radar sensor 202c, the microphone 202d, the autonomous vehicle computing 202f, the safety controller 202g, and / or the drive-by-wire (DBW) system 202h. For example, the communication device 202e may include at least one device configured to communicate with the camera 202a, the LiDAR sensor 202b, the Radar sensor 202c, the microphone 202d, the autonomous vehicle computing 202f, the safety controller 202g, and / or the drive-by-wire (DBW) system 202h. Figure 3 The communication device 202e may be a device that is the same as or similar to the communication interface 314 of the vehicle. In some embodiments, the communication device 202e includes a vehicle-to-vehicle (V2V) communication device (eg, a device for enabling wireless communication of data between vehicles).
[0043] Autonomous vehicle computing 202f includes at least one device configured to communicate with camera 202a, LiDAR sensor 202b, Radar sensor 202c, microphone 202d, communication device 202e, safety controller 202g, and / or DBW system 202h. In some examples, autonomous vehicle computing 202f includes devices such as client devices, mobile devices (e.g., cellular phones and / or tablet computers, etc.), and / or servers (e.g., computing devices including one or more central processing units and / or graphics processing units, etc.). In some embodiments, autonomous vehicle computing 202f is the same or similar to autonomous vehicle computing 400 described herein. Additionally or alternatively, in some embodiments, autonomous vehicle computing 202f is configured to communicate with an autonomous vehicle system (e.g., with Figure 1 remote AV system 114 of the same or similar autonomous vehicle system), a fleet management system (e.g., Figure 1 of the same or similar queue management system as the queue management system 116), V2I devices (e.g., Figure 1 V2I device 110 that is the same as or similar to V2I device 110) and / or V2I system (e.g., Figure 1The V2I system 118 may communicate with the same or similar V2I system.
[0044] Safety controller 202g includes at least one device configured to communicate with camera 202a, LiDAR sensor 202b, Radar sensor 202c, microphone 202d, communication device 202e, autonomous vehicle computer 202f, and / or DBW system 202h. In some examples, safety controller 202g includes one or more controllers (electrical controllers and / or electromechanical controllers, etc.) configured to generate and / or transmit control signals to operate one or more devices of vehicle 200 (e.g., powertrain control system 204, steering control system 206, and / or braking system 208, etc.). In some embodiments, safety controller 202g is configured to generate control signals that take precedence over (e.g., override) control signals generated and / or transmitted by autonomous vehicle computer 202f.
[0045] The DBW system 202h includes at least one device configured to communicate with the communication device 202e and / or the autonomous vehicle computing 202f. In some examples, the DBW system 202h includes one or more controllers (e.g., electrical controllers and / or electromechanical controllers, etc.) configured to generate and / or transmit control signals to operate one or more devices of the vehicle 200 (e.g., powertrain control system 204, steering control system 206, and / or braking system 208, etc.). Additionally or alternatively, one or more controllers of the DBW system 202h are configured to generate and / or transmit control signals to operate at least one different device of the vehicle 200 (e.g., turn signals, headlights, door locks, and / or windshield wipers, etc.).
[0046] The powertrain control system 204 includes at least one device configured to communicate with the DBW system 202h. In some examples, the powertrain control system 204 includes at least one controller and / or actuator, etc. In some embodiments, the powertrain control system 204 receives control signals from the DBW system 202h, and the powertrain control system 204 causes the vehicle 200 to make longitudinal vehicle movements such as starting to move forward, stopping to move forward, starting to move backward, stopping to move backward, accelerating in a certain direction, decelerating in a certain direction, etc., or to make lateral vehicle movements such as turning left and / or turning right. In an example, the powertrain control system 204 increases, maintains the same, or decreases the energy (e.g., fuel and / or electricity, etc.) provided to the motor of the vehicle, thereby rotating or not rotating at least one wheel of the vehicle 200.
[0047] The steering control system 206 includes at least one device configured to rotate one or more wheels of the vehicle 200. In some examples, the steering control system 206 includes at least one controller and / or actuator, etc. In some embodiments, the steering control system 206 rotates the two front wheels and / or the two rear wheels of the vehicle 200 to the left or right to turn the vehicle 200 left or right. In other words, the steering control system 206 causes the activity required for the adjustment of the y-axis component of the vehicle motion.
[0048] Braking system 208 includes at least one device configured to actuate one or more brakes to decelerate and / or hold vehicle 200 stationary. In some examples, braking system 208 includes at least one controller and / or actuator configured to cause one or more calipers associated with one or more wheels of vehicle 200 to close on respective rotors of vehicle 200. Additionally or alternatively, in some examples, braking system 208 includes an automatic emergency braking (AEB) system and / or a regenerative braking system, etc.
[0049] In some embodiments, vehicle 200 includes at least one platform sensor (not explicitly illustrated) for measuring or inferring an attribute of a state or condition of vehicle 200. In some examples, vehicle 200 includes platform sensors such as a global positioning system (GPS) receiver, an inertial measurement unit (IMU), wheel rate sensors, wheel brake pressure sensors, wheel torque sensors, engine torque sensors, and / or steering angle sensors. Although braking system 208 is Figure 2 Although illustrated as being located on the proximal side of the vehicle 200 , the braking system 208 may be located anywhere in the vehicle 200 .
[0050] Reference now Figure 3, a schematic diagram of an example device 300. As illustrated, device 300 includes a processor 304, a memory 306, a storage component 308, an input interface 310, an output interface 312, a communication interface 314, and a bus 302. In some embodiments, device 300 corresponds to: at least one device of vehicle 102 (e.g., at least one device of a system of vehicle 102); at least one device of vehicle 102 (e.g., at least one device of a system of vehicle 102); and / or one or more devices of network 112 (e.g., one or more devices of a system of network 112). In some embodiments, one or more devices of vehicle 102 (e.g., one or more devices of a system of vehicle 102), and / or one or more devices of network 112 (e.g., one or more devices of a system of network 112) include at least one device 300 and / or at least one component of device 300. As Figure 3 As shown, apparatus 300 includes a bus 302 , a processor 304 , a memory 306 , a storage component 308 , an input interface 310 , an output interface 312 , and a communication interface 314 .
[0051] The bus 302 includes components that permit communication between components of the device 300. In some cases, the processor 304 includes a processor (e.g., a central processing unit (CPU), a graphics processing unit (GPU), and / or an accelerated processing unit (APU), etc.), a microphone, a digital signal processor (DSP), and / or any processing component that can be programmed to perform at least one function (e.g., a field programmable gate array (FPGA) and / or an application specific integrated circuit (ASIC), etc.). The memory 306 includes a random access memory (RAM), a read-only memory (ROM), and / or another type of dynamic and / or static storage device (e.g., flash memory, magnetic memory, and / or optical memory, etc.) that stores data and / or instructions for use by the processor 304.
[0052] Storage component 308 stores data and / or software related to the operation and use of device 300. In some examples, storage component 308 includes a hard disk (e.g., a magnetic disk, an optical disk, a magneto-optical disk, and / or a solid-state disk, etc.), a compact disk (CD), a digital versatile disk (DVD), a floppy disk, a cassette, a tape, a CD-ROM, a RAM, a PROM, an EPROM, a FLASH-EPROM, an NV-RAM, and / or another type of computer-readable medium, and a corresponding drive.
[0053] The input interface 310 includes components that permit the device 300 to receive information, such as via user input (e.g., a touch screen display, a keyboard, a keypad, a mouse, a button, a switch, a microphone, and / or a camera, etc.). Additionally or alternatively, in some embodiments, the input interface 310 includes a sensor for sensing information (e.g., a global positioning system (GPS) receiver, an accelerometer, a gyroscope, and / or an actuator, etc.). The output interface 312 includes components for providing output information from the device 300 (e.g., a display, a speaker, and / or one or more light emitting diodes (LEDs), etc.).
[0054] In some embodiments, communication interface 314 includes a transceiver-like component (e.g., a transceiver and / or a separate receiver and transmitter, etc.) that permits device 300 to communicate with other devices via a wired connection, a wireless connection, or a combination of a wired connection and a wireless connection. In some examples, communication interface 314 permits device 300 to receive information from another device and / or provide information to another device. In some examples, communication interface 314 includes an Ethernet interface, an optical interface, a coaxial interface, an infrared interface, a radio frequency (RF) interface, a universal serial bus (USB) interface, interface and / or cellular network interface, etc.
[0055] In some embodiments, the device 300 performs one or more processes described herein. The device 300 performs these processes based on the processor 304 executing software instructions stored by a computer-readable medium such as a memory 306 and / or a storage component 308. Computer-readable media (e.g., non-transitory computer-readable media) are defined herein as non-transitory memory devices. Non-transitory memory devices include storage space located within a single physical storage device or storage space distributed across multiple physical storage devices.
[0056] In some embodiments, the software instructions are read into the memory 306 and / or storage component 308 from another computer-readable medium or from another device via the communication interface 314. The software instructions stored in the memory 306 and / or storage component 308, when executed, cause the processor 304 to perform one or more processes described herein. Additionally or alternatively, hardwired circuitry is used in place of or in combination with the software instructions to perform one or more processes described herein. Therefore, unless expressly stated otherwise, the embodiments described herein are not limited to any specific combination of hardware circuitry and software.
[0057] The memory 306 and / or the storage component 308 include a data storage unit or at least one data structure (e.g., a database, etc.). The device 300 can receive information from the data storage unit or at least one data structure in the memory 306 or the storage component 308, store information in the data storage unit or at least one data structure, communicate information to the data storage unit or at least one data structure, or search for information stored in the data storage unit or at least one data structure. In some examples, the information includes network data, input data, output data, or any combination thereof.
[0058] In some embodiments, the device 300 is configured to execute software instructions stored in the memory 306 and / or the memory of another device (e.g., another device that is the same as or similar to the device 300). As used herein, the term "module" refers to at least one instruction stored in the memory 306 and / or the memory of another device, which, when executed by the processor 304 and / or the processor of another device (e.g., another device that is the same as or similar to the device 300), causes the device 300 (e.g., at least one component of the device 300) to perform one or more processes described herein. In some embodiments, the module is implemented in software, firmware, and / or hardware, etc.
[0059] supply Figure 3 The number and arrangement of components illustrated are examples. Figure 3 The device 300 may include additional components, fewer components, different components, or differently arranged components than those illustrated. Additionally or alternatively, a set of components (e.g., one or more components) of the device 300 may perform one or more functions described as being performed by another component or set of components of the device 300.
[0060] Reference now Figure 4A, illustrates an example block diagram of an autonomous vehicle computing 400 (sometimes referred to as an "AV stack"). As illustrated, the autonomous vehicle computing 400 includes a perception system 402 (sometimes referred to as a perception module), a planning system 404 (sometimes referred to as a planning module), a positioning system 406 (sometimes referred to as a positioning module), a control system 408 (sometimes referred to as a control module), and a database 410. In some embodiments, the perception system 402, the planning system 404, the positioning system 406, the control system 408, and the database 410 are included in and / or implemented in an automatic navigation system of a vehicle (e.g., the autonomous vehicle computing 202f of the vehicle 200). Additionally or alternatively, in some embodiments, the perception system 402, the planning system 404, the positioning system 406, the control system 408, and the database 410 are included in one or more independent systems (e.g., one or more systems that are the same or similar to the autonomous vehicle computing 400, etc.). In some examples, the perception system 402, planning system 404, positioning system 406, control system 408, and database 410 are included in one or more independent systems located in the vehicle and / or at least one remote system as described herein. In some embodiments, any and / or all of the systems included in the autonomous vehicle computing 400 are implemented in software (e.g., software instructions stored in a memory), computer hardware (e.g., by a microprocessor, microcontroller, application specific integrated circuit (ASIC) and / or field programmable gate array (FPGA), etc.), or a combination of computer software and computer hardware. It will also be understood that in some embodiments, the autonomous vehicle computing 400 is configured to communicate with a remote system (e.g., an autonomous vehicle system that is the same or similar to the remote AV system 114, a fleet management system 116 that is the same or similar to the fleet management system 116, and / or a V2I system that is the same or similar to the V2I system 118, etc.).
[0061] In some embodiments, the perception system 402 receives data associated with at least one physical object in the environment (e.g., data used by the perception system 402 to detect at least one physical object) and classifies the at least one physical object. In some examples, the perception system 402 receives image data captured by at least one camera (e.g., camera 202a), the image being associated with (e.g., representing) one or more physical objects within the field of view of the at least one camera. In such examples, the perception system 402 classifies the at least one physical object based on one or more groups of physical objects (e.g., bicycles, vehicles, traffic signs, and / or pedestrians, etc.). In some embodiments, based on the classification of the physical object by the perception system 402, the perception system 402 transmits data associated with the classification of the physical object to the planning system 404.
[0062] In some embodiments, the planning system 404 receives data associated with a destination and generates data associated with at least one route (e.g., route 106) along which a vehicle (e.g., vehicle 102) can travel toward the destination. In some embodiments, the planning system 404 periodically or continuously receives data (e.g., data associated with the classification of physical objects described above) from the perception system 402, and the planning system 404 updates at least one trajectory or generates at least one different trajectory based on the data generated by the perception system 402. In other words, the planning system 404 can perform tasks related to the tactical functions required to operate the vehicle 101 in traffic on the road. Tactical efforts involve maneuvering the vehicle in traffic during the journey, which includes but is not limited to deciding whether and when to overtake another vehicle, change lanes, or select an appropriate rate, acceleration, deceleration, etc. In some embodiments, the planning system 404 receives data associated with an updated position of the vehicle (e.g., vehicle 102) from the positioning system 406, and the planning system 404 updates at least one trajectory or generates at least one different trajectory based on the data generated by the positioning system 406.
[0063] In some embodiments, the positioning system 406 receives data associated with (e.g., representing) a location of a vehicle (e.g., vehicle 102) in an area. In some examples, the positioning system 406 receives LiDAR data associated with at least one point cloud generated by at least one LiDAR sensor (e.g., LiDAR sensor 202b). In some examples, the positioning system 406 receives data associated with at least one point cloud from multiple LiDAR sensors, and the positioning system 406 generates a combined point cloud based on each point cloud. In these examples, the positioning system 406 compares the at least one point cloud or the combined point cloud with a two-dimensional (2D) and / or three-dimensional (3D) map of the area stored in the database 410. Then, based on the positioning system 406 comparing the at least one point cloud or the combined point cloud with the map, the positioning system 406 determines the position of the vehicle in the area. In some embodiments, the map includes a combined point cloud of the area generated before the navigation of the vehicle. In some embodiments, the map includes, but is not limited to, a high-precision map of roadway geometry, a map describing road network connectivity attributes, a map describing roadway physical attributes (such as traffic speed, traffic flow, number of vehicle and bicycle traffic lanes, lane width, lane traffic direction or type and location of lane markings, or a combination thereof, etc.), and a map describing the spatial location of road features (such as crosswalks, traffic signs, or various types of other driving signals, etc.) In some embodiments, the map is generated in real time based on data received by the perception system.
[0064] In another example, positioning system 406 receives global navigation satellite system (GNSS) data generated by a global positioning system (GPS) receiver. In some examples, positioning system 406 receives GNSS data associated with a location of a vehicle in a region, and positioning system 406 determines the latitude and longitude of the vehicle in the region. In such an example, positioning system 406 determines the position of the vehicle in the region based on the latitude and longitude of the vehicle. In some embodiments, positioning system 406 generates data associated with the position of the vehicle. In some examples, based on positioning system 406 determining the position of the vehicle, positioning system 406 generates data associated with the position of the vehicle. In such an example, the data associated with the position of the vehicle include data associated with one or more semantic attributes corresponding to the position of the vehicle.
[0065] In some embodiments, the control system 408 receives data associated with at least one trajectory from the planning system 404, and the control system 408 controls the operation of the vehicle. In some examples, the control system 408 receives data associated with at least one trajectory from the planning system 404, and the control system 408 controls the operation of the vehicle by generating and transmitting control signals to operate the powertrain control system (e.g., DBW system 202h and / or powertrain control system 204, etc.), the steering control system (e.g., steering control system 206) and / or the braking system (e.g., braking system 208). For example, the control system 408 is configured to perform operational functions such as lateral vehicle motion control or longitudinal vehicle motion control. Lateral vehicle motion control causes the activity required to adjust the y-axis component of the vehicle motion. Longitudinal vehicle motion control causes the activity required to adjust the x-axis component of the vehicle motion. In the example, in the case where the trajectory includes a left turn, the control system 408 transmits a control signal to cause the steering control system 206 to adjust the steering angle of the vehicle 200, thereby turning the vehicle 200 left. Additionally or alternatively, the control system 408 generates and transmits control signals to cause other devices of the vehicle 200 (eg, headlights, turn signals, door locks, and / or windshield wipers, etc.) to change states.
[0066] In some embodiments, perception system 402, planning system 404, positioning system 406, and / or control system 408 implement at least one machine learning model (e.g., at least one multi-layer perceptron (MLP), at least one convolutional neural network (CNN), at least one recurrent neural network (RNN), at least one autoencoder, and / or at least one transformer, etc.). In some examples, perception system 402, planning system 404, positioning system 406, and / or control system 408 implement at least one machine learning model alone or in combination with one or more of the above systems. In some examples, perception system 402, planning system 404, positioning system 406, and / or control system 408 implement at least one machine learning model as part of a pipeline (e.g., a pipeline for identifying one or more objects located in an environment, etc.).
[0067] Database 410 stores data transmitted to, received from, and / or updated by perception system 402, planning system 404, positioning system 406, and / or control system 408. In some examples, database 410 includes a storage component (e.g., a storage component) for storing data and / or software related to operations and using at least one system of autonomous vehicle computing 400. Figure 3In some embodiments, database 410 stores data associated with a 2D and / or 3D map of at least one area. In some examples, database 410 stores data associated with a 2D and / or 3D map of a portion of a city, portions of multiple cities, multiple cities, counties, states, and / or countries (State) (e.g., a country), etc. In such an example, a vehicle (e.g., a vehicle that is the same or similar to vehicle 102 and / or vehicle 200) can be driven along one or more drivable areas (e.g., single-lane roads, multi-lane roads, highways, remote roads, and / or off-road roads, etc.) and cause at least one LiDAR sensor (e.g., a LiDAR sensor that is the same or similar to LiDAR sensor 202b) to generate data associated with an image representing an object included in the field of view of the at least one LiDAR sensor.
[0068] In some embodiments, database 410 can be implemented across multiple devices. In some examples, database 410 includes a vehicle (e.g., a vehicle that is the same or similar to vehicle 102 and / or vehicle 200), an autonomous vehicle system (e.g., an autonomous vehicle system that is the same or similar to remote AV system 114), a fleet management system (e.g., a vehicle that is the same or similar to remote AV system 114), and a fleet management system (e.g., a vehicle that is the same or similar to remote AV system 114). Figure 1 The same or similar queue management system as the queue management system 116 of FIG. 1 and / or the V2I system (e.g., Figure 1 The V2I system 118 is the same as or similar to the V2I system 118).
[0069] Reference now Figure 4B , a diagram illustrating an implementation of a machine learning model. More specifically, a diagram illustrating an implementation of a convolutional neural network (CNN) 420. For purposes of illustration, the following description of CNN 420 will be with respect to implementing CNN 420 by perception system 402. However, it will be understood that in some examples, CNN 420 (e.g., one or more components of CNN 420) is implemented by other systems (such as planning system 404, positioning system 406, and / or control system 408, etc.) other than or in addition to perception system 402. Although CNN 420 includes certain features as described herein, these features are provided for purposes of illustration and are not intended to limit the present disclosure.
[0070] CNN 420 includes a plurality of convolutional layers including a first convolutional layer 422, a second convolutional layer 424, and a convolutional layer 426. In some embodiments, CNN 420 includes a subsampling layer 428 (sometimes referred to as a pooling layer). In some embodiments, subsampling layer 428 and / or other subsampling layers have a dimension that is smaller than the dimension of the upstream system (i.e., the number of nodes). With the subsampling layer 428 having a dimension that is smaller than the dimension of the upstream layer, CNN 420 merges the amount of data associated with the initial input and / or output of the upstream layer, thereby reducing the amount of computation required for CNN 420 to perform downstream convolution operations. Additionally or alternatively, with the subsampling layer 428 being associated with (e.g., configured to perform) at least one subsampling function (as described below with respect to Figure 4C and Figure 4D As described above, CNN 420 incorporates the amount of data associated with the initial input.
[0071] The perception system 402 performs the convolution operation based on the perception system 402 providing respective inputs and / or outputs associated with each of the first convolution layer 422, the second convolution layer 424, and the convolution layer 426 to generate respective outputs. In some examples, the perception system 402 implements the CNN 420 based on the perception system 402 providing data as input to the first convolution layer 422, the second convolution layer 424, and the convolution layer 426. In such examples, the perception system 402 provides data as input to the first convolution layer 422, the second convolution layer 424, and the convolution layer 426 based on the perception system 402 receiving data from one or more different systems (e.g., one or more systems of a vehicle that is the same or similar to the vehicle 102, a remote AV system that is the same or similar to the remote AV system 114, a queue management system that is the same or similar to the queue management system 116, and / or a V2I system that is the same or similar to the V2I system 118, etc.). The following is about Figure 4C Includes a detailed description of the convolution operation.
[0072] In some embodiments, the perception system 402 provides data associated with the input (referred to as the initial input) to the first convolutional layer 422, and the perception system 402 generates data associated with the output using the first convolutional layer 422. In some embodiments, the perception system 402 provides the output generated by the convolutional layer as input to a different convolutional layer. For example, the perception system 402 provides the output of the first convolutional layer 422 as input to the subsampling layer 428, the second convolutional layer 424, and / or the convolutional layer 426. In such an example, the first convolutional layer 422 is referred to as an upstream layer, and the subsampling layer 428, the second convolutional layer 424, and / or the convolutional layer 426 are referred to as downstream layers. Similarly, in some embodiments, the perception system 402 provides the output of the subsampling layer 428 to the second convolutional layer 424 and / or the convolutional layer 426, and in this example, the subsampling layer 428 will be referred to as the upstream layer, and the second convolutional layer 424 and / or the convolutional layer 426 will be referred to as the downstream layer.
[0073] In some embodiments, before the perception system 402 provides the input to the CNN 420, the perception system 402 processes the data associated with the input provided to the CNN 420. For example, the perception system 402 processes the data associated with the input provided to the CNN 420 based on the perception system 402 normalizing the sensor data (e.g., image data, LiDAR data, and / or Radar data, etc.).
[0074] In some embodiments, CNN 420 generates an output based on perception system 402 performing convolution operations associated with each convolution layer. In some examples, CNN 420 generates an output based on perception system 402 performing convolution operations associated with each convolution layer and the initial input. In some embodiments, perception system 402 generates an output and provides the output to fully connected layer 430. In some examples, perception system 402 provides the output of convolution layer 426 to fully connected layer 430, wherein fully connected layer 430 includes data associated with multiple feature values referred to as F1, F2, ..., FN. In this example, the output of convolution layer 426 includes data associated with multiple output feature values representing predictions.
[0075] In some embodiments, perception system 402 identifies a prediction from the plurality of predictions based on perception system 402 identifying a feature value associated with a highest likelihood of being a correct prediction from the plurality of predictions. For example, where fully connected layer 430 includes feature values F1, F2, ..., FN and F1 is the largest feature value, perception system 402 identifies the prediction associated with F1 as the correct prediction from the plurality of predictions. In some embodiments, perception system 402 trains CNN 420 to generate the predictions. In some examples, perception system 402 trains CNN 420 to generate the predictions based on perception system 402 providing training data associated with the predictions to CNN 420.
[0076] Reference now Figure 4C and Figure 4D , a diagram illustrating an example operation of CNN 440 utilizing perception system 402. In some embodiments, CNN 440 (e.g., one or more components of CNN 440) is coupled to CNN 420 (e.g., one or more components of CNN 420) (see Figure 4B ) are the same or similar.
[0077] At step 450, the perception system 402 provides data associated with the image as input to the CNN 440 (step 450). For example, as illustrated, the perception system 402 provides data associated with the image to the CNN 440, where the image is a grayscale image represented as values stored in a two-dimensional (2D) array. In some embodiments, the data associated with the image may include data associated with a color image represented as values stored in a three-dimensional (3D) array. Additionally or alternatively, the data associated with the image may include data associated with an infrared image and / or a Radar image, etc.
[0078] At step 455, CNN 440 performs a first convolution function. For example, CNN 440 performs a first convolution function based on CNN 440 providing a value representing an image as an input to one or more neurons (not explicitly illustrated) included in first convolution layer 442. In this example, the value representing the image may correspond to a value of a region (sometimes referred to as a receptive field) representing the image. In some embodiments, each neuron is associated with a filter (not explicitly illustrated). The filter (sometimes referred to as a kernel) may be represented as an array of values corresponding in size to the value provided as input to the neuron. In one example, the filter may be configured to identify edges (e.g., horizontal lines, vertical lines, and / or straight lines, etc.). In successive convolution layers, the filters associated with the neurons may be configured to continuously identify more complex patterns (e.g., arcs and / or objects, etc.).
[0079] In some embodiments, CNN 440 performs a first convolution function based on CNN 440 multiplying the values of each neuron provided as input to one or more neurons included in the first convolution layer 442 by the values of the filters corresponding to each neuron in the same or more neurons. For example, CNN 440 may multiply the values of each neuron provided as input to one or more neurons included in the first convolution layer 442 by the values of the filters corresponding to each neuron in the one or more neurons to generate a single value or an array of values as output. In some embodiments, the collective output of the neurons of the first convolution layer 442 is referred to as a convolution output. In some embodiments, when each neuron has the same filter, the convolution output is referred to as a feature map.
[0080] In some embodiments, CNN 440 provides the output of each neuron of the first convolutional layer 442 to the neurons of the downstream layer. For clarity, the upstream layer may be a layer that transmits data to a different layer (referred to as the downstream layer). For example, CNN 440 may provide the output of each neuron of the first convolutional layer 442 to the corresponding neurons of the subsampling layer. In the example, CNN 440 provides the output of each neuron of the first convolutional layer 442 to the corresponding neurons of the first subsampling layer 444. In some embodiments, CNN 440 adds a bias value to the aggregate set of all values provided to each neuron of the downstream layer. For example, CNN 440 adds a bias value to the aggregate set of all values provided to each neuron of the first subsampling layer 444. In such an example, CNN 440 determines the final value to be provided to each neuron of the first subsampling layer 444 based on the aggregate set of all values provided to each neuron and the activation function associated with each neuron of the first subsampling layer 444.
[0081] At step 460, CNN 440 performs a first subsampling function. For example, based on CNN 440 providing the values output by first convolutional layer 442 to the corresponding neurons of first subsampling layer 444, CNN 440 may perform the first subsampling function. In some embodiments, CNN 440 performs the first subsampling function based on an aggregation function. In an example, CNN 440 performs the first subsampling function based on CNN 440 determining the maximum input (referred to as a maximum pooling function) among the values provided to a given neuron. In another example, CNN 440 performs the first subsampling function based on CNN 440 determining the average input (referred to as an average pooling function) among the values provided to a given neuron. In some embodiments, based on CNN 440 providing values to the respective neurons of first subsampling layer 444, CNN 440 generates an output, which is sometimes referred to as a subsampled convolution output.
[0082] At step 465, CNN 440 performs a second convolution function. In some embodiments, CNN 440 performs the second convolution function in a manner similar to how CNN 440 performs the first convolution function described above. In some embodiments, CNN 440 performs the second convolution function based on CNN 440 providing the value output by first subsampling layer 444 as input to one or more neurons (not explicitly illustrated) included in second convolution layer 446. In some embodiments, as described above, each neuron of second convolution layer 446 is associated with a filter. As described above, the filter (one or more) associated with second convolution layer 446 can be configured to recognize more complex patterns than the filter associated with first convolution layer 442.
[0083] In some embodiments, the CNN 440 performs a second convolution function based on the CNN 440 multiplying the value of each neuron provided as input to the one or more neurons included in the second convolution layer 446 by the value of the filter corresponding to each neuron of the one or more neurons. For example, the CNN 440 may multiply the value of each neuron provided as input to the one or more neurons included in the second convolution layer 446 by the value of the filter corresponding to each neuron of the one or more neurons to generate a single value or a value array as an output.
[0084] In some embodiments, the CNN 440 provides the output of each neuron of the second convolutional layer 446 to the neurons of the downstream layer. For example, the CNN 440 may provide the output of each neuron of the first convolutional layer 442 to the corresponding neurons of the subsampling layer. In an example, the CNN 440 provides the output of each neuron of the first convolutional layer 442 to the corresponding neurons of the second subsampling layer 448. In some embodiments, the CNN 440 adds a bias value to the aggregate set of all values provided to each neuron of the downstream layer. For example, the CNN 440 adds a bias value to the aggregate set of all values provided to each neuron of the second subsampling layer 448. In such an example, the CNN 440 determines the final value provided to each neuron of the second subsampling layer 448 based on the aggregate set of all values provided to each neuron and the activation function associated with each neuron of the second subsampling layer 448.
[0085] At step 470, CNN 440 performs a second subsampling function. For example, based on CNN 440 providing the values output by second convolutional layer 446 to corresponding neurons of second subsampling layer 448, CNN 440 may perform a second subsampling function. In some embodiments, based on CNN 440 using an aggregation function, CNN 440 performs a second subsampling function. In an example, as described above, based on CNN 440 determining the maximum input or average input among the values provided to a given neuron, CNN 440 performs a first subsampling function. In some embodiments, based on CNN 440 providing values to respective neurons of second subsampling layer 448, CNN 440 generates an output.
[0086] At step 475, CNN 440 provides the output of each neuron of second subsampling layer 448 to fully connected layer 449. For example, CNN 440 provides the output of each neuron of second subsampling layer 448 to fully connected layer 449 so that fully connected layer 449 generates output 480. In some embodiments, fully connected layer 449 is configured to generate output associated with prediction (sometimes referred to as classification). The prediction may include an indication that the objects included in the image provided as input to CNN 440 include objects and / or sets of objects, etc. In some embodiments, perception system 402 performs one or more operations and / or provides data associated with the prediction to various systems described herein.
[0087] Figure 5 A block diagram of an architecture 500 for managing traffic light detection according to one or more embodiments is shown. In an embodiment, the architecture 500 is implemented in an autonomous system of a vehicle. In some examples, the vehicle is Figure 2 , and the architecture 500 is implemented by the autonomous system 202 (e.g., fully, partially, etc.) of the vehicle 200. The architecture 500 is configured to manage traffic light detections at intersections by cross-checking traffic light detections derived from different traffic light detection (TLD) systems to make reliable decisions.
[0088] Architecture 500 includes a perception system 510 (in some embodiments, perception system 510 and Figure 4A ) and a planning system 520 (in some embodiments, the planning system 520 is the same as or similar to the perception system 402 shown in Figure 4A). The perception system 510 selectively obtains area information of at least one intersection from the map mapping database 501, for example, based on the current location of the vehicle and / or the route of the vehicle. The map mapping database 501 stores a data structure that associates each intersection with a traffic light at the intersection and a corresponding state. Based on the area information and the traffic light detection data, the perception system 510 determines traffic light information 515, for example, the state of the intersection or the state of the traffic light for the upcoming section of the road at the intersection. The perception system 510 provides the traffic light information 515 to the planning system 520 to determine the action to be taken by the vehicle when it arrives at the intersection. The action to be taken may be, for example, to stop, slow down, or continue or accelerate at the current speed, as well as other appropriate actions. The planning system 520 determines the area information of at least one intersection from the map mapping database 501 based on the traffic light information 515 and other data (for example, from Figure 4A The vehicle operates according to the action determined by the control system, which is connected to the control system such as Figure 4A The control system 408 shown is the same or similar.
[0089] In some embodiments, the control system receives data associated with the determined action from the planning system 520, and the control system controls the operation of the vehicle. In some examples, the control system controls the operation of the vehicle by generating a control signal based on the data associated with the determined action and transmitting the control signal to operate a powertrain control system (in some embodiments, the powertrain control system is the same or similar to the DBW system 202h or the powertrain control system 204), a steering control system (in some embodiments, the steering control system is the same or similar to the steering control system 206) and / or a braking system (in some embodiments, the braking system is the same or similar to the braking system 208). For example, when an operation in an unexpected manner (e.g., running a red light) is detected and the determined action is to slow down or stop, the control system generates a control signal and transmits the control signal to the braking system to slow down or prepare to stop; when a green light is detected and the determined action is to continue or accelerate at the current rate, the control system generates a control signal and transmits the control signal to the powertrain control system to maintain the current rate or accelerate. In this way, reliable cross-check traffic light detection enables the vehicle to be driven safely even when the forward-looking TLD system fails.
[0090] In one embodiment, architecture 500 includes, for example, Figure 4A The map mapping database 501 is implemented in the database 410 shown. In another embodiment, the map mapping database 501 is external to the architecture 500 and is stored on a server (e.g., Figure 1The map mapping database 501 includes road network information, for example, a high-precision map with roadway geometry attributes, a map describing road network connectivity attributes, a map describing roadway physical attributes (such as traffic speed, traffic flow, the number of motor vehicle and bicycle traffic lanes, lane width, lane traffic direction or the type and location of lane markings, or a combination thereof, etc.), and a map describing the spatial location of an area of interest (such as an intersection, a crosswalk, a traffic sign, or various types of other travel signals, etc.). In an embodiment, a high-precision map is constructed by adding data to a low-precision map via automatic or manual annotation. For illustrative purposes only, an intersection is described herein as an example of an area of interest.
[0091] The map mapping database 501 includes the area information of the intersection in the map. In one embodiment, the area information of the intersection includes an intersection identifier (ID), a series of states for the intersection representing the behavior of the traffic light at the intersection, information related to the road segment at the intersection, and information related to the traffic light at the intersection.
[0092] Figure 6 Schematic diagram showing an example traffic light detection 600 at an intersection 602. In some embodiments, intersection 602 is an intersection traversed by a vehicle (which, in some embodiments, is the same or similar to vehicle 200) and corresponds to an area of interest for the vehicle. Figure 6 As shown, the intersection 602 is associated with four road segments 610, 620, 630, 640 around the center of the intersection 602. Each road segment represents an area around the intersection and is associated with, for example, two paths with opposite or angled traffic directions. Each path includes one or more than one lane.
[0093] In some embodiments, Figure 6As shown, the road segment 610 is associated with a first path 612a having a first path direction and a second path 612b having a second path direction opposite to the first path direction. The road segment 620 is associated with a third path 622a having a third path direction and a fourth path 622b having a fourth path direction opposite to the third path direction. The road segment 630 is associated with a fifth path 632a having a fifth path direction and a sixth path 632b having a sixth path direction opposite to the fifth path direction. The road segment 640 is associated with a seventh path 642a having a seventh path direction and an eighth path 642b having an eighth path direction opposite to the seventh path direction. In some cases, the sixth path direction is the same as the first path direction, and the sixth path 632b is an extension of the first path 612a; the eighth path direction is the same as the third path direction, and the eighth path 642b is an extension of the third path 622a; the second path direction is the same as the fifth path direction, and the second path 612b is an extension of the fifth path 632a; the fourth path direction is the same as the seventh path direction, and the fourth path 622b is an extension of the seventh path 642a. Path pairs 632b and 612a, 642b and 622a, 612b and 632a, 622b and 642a are separated by intersection 602.
[0094] At (or around) each road segment there is one or more traffic lights located there and configured to control vehicle movement for traffic on that road segment and / or one or more other road segments. Figure 6 As shown, there are two traffic lights 614a, 614b at (or around) road section 610; there are two traffic lights 624a, 624b at (or around) road section 620; there are two traffic lights 634a, 634b at (or around) road section 630; there are two traffic lights 644a, 644b at (or around) road section 640. Each traffic light includes three bulbs, such as red, yellow and green. In one embodiment, the traffic light includes an arrow, such as a left arrow, a right arrow, an upward arrow or a downward arrow.
[0095] Each traffic light is positioned toward a road segment for directing the movement of traffic (e.g., including vehicles and / or pedestrians) associated with (e.g., from) the road segment. For example, traffic light 614a is positioned at road segment 610 and is directed toward vehicles traveling on path 612a associated with road segment 610, while traffic light 634b is positioned at road segment 630 and is also directed toward vehicles traveling on path 612a for directing the movement of traffic associated with road segment 630. Thus, a road segment may be associated with a traffic light positioned at the road segment and also associated with a traffic light positioned at one or more other road segments at the same intersection 602, and the traffic light is positioned toward the road segment for directing the movement of traffic associated with the road segment.
[0096] The vehicle 200 is traveling along a route (e.g., route 601 approaching an intersection 602 from a road segment 610). To make driving decisions, in one embodiment, the vehicle 200 monitors the behavior of traffic lights (i.e., traffic lights 614a and / or 634b) governing traffic movement on the road segment 610 at the intersection 602 and / or the behavior of crossing traffic lights (e.g., traffic lights 624a, 644b for an intersection road segment 620 and / or traffic lights 624b, 642b for an intersection road segment 640).
[0097] Return to reference Figure 5 , the perception system 510 includes a map information extractor 512. The map information extractor 512 extracts regional information for one or more regions of interest for a vehicle. In one example, based on the current location and / or current route of the vehicle, the map information extractor 512 extracts regional information for one or more intersections around the current location and / or current route of the vehicle. The regional information of the intersection includes an intersection identifier (ID), a series of states for the intersection representing the behavior of the traffic lights at the intersection, information on the road section in the intersection, and information on the traffic lights for the road section. The information on the traffic lights for the road section includes a series of states for each traffic light and a predetermined duration for each state. The states of the traffic lights for different road sections at the same intersection are coordinated with each other. For example, if the traffic light (e.g., 634b) for the road section (e.g., 610) is in the green state, the first crossing traffic light (e.g., 644b) for the first intersection road section (e.g., 620) and the second crossing traffic light (e.g., 624b) for the second intersection road section (e.g., 640) are in the red state. If the traffic light for the road section is in the red state, the first crossing traffic light for the first intersection road section and the second crossing traffic light for the second intersection road section are in the green state. A series of states of the traffic light changes in a loop form in which the first state starts at the end of the last state.
[0098] In one embodiment, Figure 5 As shown, the architecture 500 includes a forward-looking traffic light detection (TLD) system 502 for sensing or measuring attributes of the vehicle environment in front of the vehicle. In one example, the forward-looking TLD system 502 uses a forward-looking camera system 504 to obtain information about traffic lights, street signs, and other physical objects that provide visual navigation information. The forward-looking camera system 504 has a field of view (FOV) (e.g., Figure 6 The forward-looking camera system 504 includes one or more forward-looking cameras (e.g., Figure 7 CAM_M_F 711a, CAM_N_F 711b shown). Each camera (e.g., using a wide-angle lens or a fisheye lens) has a wide field of view to obtain information about as many physical objects that provide visual navigation information as possible, so that the vehicle can access all relevant navigation information provided by these objects. For example, the viewing angle of the TLD system is about 120 degrees or more. In one example, CAM_M_F represents a forward-looking camera with a medium field of view (FOV) ranging, for example, from 5 meters to 50 meters, and CAM_N_F represents a forward-looking camera with a narrow field of view ranging, for example, from 50 meters to 150 meters (or 200 meters). In some embodiments, in response to determining that the distance from the vehicle to the intersection meets, for example, not greater than a predetermined threshold (e.g., 50 meters), the forward-looking camera system 504 switches from a first forward-looking camera (e.g., CAM_N_F) to a second forward-looking camera (e.g., CAM_M_F).
[0099] like Figure 7 As discussed in further detail in , forward-looking TLD system 502 receives one or more forward-looking images from forward-looking camera system 504. The one or more forward-looking images include an image of at least one traffic light (e.g., traffic light 634b and / or traffic light 614a) for a road segment (e.g., road segment 610) that the vehicle is approaching. In one example, forward-looking TLD system 502 provides image data of the one or more forward-looking images as forward-looking TLD data 503 to perception system 510, and perception system 510 generates the image data based on, for example, an image processing algorithm (such as a feature extraction algorithm) or a machine learning model (such as a Figure 4B CNN 420 or Figure 4C and Figure 4DThe perception system 510 processes one or more forward-looking images to process the image data to determine the state of the traffic light for the road segment. For example, the perception system 510 uses a machine learning model that receives the image data as input and generates a prediction representing the state of the traffic light for the road segment as output. In one example, the forward-looking TLD system 502 determines the state of the traffic light for the road segment based on the one or more forward-looking images as forward-looking TLD data 503, and provides the forward-looking TLD data 503 to the perception system 510.
[0100] In some embodiments, the perception system 510 includes a traffic light information (TLI) generator 514 for receiving forward-looking TLD data 503 from the forward-looking TLD system 502. The TLI generator 514 generates traffic light information 515 based on the TLD data 503 and / or information about the traffic light for the road segment from the map information extractor 512. The traffic light information 515 includes a current state of the traffic light for the road segment based on the TLD data 503. The current state is a red state, a green state, or a yellow state. In one example, the traffic light information 515 includes a remaining time for the current state of the traffic light for the road segment. The TLI generator 514 determines the remaining time based on (i) a predetermined duration for the current state obtained from the map information extractor 512 and (ii) a time point of a state change immediately before the current state. In one example, the TLI generator 514 continuously generates the traffic light information 515 and determines a time point of a state change immediately before a current state based on previously generated traffic light information and uses the time point to determine a remaining time for the current state.
[0101] The perception system 510 provides the traffic light information 515 to the planning system 520. In some embodiments, the planning system 520 updates the route based on the traffic light information 515 and provides the perception system 510 (e.g., the map information extractor 512) with the planned route 525. The perception system 510 updates the area information of one or more intersections obtained from the map mapping database 501 based on the planned route from the planning system 520.
[0102] The vehicle (e.g., the planning system 520) generates a traffic light based on the traffic light information 515 (e.g., the current state of the traffic light and the remaining time of the current state), the traffic light information (e.g., a series of states and the duration for each state from the map information extractor 512), and the information from the vehicle (e.g., Figure 2The vehicle 200 shown in FIG. 200 ) determines the current distance to the intersection (e.g., a stop sign in front of the vehicle), the route of the vehicle, and the current speed of the vehicle to determine the state of the traffic light when the vehicle arrives at the intersection from the road segment and the time remaining in the state when the vehicle arrives at the intersection. The vehicle (e.g., the planning system 520) determines which action to take for the vehicle based on the state when the vehicle arrives at the intersection and the time remaining in the state at this time, such as stopping, slowing down, continuing at the current speed, or accelerating the current speed. The planning system 520 determines the state of the traffic light when the vehicle arrives at the intersection based on the state when the vehicle arrives at the intersection and the time remaining in the state at this time. The planning system 520 determines the state of the traffic light when the vehicle arrives at the intersection based on the state when the vehicle arrives at the intersection and the time remaining in the state at this time. Figure 4A The action is determined by the control system (e.g., Figure 4A The control system 408 shown causes the vehicle to operate according to the action.
[0103] like Figure 7 As further illustrated in detail in , in some cases, if the forward-looking TLD system 502 fails or fails to function properly, the red state of the traffic light is erroneously derived as a green state, which may cause the vehicle to operate in an undesirable manner (e.g., run a red traffic light) at risk. In the example, running a red light refers to a vehicle traveling from a road segment into an intersection when the traffic light that governs the road segment is in a red state.
[0104] In some embodiments, Figure 5As shown, the architecture 500 includes a side-view traffic light detection (TLD) system 506, which is to be used as a second TLD system independent of the front-view TLD system 502 to cross-check the front-view TLD data 503 generated by the front-view TLD system 502. The side-view TLD system 506 is used to derive the state of the traffic light for the vehicle based on the crossing traffic information (e.g., crossing traffic events and / or behaviors), because the crossing traffic information is related to the state of the crossing traffic light, and the state of the crossing traffic light is coordinated with the state of the traffic light. In one example, the crossing traffic includes one or more objects (e.g., other vehicles, pedestrians, or animals) at one or more cross roads that are different from the road that the vehicle is approaching. In one example, one or more parameters of the object in the crossing traffic are detected by a side Radar sensor (as described below). One or more parameters include rate (or speed) and distance (or range). Crossing traffic events and / or behaviors include whether visibility of the crossing traffic is obscured, whether the crossing traffic is stopped (or whether the crossing traffic has a velocity of zero or substantially equal to zero), whether the crossing traffic is approaching the intersection or moving away from the road the vehicle is approaching when the distance between the crossing traffic and the intersection is greater than a limit (e.g., a limit listed in a stopping distance table), and / or whether the crossing traffic is decelerating as the velocity is reduced.
[0105] like Figure 5 As shown, in some embodiments, the side-view TLD system 506 includes a side-view camera system 508a (e.g., Figure 2 camera 202a), LiDAR system 508b (e.g., Figure 2 LiDAR sensor 202b) and Radar system 508c (e.g., Figure 2 Radar sensor 202c). In some examples, for example, Figure 7 As illustrated, the side-view camera system 508a includes one or more left / right side view cameras (e.g., CAM_F_L 721, CAM_F_R 722); the LiDAR system 508b includes one or more left / right side LiDAR sensors (e.g., LiDAR_F_L 723, LiDAR_F_R 724); and the Radar system 508c includes one or more side Radar sensors (e.g., RADAR_F_L 725, RADAR_F_R 726). In one example, the side-view TLD system 506 has a left field of view (FOV) (e.g., as shown in FIG. 1 ) using the left sensors (e.g., CAM_F_L 721, LiDAR_F_L 723, and RADAR_F_L 725). Figure 6In one example, the side-view TLD system 506 has a right field of view (FOV) (e.g., as shown in FIG. 6 ) using the right sensors (e.g., CAM_F_R 722, LiDAR_F_R 724, and RADAR_F_R 726). Figure 6 656). Figure 6 As illustrated, in one example, the left FOV 654 covers information of crossing traffic 626 at the intersection 620 , and the right FOV 656 covers information of crossing traffic 646 at the intersection 640 .
[0106] The side look TLD system 506 receives side look sensor data (e.g., side look images) from the side look camera system 508a, receives side look sensor data (e.g., LiDAR sensor data) from the LiDAR system 508b, and receives side look sensor data (e.g., Radar sensor data) from the Radar system 508c. The side look TLD system 506 generates side look TLD data 505 based on the side look sensor data and provides the side look TLD data 505 to the perception system 510 (e.g., TLI generator 514). In one example, the side look TLD data 505 includes the side look sensor data, and the TLI generator 514 derives the road segment that governs the vehicle being approached (e.g., Figure 6 In one example, the side view TLD data 505 includes the state of the traffic light derived by the side view TLD system 506 based on the side view sensor data.
[0107] In one embodiment, for example, Figure 7 As further illustrated in detail in , the collected side sensor data is used to detect different types of traffic information at various locations at the intersection (e.g., at a road section perpendicular to the driving direction of the vehicle). As described above, the detected traffic information may include crossing traffic events and / or behaviors. For example, the detected traffic information includes whether the visibility of the crossing traffic is blocked, whether the speed of the crossing traffic is not greater than zero (e.g., ), whether the vehicle is approaching the intersection, and / or whether crossing traffic is slowing down. Based on the detected traffic information, the vehicle (e.g., TLI generator 514 or side-view TLD system 506) derives the state of the traffic light at the intersection ahead of the vehicle, for example, by determining whether the field of view of crossing traffic is blocked, whether the crossing traffic is stopping or moving, whether the distance between the crossing traffic and the intersection is decreasing or increasing, and / or whether the crossing traffic is slowing down or speeding up.
[0108] In one embodiment, for example, Figure 7As further described in detail in FIG. 5 , a traffic light information (TLI) generator 514 receives both the forward-looking TLD data 503 and the side-looking TLD data 505 and generates traffic light information 515 based on the TLD data 503 and / or the side-looking TLD data 505, which is used, for example, to cross-check the state of a traffic light (e.g., a green state) derived based on the forward-looking TLD data 503 with the state of a traffic light derived based on the side-looking TLD data 505. The perception system 510 then provides the traffic light information 515 to the planning system 520 for further processing, as described above.
[0109] Figure 7 is used for management (with Figure 6 The intersection 602 is the same or similar to the intersection (with Figure 2 or Figure 6 Flow chart of the process 700 of traffic light detection of a vehicle (same or similar to the vehicle 200). Figure 5 The architecture 500 of FIG. 500 describes the process 700. In some embodiments, the process 700 is performed by a computing device (eg, in whole and / or in part, etc.) including at least one processor. The computing device and Figure 3 The computing device may be included in an autonomous system (e.g., Figure 2 In some embodiments, the autonomous system includes a perception system (e.g., Figure 4A The sensing system 402 shown or Figure 5 4 ), a planning system (e.g., the planning system 404 shown in FIG. 4 or the planning system 510 shown in FIG. 5 ), and a planning system (e.g., the planning system 404 shown in FIG. 5 ). Figure 5 The perception system may include a traffic light information generator (e.g., a traffic light information generator). Figure 5 One or more steps of process 700 are performed by the traffic light information generator.
[0110] Block 710 illustrates the use of a forward-looking TLD system (e.g., Figure 5 The forward-looking TLD system 502 includes two forward-looking cameras CAM_M_F 711a and CAM_N_F 711b. Based on one or more forward-looking images captured from the two forward-looking cameras, for example, a forward-looking TLD system (such as Figure 5 502, etc.) or by a sensory system (such as Figure 5 510, etc.) perform traffic light detection (TLD) (712) to derive dominant road segments (e.g., Figure 6The traffic light of the road segment 610 (for example, Figure 6 The state of the traffic light 634b and / or 614a) is shown in Figure 7, wherein the vehicle is approaching an intersection on the road segment or the road segment is associated with the operation of the vehicle. The state of the traffic light is represented by a circle 715.
[0111] In some cases, the state of a traffic light is derived as a red state or a yellow state. In those cases, even if the actual state of the traffic light (e.g., the ground truth state of a real-world traffic light) is a green state, the perception system determines that the state of the traffic light is a red state or a yellow state. The control system of the vehicle operates according to the red state or the yellow state, which causes the vehicle to slow down or otherwise stop. In these cases, undesirable scenarios do not occur (e.g., running a red traffic light does not occur). The perception system reports the derived state of the traffic light (red or yellow state) directly to the planning system (e.g., Figure 5 The planning system 520 of the present invention (716) and the process 700 ends (717). In some cases, the perception system checks whether the derived state of the traffic light is a red state or a yellow state (718), for example, by comparing the derived state of the traffic light (715) with a predetermined state (e.g., red or yellow). If the derived state of the traffic light is a red state or a yellow state, the state of the traffic light reported to the planning system (716) is represented by a circle 719.
[0112] In some cases, the state of the traffic light is derived as a green state. The actual state of the traffic light can be a green state or a red or yellow state. If the forward-looking TLD system fails or fails to function properly, the actual state of the traffic light is a red state 713 and the forward-looking TLD system or the perception system derives the state of the traffic light as a green state 714, which is considered an incorrect TLD or a low-confidence TLD. In those cases, in response to determining that the derived state of the traffic light is not a red state or a yellow state (718), the process 700 enters step 760, which is used to compare the derived state of the traffic light (712) with the state from the side-looking TLD system (e.g., a system that is independent of the forward-looking TLD system) and the side-looking TLD system (e.g., a system that is independent of the forward-looking TLD system). Figure 5 The state of the traffic light derived by the side-view TLD system 506 is cross-checked.
[0113] Block 720 shows one or more steps performed using the side-view TLD system and the perception system. Figure 7 As shown, the side-view TLD system includes a side-view camera system (e.g., Figure 5 508a), LiDAR systems (e.g., Figure 5 508b) and Radar systems (e.g. Figure 5508c), the side view camera system includes left / right side view cameras CAM_F_L 721, CAM_F_R 722, the LiDAR system includes left / right side LiDAR sensors LiDAR_F_L 723, LiDAR_F_R 724, and the Radar system includes left / side side Radar sensors RADAR_F_L 725, RADAR_F_R 726.
[0114] In one embodiment, the perception system is based on sensor data from a side-looking TLD system, for example using a sensor tracking algorithm or a machine learning model such as Figure 4B CNN 420 or Figure 4C and Figure 4D CNN 440, etc.) to perform perception step 730 to infer one or more intersection road segments adjacent to the road segment (e.g., Figure 6 Crossing traffic information (e.g., Figure 6 Crossing traffic within the field of view 654, 656 shown). In one example, the sensor tracking algorithm includes a nearest neighbor algorithm, a probabilistic data association algorithm, a multi-hypothesis tracking algorithm, or an interactive multiple model (IMM). The sensor data includes side view image data from a side view camera system, LiDAR sensor data from a LiDAR system, and / or Radar sensor data from a Radar system indicating crossing traffic information. The perception system determines different types of crossing traffic information at perception step 730, including: whether visibility of crossing traffic is blocked (732), whether the speed of crossing traffic is not greater than zero (734), whether the vehicle is approaching an intersection (736), and whether crossing traffic is decelerating (738). Since these types of crossing traffic information are related to the state of the crossing traffic light, and the state of the crossing traffic light is coordinated with the state of the traffic light to be derived, the state of the traffic light can be determined based on the inferred crossing traffic information.
[0115] The perception system derives the state of the traffic light (740) to obtain a derived state of the traffic light 750. Figure 7As shown, the perception system determines whether the field of view (FOV) of the crossing traffic is blocked (742). If the FOV of the crossing traffic is blocked, the perception system determines the state of the traffic light as red (752). If the FOV of the crossing traffic is not blocked, the perception system further determines whether the rate of the crossing traffic is not greater than zero (e.g., zero or substantially equal to zero) (744). If the rate of the crossing traffic is not greater than zero, the perception system determines the state of the traffic light as green (754). If the rate of the crossing traffic is greater than zero, the perception system further determines whether the distance between the crossing traffic and the intersection is greater than a limit (746). In one example, the distance limit is a predetermined limit in a stop distance table. If the distance is greater than the limit, the perception system determines the state of the traffic light as green (756). If the distance is less than or equal to the limit, the perception system determines whether the rate of the crossing traffic is decreasing (748). If the rate of the crossing traffic is decreasing, the perception system determines the state of the traffic light as green (758). If the rate of crossing traffic is increasing, the perception system determines the state of the traffic light to be red ( 759 ).
[0116] At step 760 , the perception system performs a cross-check by checking whether the state derived from the forward-looking TLD system is the same as (or matches) the state derived from the side-looking TLD system, and determines the traffic light information at the intersection based on the result of the check.
[0117] If the state derived from the side view TLD system (e.g., green state 754, 756, 758) is determined to be the same as the green state derived from the front view TLD system, the perception system determines that the current state of the traffic light at the intersection is the green state. The green state of the traffic light is represented by circle 719 and is reported to the planning system (716).
[0118] If the state derived from the side view TLD system (e.g., red state 752, 759) is determined to be different from the green state derived from the front view TLD system, the perception system determines that the current state of the traffic light at the intersection is the red state (762), which prevents the vehicle from mistakenly or riskily operating in an undesirable manner (such as running a red traffic light, etc.). The red state of the traffic light determined at step 762 is represented by circle 719 and is reported to the planning system at step 716.
[0119] In some embodiments, the perception system determines whether the duration for which the state of the traffic light (due to the failure of the cross check) is assumed to be the red state is greater than a predetermined duration (e.g., 3 seconds). If the duration is not greater than the predetermined duration, the faulty TLD system returns to work, and the process 700 proceeds to the end (717). If the duration is greater than the predetermined duration, this means that the faulty TLD system is still not working properly, and the perception system initiates external support (766), for example, by triggering an alarm signal for manual support to the operator of the vehicle, and then the process 700 proceeds to the end (717).
[0120] Reference now Figure 8 , illustrated is a flow chart of a process 800 for managing traffic light detection, in particular by cross-checking the results of a first traffic light detection (TLD) system with the results of a separate and independent second TLD system. In some embodiments, the process 800 is performed by a computing device (e.g., fully and / or partially, etc.) including at least one processor. The computing device and Figure 3 The computing device may be included in an autonomous system (e.g., Figure 2 Additionally or alternatively, in some embodiments, other devices or groups of devices separate from the autonomous system (e.g., Figure 1 14) (e.g., completely and / or partially, etc.) to perform process 800. The steps in process 800 may be similar to those in Figure 7 The steps in the described process 700 correspond.
[0121] In some embodiments, the autonomous system includes a sensory system (e.g., Figure 4A The sensing system 402 shown or Figure 5 4 ), a planning system (e.g., the planning system 404 shown in FIG. 4 or the planning system 510 shown in FIG. 5 ), and a planning system (e.g., the planning system 404 shown in FIG. 5 ). Figure 5 The perception system may include a traffic light information generator (e.g., a traffic light information generator). Figure 5 The traffic light information generator 514 of the embodiment of the present invention may be used to generate one or more steps of the process 800.
[0122] refer to Figure 8, the autonomous system derives a first state of a traffic light at an intersection that the vehicle is approaching based on first detection data acquired by a first traffic light detection (TLD) system (802). In some embodiments, step 802 corresponds to step 702. The first state of the traffic light can be red, yellow, or green.
[0123] The intersection can be Figure 6 The intersection 602 is shown. In some embodiments, the intersection is connected to multiple road segments (e.g., Figure 6 The segment includes a first path segment that the vehicle is approaching (e.g., Figure 6 610), and traffic lights (e.g., Figure 6 The traffic light 634b) controls the movement of the vehicle at the first path segment. The road segment also includes at least one intersection segment adjacent to the first path segment (e.g., Figure 6 620, 640), and crossing a traffic light (e.g., Figure 6 The traffic lights 644b, 624b) of the at least one intersection segment control the corresponding crossing traffic movement at each intersection segment of the at least one intersection segment. The crossing traffic lights for at least one intersection segment are coordinated with the traffic lights for the first path segment. For example, if the crossing traffic light is red, the traffic light for the first path segment is green; if the crossing traffic light is green, the traffic light for the first path segment is red.
[0124] In some embodiments, the first TLD system includes at least one forward-looking camera (e.g., Figure 7 CAM_M_F 711a and / or CAM_N_F 711b) described in the forward-looking TLD system (e.g., Figure 5 The first detection data acquired by the first TLD system includes at least one forward-looking image of the intersection, which may include, for example, Figure 6 An image of a traffic light in front of the vehicle within the field of view 652 is shown. The autonomous system uses image processing algorithms and machine learning models (e.g., Figure 4B CNN 420 or Figure 4C and Figure 4D At least one of the CNNs 440) is configured to derive a first state of the traffic light based on at least one forward-looking image of the intersection.
[0125] refer to Figure 8 The autonomous system derives a second state of the traffic light at the intersection based on second detection data acquired by a second TLD system independent of the first TLD system (804). The second state can be red or green. In one example, the second TLD system is a side view TLD system (e.g., Figure 5 Side-view TLD system 506).
[0126] In some embodiments, for example, Figure 7 As illustrated, the second TLD system includes at least one side-view camera (e.g., such as Figure 7 CAM_F_L 721, CAM_F_R 722, etc.), at least one LiDAR sensor (e.g., such as Figure 7 left / right side of LIDAR_F_L 723, LIDAR_F_R 724, etc.) and at least one Radar sensor (e.g., such as Figure 7 The second detection data includes crossing traffic (e.g., Figure 6 Side view camera images, LiDAR sensor data and / or Radar sensor data of crossing traffic within the field of view 654, 656 shown.
[0127] In some embodiments, an autonomous system (e.g., using Figure 7 The sensor tracking algorithm exemplified in Figure 4B CNN 420 or Figure 4C and Figure 4D The sensor tracking algorithm may be used to infer the crossing traffic information based on the second detection data, and derive the second state of the traffic light based on the inferred crossing traffic information. For example, the sensor tracking algorithm receives the second detection data as input, and outputs the crossing traffic information or the second state of the traffic light as output.
[0128] The crossing traffic information includes crossing traffic events and / or behaviors on at least one cross section at the intersection. Figure 7 As illustrated in step 730, the crossing traffic information includes at least one of the following items: whether the field of view of a crossing traffic is obscured, the crossing traffic including one or more other vehicles at an intersection section that is different from the section that the vehicle is approaching; whether the crossing traffic is stopped; whether the crossing traffic is approaching the intersection; and whether the crossing traffic is decelerating.
[0129] In some embodiments, for example, Figure 7 As illustrated in step 740, if the field of view of the crossing traffic is obscured, the autonomous system determines the second state of the traffic light to be red; if the field of view of the crossing traffic is not obscured, the autonomous system further determines at least one of the following items: whether the speed of the crossing traffic is not greater than zero (e.g., zero or substantially equal to zero), whether the distance between the crossing traffic and the intersection is decreasing, and whether the crossing traffic is decelerating.
[0130] In some embodiments, if the rate of the crossing traffic is not greater than zero, the autonomous system determines the second state of the traffic light as green; if the rate of the crossing traffic is greater than zero, the autonomous system further determines whether the distance between the crossing traffic and the intersection is greater than a predetermined limit (e.g., a limit listed in a stopping distance table). If the distance is greater than the predetermined limit, the autonomous system determines the second state of the traffic light as green; if the distance is less than or equal to the predetermined limit, the autonomous system determines whether the crossing traffic is decelerating (or whether the rate of the crossing traffic is decreasing). If the crossing traffic is decelerating, the autonomous system determines the second state of the traffic light as green; if the crossing traffic is accelerating (or the rate of the crossing traffic is increasing), the autonomous system determines the second state of the traffic light as red.
[0131] In some embodiments, in response to determining that the distance from the vehicle to the intersection meets, for example, no greater than a predetermined threshold (e.g., 50 meters), the autonomous system initiates at least one of the first TLD system and the second TLD system to detect a traffic light at the intersection.
[0132] Continue to refer Figure 8 The autonomous system determines traffic light information at the intersection based on at least one of (i) the first state and (ii) whether the first state is the same as the second state (806).
[0133] In some embodiments, the autonomous system first determines whether the first state derived from the first TLD system is a specified state (e.g., green). If the first state (e.g., red or yellow) is different from the specified state, the autonomous system bypasses the cross-check of the first state with the second state and determines the traffic light information at the intersection based on the first state. If the first state is the specified state (e.g., green), in order to avoid a malfunction of the first TLD system, the autonomous system continues to cross-check by checking whether the first state derived from the first TLD system is the same (or matches) as the second state derived from the second TLD system, and determines the traffic light information at the intersection based on the result of the check.
[0134] In some embodiments, if the designated state (e.g., green) is determined to be the same as the second state (e.g., green), the autonomous system determines that the current state of the traffic light at the intersection is the designated state. If the designated state (e.g., green) is determined to be different from the second state (e.g., red), the autonomous system determines that the current state of the traffic light at the intersection is the red state, which can prevent the vehicle from operating in an unexpected manner (e.g., running a red traffic light) by mistake or risk.
[0135] In some embodiments, if the duration that the specified state is determined to be different from the second state is greater than a predetermined duration (eg, 3s), the autonomous system initiates external support, such as by triggering an alarm signal for manual support to an operator of the vehicle.
[0136] In some embodiments, the traffic light information includes: the current state of the traffic light at the intersection and the remaining time for the current state of the traffic light at the intersection. In some examples, the autonomous system determines the remaining time for the current state of the traffic light at the intersection based on the time point of the state change of the traffic light immediately before the current state and a predetermined duration (e.g., 20s) for the current state of the traffic light. For example, the rules or protocols for managing traffic lights at intersections set how a series of states of the traffic light are presented and how long each state lasts. The rules or protocols can be stored in a database (e.g., Figure 4A Database 410 or Figure 5 in the map mapping database 501).
[0137] Continue to refer Figure 8 , the autonomous system causes the vehicle to operate according to the determined traffic light information at the intersection (808). For example, if the determined traffic light is red, the vehicle (e.g., the control system) can determine whether to slow down or stop based on the remaining time for the red state and / or the travel time to the intersection. If the determined traffic light is green, the vehicle can determine whether to keep moving or accelerate through the intersection based on the remaining time for the green state and / or the travel time to the intersection.
[0138] Further non-limiting aspects or embodiments are set forth in the following numbered clauses:
[0139] Clause 1: A method comprising: using at least one processor to derive a first state of a traffic light at an intersection that a vehicle is approaching based on first detection data obtained by a first traffic light detection system, i.e., a first TLD system; using the at least one processor to derive a second state of the traffic light at the intersection based on second detection data obtained by a second TLD system independent of the first TLD system; using the at least one processor to determine traffic light information at the intersection based on at least one of the first state and a check result of whether the first state is the same as the second state; and using the at least one processor to cause the vehicle to operate according to the determined traffic light information at the intersection.
[0140] Clause 2: The method according to clause 1 further includes: determining whether the first state is a specified state.
[0141] Clause 3: The method according to Clause 2, comprising: in response to determining that the first state is different from the designated state, determining traffic light information at the intersection based on the first state.
[0142] Clause 4: The method according to Clause 2 includes: in response to determining that the first state is the designated state, checking whether the first state is the same as the second state; and determining the traffic light information at the intersection based on the checking result.
[0143] Clause 5: The method according to Clause 4, wherein determining the traffic light information at the intersection based on the inspection result includes: in response to determining that the designated state is the same as the second state, determining that the current state of the traffic light at the intersection is the designated state.
[0144] Clause 6: The method according to Clause 4, wherein determining the traffic light information at the intersection based on the inspection result includes: in response to determining that the designated state is different from the second state, determining that the current state of the traffic light at the intersection is a red state.
[0145] Clause 7: The method according to clause 6 further includes: in response to determining that the following duration is greater than a predetermined duration, initiating external support using the at least one processor, wherein the duration is determined to be a duration that the specified state is different from the second state.
[0146] Clause 8: A method according to any one of clauses 1 to 7, wherein the first TLD system includes at least one forward-looking camera, and wherein the second TLD system includes at least one of at least one side-looking camera, at least one LiDAR sensor, and at least one Radar sensor.
[0147] Clause 9: A method according to clause 8, wherein the first detection data includes at least one forward-looking image of the intersection, and wherein deriving the first state of the traffic light at the intersection based on the first detection data includes: using at least one of an image processing algorithm and a machine learning model to derive the first state of the traffic light based on at least one forward-looking image of the intersection.
[0148] Clause 10: The method according to clause 8 or 9, wherein deriving the second state of the traffic light at the intersection based on the second detection data comprises: inferring crossing traffic information based on the second detection data, and deriving the second state of the traffic light based on the inferred crossing traffic information.
[0149] Clause 11: A method according to clause 10, wherein the crossing traffic information includes at least one of the following items: whether the field of view of crossing traffic is obscured, the crossing traffic including one or more other vehicles at an intersection section different from the section that the vehicle is approaching; whether the crossing traffic is stopped; whether the distance between the crossing traffic and the vehicle is greater than a limit; and whether the crossing traffic is decelerating (or whether the speed of the crossing traffic is decreasing).
[0150] Clause 12: A method according to clause 11, wherein the second TLD system includes at least one of at least one side-view camera, at least one LiDAR sensor, and at least one Radar sensor, and wherein inferring crossing traffic information based on the second detection data includes at least one of the following items: inferring whether the field of view of the crossing traffic is blocked based on at least one of the detection data of the at least one side-view camera and the detection data of the at least one LiDAR sensor, inferring whether the crossing traffic is stopped based on at least one of the detection data of the at least one side-view camera, the detection data of the at least one LiDAR sensor, and the detection data of the at least one Radar sensor, inferring whether the crossing traffic is approaching the intersection based on at least one of the detection data of the at least one LiDAR sensor and the detection data of the at least one Radar sensor, and inferring whether the crossing traffic is decelerating based on at least one of the detection data of the at least one LiDAR sensor and the at least one Radar sensor.
[0151] Clause 13: A method according to clause 11 or 12, wherein deriving the second state of the traffic light based on the inferred crossing traffic information includes: determining the second state of the traffic light to be red based on a determination that the field of view of the crossing traffic is blocked; and determining at least one of the following items based on a determination that the field of view of the crossing traffic is not blocked: whether the speed of the crossing traffic is not greater than zero, whether the distance between the crossing traffic and the intersection is decreasing, and whether the crossing traffic is decelerating.
[0152] Clause 14: A method according to clause 13, wherein deriving the second state of the traffic light based on the inferred crossing traffic information includes at least one of the following items: determining the second state of the traffic light as green based on a determination that the speed of the crossing traffic is not greater than zero; determining whether the distance between the crossing traffic and the intersection is decreasing based on a determination that the speed of the crossing traffic is greater than zero; determining the second state of the traffic light as green based on a determination that the distance is decreasing; determining whether the crossing traffic is decelerating based on a determination that the distance is increasing; determining the second state of the traffic light as green based on a determination that the crossing traffic is decelerating; and determining the second state of the traffic light as red based on a determination that the crossing traffic is accelerating.
[0153] Clause 15: A method according to any one of clauses 1 to 14, wherein the intersection is associated with multiple road segments, the multiple road segments comprising: a first path segment that the vehicle is approaching, the traffic light being used to control the movement of the vehicle at the first path segment; and at least one intersection segment adjacent to the first path segment, a crossing traffic light being used to control the corresponding crossing traffic movement at each intersection segment in the at least one intersection segment, the crossing traffic light being coordinated with the traffic light for the first path segment.
[0154] Clause 16: The method according to any one of clauses 1 to 15, wherein the traffic light information comprises: a current state of the traffic light at the intersection and a remaining time for the current state of the traffic light at the intersection.
[0155] Clause 17: A method according to Clause 16, wherein determining the traffic light information at the intersection includes: determining the remaining time for the current state of the traffic light at the intersection based on the time point of the state change of the traffic light immediately before the current state and the predetermined duration for the current state of the traffic light.
[0156] Clause 18: The method according to any one of clauses 1 to 17 further includes: in response to determining that the distance from the vehicle to the intersection meets a predetermined threshold, using the at least one processor to initiate at least one of the first TLD system and the second TLD system to detect the traffic light at the intersection.
[0157] Clause 19: A system comprising: at least one processor, and at least one non-transitory storage medium storing instructions, which, when executed by the at least one processor, cause the at least one processor to perform the method according to any one of clauses 1 to 18.
[0158] Clause 20: At least one non-transitory storage medium storing instructions which, when executed by at least one processor, cause the at least one processor to perform the method of any one of clauses 1 to 18.
[0159] In the previous description, aspects and embodiments of the present disclosure have been described with reference to many specific details, which may vary depending on the implementation. Therefore, the description and the accompanying drawings should be regarded as illustrative, not restrictive. The only and exclusive indication of the scope of the invention, and the applicant's expectation that the content of the scope of the invention is the literal and equivalent scope of the claims issued from this application in the specific form of the claims, including any subsequent amendments. Any definition of the terms used to be included in such claims that are clearly set forth herein should be based on the meaning of such terms as used in the claims. In addition, when the term "also includes" is used in the previous description or the attached claims, the following of the phrase may be an additional step or entity, or a sub-step / sub-entity of the previously described step or entity.
Claims
1. A method comprising: Using at least one processor, deriving a first state of a traffic light at an intersection approached by a vehicle based on first detection data acquired by a first traffic light detection system, namely, a first TLD system; deriving, using the at least one processor, a second state of the traffic light at the intersection based on second detection data acquired by a second TLD system independent of the first TLD system; determining, using the at least one processor, traffic light information at the intersection based on at least one of the first state and a result of a check on whether the first state is the same as the second state; as well as Using the at least one processor, the vehicle is caused to operate according to the determined traffic light information at the intersection.
2. The method according to claim 1, further comprising: It is determined whether the first state is a specified state.
3. The method according to claim 2, comprising: In response to determining that the first state is different from the designated state, traffic light information at the intersection is determined based on the first state.
4. The method according to claim 2, comprising: In response to determining that the first state is the specified state, checking whether the first state is the same as the second state; as well as Traffic light information at the intersection is determined based on the checking result.
5. The method according to claim 4, wherein: Determining the traffic light information at the intersection based on the inspection result includes: In response to determining that the designated state is the same as the second state, it is determined that the current state of the traffic light at the intersection is the designated state.
6. The method according to claim 4, wherein: Determining the traffic light information at the intersection based on the inspection result includes: In response to determining that the designated state is different from the second state, it is determined that the current state of the traffic light at the intersection is a red state.
7. The method according to claim 6, further comprising: In response to a determination that a duration during which the specified state is determined to be different from the second state is greater than a predetermined duration, initiating external support using the at least one processor.
8. The method according to any one of claims 1 to 7, wherein: The first TLD system includes at least one forward-looking camera, and Wherein, the second TLD system includes at least one of at least one side-view camera, at least one LiDAR sensor and at least one Radar sensor.
9. The method according to claim 8, wherein: The first detection data includes at least one front-view image of the intersection, and Wherein, deriving the first state of the traffic light at the intersection according to the first detection data includes: A first state of the traffic light is derived based on at least one forward-view image of the intersection using at least one of an image processing algorithm, a sensor tracking algorithm, and a machine learning model.
10. The method according to claim 8 or 9, wherein: Deriving the second state of the traffic light at the intersection according to the second detection data comprises: inferring crossing traffic information based on the second detection data, and A second state of the traffic light is derived based on the inferred crossing traffic information.
11. The method according to claim 10, wherein: The crossing traffic information includes at least one of the following items: whether the field of view of crossing traffic is obstructed, the crossing traffic including one or more other vehicles at a different intersection than the one the vehicle is approaching, Whether the crossing traffic is stopped, whether the distance between the crossing traffic and the intersection is greater than a predetermined limit, and Whether the rate of crossing traffic is decreasing.
12. The method according to claim 11, wherein: The second TLD system includes at least one of at least one side-view camera, at least one LiDAR sensor, and at least one Radar sensor, and Inferring the crossing traffic information according to the second detection data includes at least one of the following items: inferring whether the field of view of the crossing traffic is obstructed based on at least one of the detection data of the at least one side-looking camera and the detection data of the at least one LiDAR sensor, inferring whether the crossing traffic is stopped based on at least one of the detection data of the at least one side-view camera, the detection data of the at least one LiDAR sensor, and the detection data of the at least one Radar sensor, inferring whether the crossing traffic is approaching the intersection based on at least one of the detection data of the at least one LiDAR sensor and the detection data of the at least one Radar sensor, and Whether the crossing traffic is decelerating is inferred based on at least one of the detection data of the at least one LiDAR sensor and the at least one Radar sensor.
13. The method according to claim 11 or 12, wherein: Deriving a second state of the traffic light based on the inferred crossing traffic information includes: Determining the second state of the traffic light to be red based on determining that the field of view of the crossing traffic is blocked; and Based on determining that the field of view of the crossing traffic is not obstructed, at least one of: whether a speed of the crossing traffic is not greater than zero, whether a distance between the crossing traffic and the intersection is greater than a predetermined limit, and whether a speed of the crossing traffic is decreasing is determined.
14. The method according to claim 13, wherein: Deriving the second state of the traffic light based on the inferred crossing traffic information includes at least one of the following: determining the second state of the traffic light to be green based on determining that the velocity of the crossing traffic is not greater than zero, determining whether a distance between the crossing traffic and the intersection is greater than the predetermined limit based on determining that the speed of the crossing traffic is greater than zero, determining the second state of the traffic light to be green based on determining that the distance is greater than the predetermined limit, based on determining that the distance is less than or equal to the predetermined limit, determining whether a rate of the crossing traffic is decreasing, determining the second state of the traffic light to be green based on determining that the speed of the crossing traffic is decreasing, and Based on determining that the velocity of the crossing traffic is increasing, the second state of the traffic light is determined to be red.
15. The method according to any one of claims 1 to 14, wherein: The intersection is associated with a plurality of road segments, the plurality of road segments comprising: a first path segment that the vehicle is approaching, the traffic light being used to control the movement of the vehicle at the first path segment, and At least one intersection adjacent to the first path segment has a crossing traffic light for controlling corresponding crossing traffic movement at each of the at least one intersection, the crossing traffic light being coordinated with the traffic light for the first path segment.
16. The method according to any one of claims 1 to 15, wherein: The traffic light information includes: A current state of the traffic light at the intersection and a remaining time for the current state of the traffic light at the intersection.
17. The method according to claim 16, wherein: Determining the traffic light information at the intersection includes: A remaining time for the current state of the traffic light at the intersection is determined based on a time point of a change in the state of the traffic light immediately before the current state and a predetermined duration for the current state of the traffic light.
18. The method according to any one of claims 1 to 17, further comprising: In response to determining that the distance from the vehicle to the intersection satisfies a predetermined threshold, at least one of the first TLD system and the second TLD system is initiated, using the at least one processor, to detect the traffic light at the intersection.
19. A system comprising: at least one processor, and At least one non-transitory storage medium storing instructions, which, when executed by the at least one processor, cause the at least one processor to perform the method according to any one of claims 1 to 18.
20. At least one non-transitory storage medium storing instructions which, when executed by at least one processor, cause the at least one processor to perform the method according to any one of claims 1 to 18.