System and method for model training and on-vehicle validation using autonomous vehicles

By recording and labeling audio data in autonomous vehicles and combining it with visual sensor data, improved labeled audio data is generated for training machine learning algorithms, solving the problems of inaccurate motion planning and inaccurate obstacle recognition, and improving the accuracy and safety of autonomous driving.

CN114764523BActive Publication Date: 2025-09-09BAIDU USA LLC
View PDF 2 Cites 0 Cited by

Patent Information

Application Number
CN202111642930.3
Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Priority Date
2021-01-12
Filing Date
2021-12-29
Publication Date
2025-09-09
Estimated Expiration
2041-12-29

AI Technical Summary

Technical Problem

The motion planning and control of existing autonomous vehicles lack consideration of the characteristic differences between different types of vehicles, resulting in inaccurate motion planning. In addition, obstacle and sound source identification relies on inaccurate manually labeled data, which affects model accuracy.

Method used

By recording and labeling audio data in autonomous vehicles, the audio samples are used to generate improved labeled data, which is combined with visual sensor data to determine the location and direction of objects and used to train machine learning algorithms to identify sound sources and obstacles.

Benefits of technology

This improves the accuracy of sound source and obstacle identification in autonomous vehicles, ensures more accurate generated models, and improves the safety and efficiency of autonomous driving.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN114764523B_ABST
    Figure CN114764523B_ABST
Patent Text Reader

Abstract

Disclosed are systems and methods for generating labeled audio data using an autonomous vehicle (ADV) while the ADV is operating in a driving environment and for performing on-board verification of the labeled audio data. The method includes recording sounds emitted by objects within the driving environment of the ADV and converting the recorded sounds into audio samples. The method further includes labeling the audio samples and improving the labeled audio samples to produce improved labeled audio data. The improved labeled audio data is used to subsequently train a machine learning algorithm to identify sound sources during autonomous driving of the ADV. The method further includes generating a performance profile for the improved labeled audio data based on at least the audio samples, the position of the object, and the relative orientation of the object. The position of the object and the relative orientation of the object are determined by a perception system of the ADV.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] Embodiments of the present disclosure generally relate to operating autonomous driving vehicles (ADVs). More particularly, embodiments of the present disclosure relate to utilizing audio recordings of autonomous driving vehicles for model training and on-board validation. Background Art

[0002] Vehicles operating in autonomous mode (e.g., without human drivers) can relieve occupants, particularly the driver, of some driving-related responsibilities. When operating in autonomous mode, the vehicle can navigate to various locations using onboard sensors, allowing the vehicle to travel with minimal human interaction or in some cases without any passengers.

[0003] Motion planning and control are critical operations in autonomous driving. However, traditional motion planning primarily estimates the difficulty of completing a given path based on curvature and velocity, without considering the differences in characteristics between different types of vehicles. Applying the same motion planning and control to all vehicle types can be inaccurate and unsmooth in some situations.

[0004] Furthermore, motion planning and control operations often require sensing surrounding obstacles or objects, as well as listening to or detecting sound sources within the driving environment. Therefore, obstacle recognition and sound source identification require labeled data (e.g., sensor data) to serve as training and testing data for machine learning models.

[0005] Unfortunately, data labeling is done manually, and due to the inherent flaws of humans, manually labeled data is not very accurate, which in turn affects the accuracy of the model. Summary of the Invention

[0006] In a first aspect, a method for generating labeled audio data using an autonomous vehicle (ADV) and performing on-board verification of the labeled audio data while the ADV is operating within a driving environment is provided, the method comprising:

[0007] Recording sounds made by objects in a driving environment of the ADV and converting the recorded sounds into audio samples;

[0008] Labeling the audio samples, and improving the labeled audio samples to generate improved labeled audio data, wherein the improved labeled audio data is used for subsequent training of a machine learning algorithm to identify sound sources during autonomous driving of an ADV; and

[0009] A performance profile of the improved labeled audio data is generated based on at least the audio samples, the positions of the objects, and the relative directions of the objects, wherein the positions of the objects and the relative directions of the objects are determined by a perception system of the ADV.

[0010] In a second aspect, a computer-implemented method is provided for performing on-board verification of labeled audio data with an autonomous vehicle (ADV) while the ADV is operating within a driving environment, the method comprising:

[0011] Recording sounds made by obstacles within a driving environment of the ADV to create audio samples;

[0012] Determining the location of obstacles and the relative direction of obstacles based on sensor data provided by the ADV's vision sensor; and

[0013] The audio samples, the locations of obstacles, and the relative directions of the obstacles are used to generate a performance profile of improved labeled audio data, wherein the improved labeled audio data is generated by labeling the audio samples and improving the labeled audio samples, and is used to subsequently train a machine learning algorithm to identify sound sources during autonomous driving of an ADV.

[0014] In a third aspect, a system for generating labeled audio data using an autonomous vehicle (ADV) and performing on-board verification of the labeled audio data while the ADV is operating in a driving environment is provided, comprising:

[0015] processor; and

[0016] A memory coupled to the processor and storing instructions, which, when executed by the processor, causes the processor to perform the operations of the method according to the first aspect.

[0017] In a fourth aspect, a system for performing in-vehicle verification of labeled audio data is provided, comprising:

[0018] processor; and

[0019] A memory coupled to the processor and storing instructions, which, when executed by the processor, causes the processor to perform the operations of the method according to the second aspect.

[0020] In a fifth aspect, a non-transitory machine-readable medium having instructions stored therein is provided, wherein the instructions, when executed by a processor, cause the processor to perform the operations of the method of the first aspect or the method of the second aspect.

[0021] In a sixth aspect, a computer program product is provided, comprising a computer program, which, when executed by a processor, causes the processor to perform the operations of the method according to the first aspect or the method according to the second aspect.

[0022] According to the present disclosure, the accuracy of the generated model can be guaranteed. BRIEF DESCRIPTION OF THE DRAWINGS

[0023] Embodiments of the present disclosure are illustrated by way of example and not limitation in the figures of the accompanying drawings in which like references indicate similar elements.

[0024] Figure 1 is a block diagram illustrating a networking system according to one embodiment.

[0025] Figure 2 is a block diagram illustrating an example of an autonomous driving vehicle (ADV) according to one embodiment.

[0026] Figures 3A-3B is a block diagram illustrating an example of an autonomous driving system for use with an autonomous driving vehicle according to one embodiment.

[0027] Figure 4 is a block diagram illustrating a system for audio recording and in-vehicle verification using an autonomous vehicle, according to one embodiment.

[0028] Figure 5 is a diagram illustrating an example driving scenario using a system for audio recording and in-vehicle verification, according to one embodiment.

[0029] Figure 6 is a flow chart illustrating a method of generating labeled audio data and in-vehicle verification of the labeled audio data according to one embodiment.

[0030] Figure 7 is a flow chart illustrating a method for in-vehicle verification of labeled audio data according to one embodiment. DETAILED DESCRIPTION

[0031] Various embodiments and aspects of the present disclosure will be described with reference to the details discussed below, and the accompanying drawings will illustrate various embodiments. The following description and the accompanying drawings are illustrative of the present disclosure and should not be construed as limiting the present disclosure. Many specific details are described to provide a comprehensive understanding of the various embodiments of the present disclosure. However, in some cases, in order to provide a brief discussion of the embodiments of the present disclosure, well-known or conventional details are not described.

[0032] Reference in the specification to "one embodiment" or "an embodiment" means that a particular feature, structure, or characteristic described in connection with the embodiment can be included in at least one embodiment of the present disclosure. The appearances of the phrase "in one embodiment" in various places in the specification are not necessarily all referring to the same embodiment.

[0033] According to one aspect, a method for generating labeled audio data using an autonomous vehicle (ADV) and performing on-board verification of the labeled audio data while the ADV is operating in a driving environment is described. The method includes recording sounds emitted by objects within the driving environment of the ADV and converting the recorded sounds into audio samples. The method further includes labeling the audio samples and improving the labeled audio samples to produce improved labeled audio data. The improved labeled audio data is used to subsequently train a machine learning algorithm to identify sound sources during autonomous driving of the ADV. The method further includes generating a performance profile of the improved labeled audio data based on at least the audio samples, the position of the object, and the relative direction of the object. The position of the object and the relative direction of the object are determined by a perception system of the ADV.

[0034] According to another aspect, a method for performing on-board verification of labeled audio data using an ADV while the ADV is operating within a driving environment is described. The method includes recording sounds emitted by obstacles within the driving environment of the ADV to create audio samples. The method further includes determining the location of the obstacle and the relative direction of the obstacle based on sensor data provided by a visual sensor of the ADV. The method further includes generating a performance profile of the improved labeled audio data using the audio samples, the location of the obstacle, and the relative direction of the obstacle. The improved labeled audio data is generated by labeling the audio samples and improving the labeled audio samples. The improved labeled audio data is also used to subsequently train a machine learning algorithm to identify sound sources during autonomous driving of the ADV.

[0035] Figure 1 is a block diagram illustrating an autonomous driving network configuration according to one embodiment of the present disclosure. Figure 1 , a network configuration 100 includes an autonomous vehicle (ADV) 101, which can be communicatively coupled to one or more servers 103-104 via a network 102. Although one ADV is shown, multiple ADVs can be coupled to each other and / or to the servers 103-104 via the network 102. The network 102 can be any type of network, such as a local area network (LAN), a wide area network (WAN) such as the Internet, a cellular network, a satellite network, or a combination thereof, wired or wireless. The servers 103-104 can be any type of server or server cluster, such as a web or cloud server, an application server, a back-end server, or a combination thereof. The servers 103-104 can be data analysis servers, content servers, traffic information servers, map and point of interest (MPOI) servers, or location servers, etc.

[0036] An ADV refers to a vehicle that can be configured to operate in an autonomous mode, in which the vehicle navigates through an environment with little or no driver input. Such an ADV may include a sensor system having one or more sensors configured to detect information about the environment in which the vehicle operates. The vehicle and its associated controller use the detected information to navigate through the environment. The ADV 101 can operate in a manual mode, a fully autonomous mode, or a partially autonomous mode.

[0037] In one embodiment, the ADV 101 includes, but is not limited to, an autonomous driving system (ADS) 110, a vehicle control system 111, a wireless communication system 112, a user interface system 113, and a sensor system 115. The ADV 101 may further include certain common components included in an ordinary vehicle, such as an engine, wheels, a steering wheel, a transmission, etc., which can be controlled by the vehicle control system 111 and / or the ADS 110 using various communication signals and / or commands, such as, for example, acceleration signals or commands, deceleration signals or commands, steering signals or commands, braking signals or commands, etc.

[0038] Components 110-115 can be communicatively coupled to each other via an interconnect, a bus, a network, or a combination thereof. For example, components 110-115 can be communicatively coupled to each other via a controller area network (CAN) bus. The CAN bus is a vehicle bus standard designed to allow microcontrollers and devices to communicate with each other in applications without a host computer. It is a message-based protocol originally designed for multi-way electrical wiring within vehicles, but is also used in many other environments.

[0039] Now refer to Figure 2In one embodiment, the sensor system 115 includes, but is not limited to, one or more cameras 211, a global positioning system (GPS) unit 212, an inertial measurement unit (IMU) 213, a radar unit 214, and a light detection and ranging (LIDAR) unit 215. The GPS system 212 may include a transceiver operable to provide information regarding the location of the ADV. The IMU unit 213 may sense the position and orientation changes of the ADV based on inertial acceleration. The radar unit 214 may represent a system that uses radio signals to sense objects within the ADV's local environment. In some embodiments, in addition to sensing objects, the radar unit 214 may additionally sense the speed and / or heading of objects. The LIDAR unit 215 may use lasers to sense objects in the ADV's environment. The LIDAR unit 215 may include one or more laser sources, a laser scanner, and one or more detectors, among other system components. The camera 211 may include one or more devices to capture images of the environment surrounding the ADV. The camera 211 may be a static camera and / or a video camera. The camera may be mechanically movable, for example, by mounting the camera on a rotating and / or tilting platform.

[0040] The sensor system 115 may also include other sensors such as a sonar sensor, an infrared sensor, a steering sensor, a throttle sensor, a brake sensor, and an audio sensor (e.g., a microphone). The audio sensor may be configured to capture sounds from the environment surrounding the ADV. The steering sensor may be configured to sense the steering angle of the steering wheel, the wheels of the vehicle, or a combination thereof. The throttle sensor and the brake sensor may sense the throttle position and brake position of the vehicle, respectively. In some cases, the throttle sensor and the brake sensor may be integrated into an integrated throttle / brake sensor.

[0041] In one embodiment, the vehicle control system 111 includes, but is not limited to, a steering unit 201, a throttle unit 202 (also referred to as an acceleration unit), and a brake unit 203. The steering unit 201 is used to adjust the direction or heading of the vehicle. The throttle unit 202 is used to control the speed of the motor or engine, which in turn controls the speed and acceleration of the vehicle. The brake unit 203 decelerates the vehicle by providing friction to slow down the wheels or tires of the vehicle. Note that Figure 2 The components shown may be implemented in hardware, software, or a combination thereof.

[0042] Return Reference Figure 1, the wireless communication system 112 allows communication between the ADV 101 and external systems, such as devices, sensors, another vehicle, etc. For example, the wireless communication system 112 can wirelessly communicate with one or more devices directly or via a communication network, such as wirelessly communicating with servers 103-104 via the network 102. The wireless communication system 112 can use any cellular communication network or wireless local area network (WLAN), such as using WiFi to communicate with another component or system. The wireless communication system 112 can communicate directly with a device (e.g., a passenger's mobile device, a display device, a speaker within the vehicle 101) using, for example, an infrared link, Bluetooth, etc. The user interface system 113 can be part of a peripheral device implemented within the vehicle 101, including, for example, a keyboard, a touch screen display device, a microphone, and a speaker, etc.

[0043] Some or all functions of the ADV 101 may be controlled or managed by the ADS 110, particularly when operating in autonomous driving mode. The ADS 110 includes the necessary hardware (e.g., processor, memory, storage device) and software (e.g., operating system, planning and routing programs) to receive information from the sensor system 115, the control system 111, the wireless communication system 112, and / or the user interface system 113, process the received information, plan a route or path from a starting point to a destination point, and then drive the vehicle 101 based on the planning and control information. Alternatively, the ADS 110 may be integrated with the vehicle control system 111.

[0044] For example, a passenger may specify a starting location and destination for a trip, for example, via a user interface. ADS 110 obtains data related to the trip. For example, ADS 110 may obtain location and route information from an MPOI server, which may be part of servers 103-104. The location server provides location services, and the MPOI server provides map services and point of interest (POIs) for certain locations. Alternatively, this location and MPOI information may be cached locally in a persistent storage device of ADS 110.

[0045] As the ADV 101 moves along the route, the ADS 110 may also obtain real-time traffic information from a traffic information system or server (TIS). Note that the servers 103-104 may be operated by a third-party entity. Alternatively, the functionality of the servers 103-104 may be integrated with the ADS 110. Based on real-time traffic information, MPOI information, and location information, as well as real-time local environment data (e.g., obstacles, objects, nearby vehicles) detected or sensed by the sensor system 115, the ADS 110 may plan an optimal route and, for example, drive the vehicle 101 according to the planned route via the control system 111 to safely and efficiently reach a designated destination.

[0046] The server 103 may be a data analysis system that performs data analysis services for various clients. In one embodiment, the data analysis system 103 includes a machine learning engine 122. Based on the improved labeled audio data 126 (described in more detail below), the machine learning engine 122 generates or trains a set of rules, algorithms, and / or prediction models 124 for various purposes, such as identifying sound sources for motion planning and control. The algorithm 124 can then be uploaded to the ADV for real-time use during autonomous driving. As described in more detail below, the improved labeled audio data 126 may include, but is not limited to, audio samples of sound sources, one or more locations of sound sources, directions of sound sources, audio sample identifiers (IDs), and the like.

[0047] Figure 3A and 3B is a block diagram illustrating an example of an autonomous driving system for use with an ADV according to one embodiment. System 300 may be implemented as Figure 1 A portion of the ADV 101 includes, but is not limited to, the ADS 110, the control system 111, and the sensor system 115. Figures 3A-3B The ADS 110 includes, but is not limited to, a positioning module 301 , a perception module 302 , a prediction module 303 , a decision module 304 , a planning module 305 , a control module 306 , a routing module 307 , an audio recorder 308 , a manual audio data labeling module 309 , an automatic improvement module 310 , and a summarization module 311 .

[0048] Some or all of the modules 301-311 may be implemented in software, hardware, or a combination thereof. For example, these modules may be installed in permanent storage 352, loaded into memory 351, and executed by one or more processors (not shown). Note that some or all of these modules may be communicatively coupled to Figure 2Some or all modules of the vehicle control system 111 may be integrated with or integrated with the vehicle control system 111. Some modules in the modules 301-311 may be integrated together as integrated modules.

[0049] The positioning module 301 determines the current location of the ADV 300 (e.g., using the GPS unit 212) and manages any data related to the user's itinerary or route. The positioning module 301 (also known as the map and route module) manages any data related to the user's itinerary or route. The user can log in and specify the starting location and destination of the itinerary, for example, via a user interface. The positioning module 301 communicates with other components of the ADV 300, such as map and route information 311, to obtain data related to the itinerary. For example, the positioning module 301 can obtain location and route information from a location server and a map and point of interest (MPOI) server. The location server provides location services, and the MPOI server provides map services and POIs for certain locations, which can be cached as part of the map and route data 311. As the ADV 300 moves along the route, the positioning module 301 can also obtain real-time traffic information from a traffic information system or server.

[0050] Based on the sensor data provided by the sensor system 115 and the positioning information obtained by the positioning module 301, the perception module 302 determines the perception of the surrounding environment. The perception information can represent the situation around the vehicle that the driver is driving as perceived by an average driver. The perception can include lane configuration, traffic light signals, the relative position of another vehicle, pedestrians, buildings, crosswalks or other traffic-related signs (e.g., stop signs, yield signs), etc. The lane configuration includes information describing one or more lanes, such as the shape of the lane (e.g., straight or curved), the width of the lane, the number of lanes in the road, one-way or two-way lanes, merging or separating lanes, exit lanes, etc.

[0051] The perception module 302 may include a computer vision system or functionality of a computer vision system for processing and analyzing images captured by one or more cameras to identify objects and / or features in the ADV environment. Objects may include, for example, traffic signals, roadway boundaries, other vehicles, pedestrians, and / or obstacles. The computer vision system may utilize object recognition algorithms, video tracking, and other computer vision techniques. In some embodiments, the computer vision system may map the environment, track objects, and estimate the speed of objects, among other things. The perception module 302 may also detect objects based on other sensor data provided by other sensors, such as radar and / or LIDAR.

[0052] For each object, prediction module 303 predicts how the object will behave in the environment. Prediction is performed based on sensory data regarding the driving environment at that point in time, given a set of map / route information 311 and traffic regulations 312. For example, if the object is a vehicle traveling in the opposite direction and the current driving environment includes an intersection, prediction module 303 will predict whether the vehicle is likely to move straight ahead or turn. If the sensory data indicates that the intersection does not have a traffic light, prediction module 303 may predict that the vehicle may have to come to a complete stop before entering the intersection. If the sensory data indicates that the vehicle is currently in a left-turn-only lane or a right-turn-only lane, prediction module 303 may predict that the vehicle is more likely to make a left or right turn, respectively.

[0053] For each object, decision module 304 makes a decision about how to handle the object. For example, given a particular object (e.g., another vehicle in an intersecting path) and metadata describing the object (e.g., speed, direction, steering angle), decision module 304 determines how to encounter the object (e.g., overtake, yield, stop, pass). Decision module 304 may make these decisions based on a set of rules, such as traffic rules or driving rules 312, which may be stored in persistent storage 352.

[0054] The routing module 307 is configured to provide one or more routes or paths from a starting point to a destination point. For a given itinerary from a starting location to a destination location, for example, received from a user, the routing module 307 obtains route and map information 311 and determines all possible routes or paths from the starting location to the destination location. The routing module 307 can generate a reference line in the form of a topographic map for each route it determines from the starting location to the destination location. A reference line is an ideal route or path free of any interference from other factors, such as other vehicles, obstacles, or traffic conditions. In other words, if there are no other vehicles, pedestrians, or obstacles on the road, the ADV should accurately or closely follow the reference line. The topographic map is then provided to the decision module 304 and / or planning module 305. The decision module 304 and / or planning module 305 examines all possible routes to select and modify one of the optimal routes based on other data provided by other modules (such as traffic conditions from the positioning module 301, the driving environment perceived by the perception module 302, and traffic conditions predicted by the prediction module 303). Depending on the particular driving circumstances at a point in time, the actual path or route used to control the ADV may be close to or different from the reference line provided by the routing module 307 .

[0055] Based on the decision for each perceived object, the planning module 305 plans a path or route for the ADV and driving parameters (e.g., distance, speed, and / or steering angle) using the reference line provided by the routing module 307 as a basis. That is, for a given object, the decision module 304 decides what to do with the object, while the planning module 305 determines how to do it. For example, for a given object, the decision module 304 may decide to pass through the object, while the planning module 305 may determine whether to pass through on the left or right side of the object. Planning and control data are generated by the planning module 305 and include information describing how the vehicle 300 will move in the next movement cycle (e.g., the next route / path segment). For example, the planning and control data may instruct the vehicle 300 to move at a speed of 30 miles per hour (mph) for 10 meters and then change to the right lane at a speed of 25 mph.

[0056] Based on the planning and control data, the control module 306 controls and drives the ADV according to the route or path defined by the planning and control data by sending appropriate commands or signals to the vehicle control system 111. The planning and control data includes sufficient information to drive the vehicle from a first point on the route or path to a second point using appropriate vehicle settings or driving parameters (e.g., throttle, brake, steering commands) at different points in time along the route or path.

[0057] In one embodiment, the planning phase is performed over multiple planning cycles (also referred to as driving cycles, such as within each time interval of 100 milliseconds (ms)). For each planning cycle or driving cycle, one or more control commands are issued based on the planning and control data. That is, for every 100 ms, the planning module 305 plans the next route segment or path segment, for example, including the target location and the time required for the ADV to reach the target location. Alternatively, the planning module 305 may also specify a specific speed, direction and / or steering angle, etc. In one embodiment, the planning module 305 plans the route segment or path segment for the next predetermined time period, such as 5 seconds. For each planning cycle, the planning module 305 plans the target position for the current cycle (e.g., the next 5 seconds) based on the target position planned in the previous cycle. The control module 306 then generates one or more control commands (e.g., throttle, brake, steering control commands) based on the planning and control data for the current cycle.

[0058] Note that the decision module 304 and the planning module 305 can be integrated into an integrated module. The decision module 304 / planning module 305 may include a navigation system or the functionality of a navigation system to determine a driving path for the ADV. For example, the navigation system may determine a series of speeds and directional headings to influence the movement of the ADV along a path that substantially avoids perceived obstacles while generally moving the ADV along a roadway-based path to the final destination. The destination may be set based on user input via the user interface system 113. The navigation system may dynamically update the driving path while the ADV is operating. The navigation system may incorporate data from a GPS system and one or more maps to determine a driving path for the ADV.

[0059] Additional references Figure 4 , Figure 4 is a block diagram illustrating a system for audio recording and in-vehicle verification using an autonomous vehicle, according to one embodiment. An audio recorder 308 can communicate with an audio sensor 411 (e.g., a microphone) of a sensor system 115 to record or capture sounds from the environment surrounding the autonomous vehicle. For example, a user of the autonomous vehicle can activate (turn on) the audio recorder 308 via the user interface system 113. In response to user input, the audio recorder 308 can record sounds (e.g., sirens) emitted by obstacles within the autonomous vehicle's driving environment (e.g., emergency vehicles such as police cars, ambulances, fire trucks, etc.) and convert them into audio samples / events 313 (e.g., audio files in a suitable format). When sufficient audio data has been recorded (e.g., a certain data size has been reached or a certain amount of time has elapsed), the user can deactivate (turn off) the audio recorder 308, and the audio samples / events 313 can be stored in a persistent storage device 352.

[0060] Using the captured or recorded audio sample 313, a user can invoke an audio data labeling module 309 (e.g., a data labeling tool or application) to manually label the audio sample 313. For example, the audio data labeling module 309 can be used to label or identify the audio sample 313 with an audio sample ID, one or more locations associated with the audio sample 313 (e.g., the location of a sound source or obstacle), a direction associated with the audio sample 313 (e.g., the relative direction of the sound source or obstacle), etc., and generate labeled audio data 314, which can be stored on a persistent storage device 352 or a remote server (e.g., server 103). Thus, the labeled audio data 314 may include the audio sample 313, the audio sample ID, one or more locations associated with the audio sample 313, a direction associated with the audio sample 313, etc. The automatic improvement module 310 can automatically improve the labeled audio data 314 for machine learning. That is, the labeled audio data 314 may need to be improved and standardized in a usable format before it can be fed into a machine learning model. For example, module 310 can perform obstacle clipping and / or rotation to improve the position and / or direction of the sound source and generate improved labeled audio data 126, which can be stored locally on persistent storage device 352 and / or uploaded to a remote server (e.g., server 103). Thus, improved labeled audio data 126 can include audio samples 313, audio sample IDs, improved positions associated with audio samples 313, and / or improved directions associated with audio samples 313. Data 126 can be used as input to a machine learning engine 122 that generates or trains a set of rules, algorithms, and / or predictive models 124 for various purposes, such as motion planning and control.

[0061] At the same time, the audio samples / events 313 and the improved labeled audio data 126 can be provided to the generalization module 311 to generate a performance profile 315 for online performance. For example, the generalization module 311 can communicate and / or operate with the perception module 302 to determine and capture one or more locations and relative directions of sound sources (or obstacles). Figure 4As shown in FIG, perception module 302 can communicate with one or more visual sensors 412 (e.g., radar, camera, and / or LIDAR) from sensor system 115 to detect objects or obstacles at different points in time based on sensor data provided by sensors 412. Using this sensor data, perception module 302 can determine the real-time relative position and orientation of an object at any point in time. Based on the relative position and orientation of the object and input audio samples 313, summarization module 311 can summarize or verify the online performance of improved labeled audio data 126 against real-time information associated with the driving environment perceived by perception module 302 (e.g., the position and orientation of obstacles) to generate performance profile 315. That is, summarization module 311 can use the relative position and orientation of the object and input audio samples 313 as reference information to compare and evaluate improved labeled audio data 126. The summarized performance data of improved labeled audio data 126 can be stored as part of performance profile 315, which can be stored locally in persistent storage device 352.

[0062] Figure 5 FIG is an example diagram illustrating a driving scenario using a system for audio recording and in-vehicle verification according to one embodiment. Figure 5 As the ADV 101 travels along the route, the audio sensor 411 (e.g., a microphone) may detect a siren 511 from an emergency vehicle 501 (e.g., a police car, an ambulance, a fire truck, etc.) and record the siren 511 as an audio sample / event. Therefore, the emergency vehicle 501 is the sound source of the siren 511.

[0063] The recorded audio samples / events can be manually labeled (e.g., using a data labeling application) with an audio sample ID, one or more locations associated with the audio sample (e.g., the location of vehicle 501), a direction associated with the audio sample (e.g., the relative direction of vehicle 501), etc. As previously described, the labeled audio data can then be improved and standardized in a usable format before being fed into a machine learning model. For example, to improve the labeled audio data, obstacle cropping and / or rotation can be performed to improve the location and / or direction associated with the audio sample. The improved labeled audio data can then be fed into a machine learning engine that generates or trains a set of rules, algorithms, and / or predictive models 124 for various purposes, such as for motion planning and control.

[0064] At the same time, the recorded audio samples / events can be provided to a summarization system (e.g., Figure 4The summarization module 311 of ADV 101 is used to summarize or analyze the performance of the improved labeled audio data by comparing it with real-time visual sensor data. For example, when ADV 101 is traveling along a route, visual sensor 412 (e.g., camera, radar, LIDAR) can detect an emergency vehicle 501 within the driving environment of ADV 101. In response to detecting vehicle 501 in the driving environment, visual sensor 412 can provide the position (e.g., x, y, z coordinates) of vehicle 501 at different points in time. Based on the provided position and the reference axis (x-axis) of ADV 101, the perception system of ADV 101 can additionally determine a direction vector from ADV 101 to vehicle 501. Based on the direction vector and reference axis of ADV 101, a direction angle can be determined, wherein the direction angle represents the relative direction of vehicle 501. Online performance of the improved labeled audio data can then be generated for the position and relative direction of vehicle 501 and the recorded audio samples / events.

[0065] Note that although Figure 5 A single vehicle 501 is shown in FIG. 1 , but in other driving scenarios, multiple vehicles 501 may be used, with the system for audio recording and in-vehicle verification being operated with those vehicles simultaneously. Furthermore, vehicle 501 may be stationary, moving toward or away from ADV 101. Similarly, ADV 101 may be stationary, moving toward or away from vehicle 501.

[0066] Figure 6 6 is a flow chart of a method for generating labeled audio data and in-vehicle verification of labeled audio data, according to one embodiment. In some embodiments, method 600 is performed by perception module 302, audio recorder 308, audio data labeling module 309, automatic improvement module 310, summarization module 311, and / or machine learning engine 122.

[0067] Reference Figure 6 At block 601, sounds (e.g., sirens) emitted by objects (e.g., emergency vehicles) within the driving environment of an autonomous vehicle are recorded and converted into audio samples. At block 602, the audio samples are labeled, and the labeled audio samples are improved to generate improved labeled audio data, wherein the improved labeled audio data is used to subsequently train a machine learning algorithm to identify sound sources during autonomous driving of the autonomous vehicle. At block 603, a performance profile of the improved labeled audio data is generated based on at least the audio samples, the position of the object, and the relative direction of the object, wherein the position of the object and the relative direction of the object are determined by a perception system of the autonomous vehicle.

[0068] Figure 7is a flow chart of a method for in-vehicle verification of labeled audio data according to one embodiment. The method or process 700 may be performed by processing logic, which may include software, hardware, or a combination thereof. For example, the process 700 may be performed by Figure 1 The ADS110 implements

[0069] refer to Figure 7 At block 701, processing logic records sounds (e.g., sirens) emitted by obstacles (e.g., emergency vehicles) within the driving environment of an ADV to create audio samples. At block 702, processing logic determines the location of the obstacle and the relative direction of the obstacle based on sensor data provided by the ADV's visual sensors (e.g., cameras, radars, LIDARs). At block 703, processing logic uses the audio samples, the location of the obstacle, and the relative direction of the obstacle to generate a performance profile of improved labeled audio data, wherein the improved labeled audio data is generated by labeling the audio samples and improving the labeled audio samples, and is used to subsequently train a machine learning algorithm to identify sound sources during autonomous driving of the ADV.

[0070] Note that some or all of the components shown and described above may be implemented in software, hardware, or a combination thereof. For example, these components may be implemented as software installed and stored in a permanent storage device, which may be loaded and executed in memory by a processor (not shown) to perform the processes or operations described throughout this application. Alternatively, these components may be implemented as executable code programmed or embedded into dedicated hardware, such as an integrated circuit (e.g., a dedicated IC or ASIC), a digital signal processor (DSP), or a field programmable gate array (FPGA), which may be accessed via corresponding drivers and / or an operating system from an application. Furthermore, these components may be implemented as specific hardware logic in a processor or processor core as part of an instruction set accessible via one or more specific instruction software components.

[0071] Some portions of the foregoing detailed description have been presented in terms of algorithms and symbolic representations of operations on data bits within a computer memory. These algorithmic descriptions and representations are the means used by those skilled in the data processing arts to most effectively convey the substance of their work to others skilled in the art. An algorithm is here, and generally, conceived to be a self-consistent sequence of operations leading to a desired result. These operations are those requiring physical manipulations of physical quantities.

[0072] It should be remembered, however, that all of these and similar terms are associated with the appropriate physical quantities and are merely convenient labels applied to these quantities. Unless otherwise indicated, as will be apparent from the foregoing discussion, it should be understood that throughout this specification, discussions using terms such as those set forth in the appended claims refer to the actions and processes of a computer system or similar electronic computing device that manipulates and transforms data represented as physical (electronic) quantities within the computer system's registers and memories into other data similarly represented as physical quantities within the computer system's memories or registers or other such information storage, transmission, or display devices.

[0073] Embodiments of the present disclosure also relate to apparatus for performing the operations herein. Such a computer program is stored in a non-transitory computer-readable medium. A machine-readable medium includes any mechanism for storing information in a form readable by a machine (e.g., a computer). For example, a machine-readable (e.g., computer-readable) medium includes a machine (e.g., computer) readable storage medium (e.g., read-only memory ("ROM"), random access memory ("RAM"), magnetic disk storage media, optical storage media, flash memory devices).

[0074] The processes or methods described in the foregoing figures may be performed by processing logic comprising hardware (e.g., circuitry, dedicated logic, etc.), software (e.g., embodied on a non-transitory computer-readable medium), or a combination of both. Although the processes or methods are described above based on some sequential operations, it should be understood that some of the operations described may be performed in a different order. Furthermore, some operations may be performed in parallel rather than sequentially.

[0075] The embodiments of the present disclosure are not described with reference to any particular programming language. It will be appreciated that a variety of programming languages ​​can be used to implement the teachings of the embodiments of the present disclosure as described herein.

[0076] In the foregoing description, embodiments of the present disclosure have been described with reference to specific exemplary embodiments thereof. It will be apparent that various modifications may be made thereto without departing from the broader spirit and scope of the present disclosure as set forth in the appended claims. Accordingly, the description and drawings are to be regarded as illustrative rather than restrictive.

Claims

1. A method for generating labeled audio data using an autonomous vehicle (ADV) and performing on-board verification of the labeled audio data while the autonomous vehicle (ADV) is operating within a driving environment, the method comprising: Recording the sounds of obstacles in the driving environment of the ADV and converting the recorded sounds into audio samples; labeling the audio sample with an audio sample identifier ID of the audio sample, one or more locations associated with the audio sample, and a direction associated with the audio sample, and improving the labeled audio sample to generate improved labeled audio data, wherein the improved labeled audio data is used for subsequent training of a machine learning algorithm to identify sound sources during autonomous driving of an ADV; and generating a performance profile for the improved labeled audio data based on at least the audio samples, the location of the obstacle, and the relative direction of the obstacle, wherein the location of the obstacle and the relative direction of the obstacle are determined by sensor data provided by a vision sensor of the ADV, the vision sensor being coupled to the perception system; wherein generating the performance profile for the improved labeled audio data comprises summarizing the improved labeled audio data with respect to the audio samples, the location of the obstacle, and the relative direction of the obstacle; The improving the marked audio sample includes performing obstacle cropping and / or rotation to improve the position and / or direction associated with the audio sample.

2. The method according to claim 1, wherein The obstruction is an emergency vehicle and the sound emitted is a siren.

3. The method according to claim 1, wherein The audio samples are manually labeled by users of ADV.

4. A computer-implemented method for performing on-board verification of labeled audio data using an autonomous vehicle (ADV) while the ADV is operating within a driving environment, the method comprising: Recording sounds made by obstacles within a driving environment of the ADV to create audio samples; determining a location of an obstacle and a relative direction of the obstacle based on sensor data provided by a vision sensor of the ADV, the vision sensor being coupled to the perception system; as well as generating a performance profile of improved labeled audio data using the audio samples, the locations of the obstacles, and the relative directions of the obstacles, wherein the improved labeled audio data is generated by labeling the audio samples with an audio sample identifier ID of the audio sample, one or more locations associated with the audio sample, and a direction associated with the audio sample and improving the labeled audio samples, and for subsequent training of a machine learning algorithm to identify sound sources during autonomous driving of the ADV; wherein generating the performance profile using the audio samples, the locations of the obstacles, and the relative directions of the obstacles comprises summarizing the improved labeled audio data with respect to the audio samples, the locations of the objects, and the relative directions of the objects; The improving the marked audio sample includes performing obstacle cropping and / or rotation to improve the position and / or direction associated with the audio sample.

5. The method according to claim 4, wherein The obstruction is an emergency vehicle and the sound emitted is a siren.

6. The method according to claim 4, wherein: The audio samples are manually labeled by users of ADV.

7. A system for in-vehicle verification of labeled audio data, comprising: processor; as well as A memory coupled to the processor and storing instructions that, when executed by the processor, cause the processor to perform the operations of the method as claimed in any one of claims 4 to 6.

8. A system for generating labeled audio data using an autonomous vehicle (ADV) and performing on-board verification of the labeled audio data while the autonomous vehicle (ADV) is operating within a driving environment, comprising: processor; as well as A memory coupled to the processor and storing instructions, which, when executed by the processor, cause the processor to perform the operations of the method as claimed in any one of claims 1 to 3.

9. A non-transitory machine-readable medium having instructions stored therein, the instructions, when executed by a processor, causing the processor to perform the operations of the method of any one of claims 1 to 3 or the method of any one of claims 4 to 6.

10. A computer program product comprising a computer program which, when executed by a processor, causes the processor to perform the operations of the method of any one of claims 1 to 3 or the method of any one of claims 4 to 6.

Citation Information

Patent Citations

  • Automatic data labelling for autonomous driving vehicles

    US20190317507A1

  • Using classified sounds and localized sound sources to operate an autonomous vehicle

    US20200241552A1