Data pipeline and deep learning system for autonomous driving

By decomposing sensor data into multiple data components and providing it at different layers of the deep learning network, the problem of reduced signal fidelity and complex conversion processes in sensor data processing is solved, and more efficient signal information utilization and autonomous driving performance are achieved.

CN120126097APending Publication Date: 2025-06-10TESLA INC
View PDF 0 Cites 0 Cited by

Patent Information

Application Number
CN202510202340.0
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Priority Date
2018-06-20
Filing Date
2019-03-20
Publication Date
2025-06-10

AI Technical Summary

Technical Problem

The existing autonomous driving deep learning system has problems such as reducing signal fidelity and complex conversion process in sensor data processing, making it difficult to maximize the signal information of sensor data and provide it to deep learning networks.

Method used

By extracting sensor data to multiple different data components (such as high-pass, low-pass, and bandpass components), and providing these components at different layers of the deep learning network to maintain target-related data and improve the utilization efficiency of signal information.

Benefits of technology

A more complete version of the sensor data is realized, the input signal quality and computing efficiency of the deep learning network are improved, and the identification and control capabilities of the autonomous driving system are enhanced.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120126097A_ABST
    Figure CN120126097A_ABST
Patent Text Reader

Abstract

The embodiment of the invention relates to a data pipeline and a deep learning system for autonomous driving. An image captured using a sensor on a vehicle is received and decomposed into a plurality of component images. Each of the plurality of component images is provided as a different input to a different layer of the plurality of layers of the artificial neural network to determine a result. The results of the artificial neural network are used to at least partially autonomously operate the vehicle.
Need to check novelty before this filing date? Find Prior Art

Description

[0001] Division Application Instructions

[0002] This application is a divisional application of a Chinese patent application with an international filing date of March 20, 2019, entering the Chinese national phase on February 19, 2021, with a national application number of 201980055004.0 and a title of "Data Pipeline and Deep Learning System for Autonomous Driving". Technical Field

[0003] The embodiments of the present application relate to a data pipeline and a deep learning system for autonomous driving. Background Art

[0004] Deep learning systems for implementing autonomous driving typically rely on captured sensor data as input. In traditional learning systems, the captured sensor data can be made compatible with the deep learning system by converting the captured data from the sensor format to a format compatible with the initial input layer of the learning system. This conversion can include compression and downsampling, which can reduce the signal fidelity of the original sensor data. Additionally, changing sensors may require a new conversion process. Therefore, a customized data pipeline is needed that can maximize the signal information from the captured sensor data and provide a higher level of signal information to the deep learning network for deep learning analysis. Brief Description of the Drawings

[0005] The various embodiments of the present invention are disclosed in the following detailed description and the drawings.

[0006] Figure 1 is a flowchart showing an embodiment of a process for performing machine learning processing using a deep learning pipeline.

[0007] Figure 2 is a flowchart showing an embodiment of a process for performing machine learning processing using a deep learning pipeline.

[0008] Figure 3 is a flowchart showing an embodiment of a process for performing machine learning processing using component data.

[0009] Figure 4 is a flowchart showing an embodiment of a process for performing machine learning processing using high-pass and low-pass component data.

[0010] Figure 5 is a flowchart showing an embodiment of a process for performing machine learning processing using high-pass, band-pass, and low-pass component data.

[0011] Figure 6 is a block diagram showing an embodiment of a deep learning system for autonomous driving. Detailed implementation manners

[0012] The present invention can be implemented in various ways, including as a process; an apparatus; a system; a composition of matter; a computer program product embodied on a computer-readable storage medium; and / or a processor, such as a processor configured to execute instructions stored on a memory coupled to the processor and / or provided by the memory. In this specification, these implementations or any other form that the present invention may take can be referred to as techniques. Generally, within the scope of the present invention, the order of steps of the disclosed processes can be changed. Unless otherwise stated, components such as processors or memories described as being configured to perform tasks can be implemented as general components temporarily configured to perform the tasks at a given time or as specific components manufactured to perform the tasks. As used herein, the term "processor" refers to one or more devices, circuits, and / or processing cores configured to process data such as computer program instructions.

[0013] The following provides a detailed description of one or more embodiments of the present invention and the accompanying drawings showing the principles of the present invention. The present invention is described in connection with such embodiments, but the present invention is not limited to any embodiment. The scope of the present invention is only limited by the claims, and the present invention covers many alternatives, modifications, and equivalent forms. Many specific details are set forth in the following description to provide a thorough understanding of the present invention. These details are provided for purposes of illustration, and the present invention can be practiced without some or all of these specific details in accordance with the claims. For clarity, technical materials known in the technical field related to the present invention are not described in detail so as not to unnecessarily obscure the present invention.

[0014] A data pipeline is disclosed that extracts sensor data and provides it as separate components to a deep learning network for autonomous driving. In some embodiments, autonomous driving is implemented using a deep learning network and input data received from sensors. For example, sensors fixed to a vehicle provide real-time sensor data of the vehicle's surrounding environment, such as visual, radar, and ultrasonic data, to a neural network to determine vehicle control responses. In some embodiments, the network is implemented using multiple layers. Based on the signal information of the data, the sensor data is extracted into two or more different data components. For example, feature and / or edge data can be extracted as different data components separately from global data such as global illumination data. The different data components retain target-related data, e.g., data that will ultimately be used by the deep learning network to identify edges and other features. In some embodiments, the different data components serve as containers for storing data highly relevant to identifying certain target features, but they do not themselves identify or detect features. The different data components extract data to ensure accurate feature detection at appropriate stages of the machine learning network. In some embodiments, the different data components can then be preprocessed to enhance the specific signal information they contain. The data components can be compressed and / or downsampled to increase resource and computational efficiency. Then, the different data components are provided to the deep learning system at different layers of the system. The deep learning network is capable of using the signal information retained during the extraction process as input to accurately identify and detect features associated with the target data (e.g., edges, objects, etc.) of the data components. For example, the feature and edge data are provided to the first layer of the network, while the global data is provided to a later layer of the network. By extracting different data components that each retain their respective target signal information, the network can process sensor data more effectively. Instead of receiving the sensor data as the initial input to the network, the most useful information is provided to the network at the most appropriate layer of the network. In some embodiments, a more complete version of the captured sensor data is analyzed by the network because the different data components can make full use of the image resolution of their respective components for their intended purposes. For example, the input for features and edges can utilize the full resolution, bit range, and bit depth for the feature and edge data, while the input for global illumination can utilize the full resolution, bit range, and bit depth for the global illumination data.

[0015] In some embodiments, images captured using sensors on a vehicle are received. For example, images are captured from a high-dynamic-range front camera. As another example, ultrasonic data is captured from a side-facing ultrasonic sensor. In some embodiments, the received images are decomposed into multiple component images. For example, feature data is extracted from the captured high-dynamic-range image. As another example, global illumination data is extracted from the captured high-dynamic-range image. As another example, high-pass, low-pass, and / or band-pass filters can be used to decompose the image. In some embodiments, each of the multiple component images is provided as a different input to a different one of multiple layers of an artificial neural network to determine a result. For example, an artificial neural network such as a convolutional neural network includes multiple layers for processing input data. Different component images obtained by decomposing the captured image are provided as inputs to different layers of the neural network. For example, the feature data is presented as an input to the first layer of the network, while the global data is presented as an input to a later layer of the network (e.g., the third layer). In some embodiments, the result of the artificial neural network is used to at least partially autonomously operate the vehicle. For example, the result of a deep learning analysis using the artificial neural network is used to control the vehicle's steering, braking, lighting, and / or warning systems. In some embodiments, the result is used to autonomously match the speed of the vehicle to traffic conditions, guide the vehicle to follow a navigation path, avoid collisions when an object is detected, summon the vehicle to a desired location, and warn the user of potential collisions, as well as other autonomous driving applications.

[0016] In some embodiments, a vehicle is attached with multiple sensors for capturing data. For example, in some embodiments, eight surround cameras are attached to the vehicle and provide 360-degree visibility within a range of up to 250 meters around the vehicle. In some embodiments, the camera sensors include a wide-angle front camera, a narrow-angle front camera, a rear camera, a front-side camera, and / or a rear-side camera. In some embodiments, ultrasonic and radar sensors are used to capture surrounding details. For example, twelve ultrasonic sensors can be fixed to the vehicle to detect both hard and soft objects. In some embodiments, forward radar is used to capture data on the surrounding environment. In various embodiments, the radar sensor is capable of capturing surrounding details despite heavy rain, fog, dust, and other vehicles. Various sensors are used to capture the environment around the vehicle, and the captured images are provided for deep learning analysis.

[0017] Determine machine learning results for autonomous driving using data captured from sensors and analyzed using the disclosed deep learning system. In various embodiments, the machine learning results are provided to a vehicle control module to implement autonomous driving features. For example, the vehicle control module can be used to control the steering, braking, warning systems, and / or lighting of a vehicle. In some embodiments, the vehicle is controlled to navigate on a road, match the speed of the vehicle to traffic conditions, keep the vehicle within a lane, automatically change lanes without driver input, transition the vehicle from one highway to another, exit the highway when approaching a destination, automatically park the vehicle, and summon the vehicle to and from a parking space, as well as other autonomous driving applications. In some embodiments, the autonomous driving features include identifying opportunities to move the vehicle to a faster lane when behind slower traffic. In some embodiments, machine learning results are used to determine when autonomous driving without driver interaction is appropriate and when it should be disabled. In various embodiments, the machine learning results are used to assist a driver in driving the vehicle.

[0018] In some embodiments, the machine learning results are used to implement an automatic parking mode in which the vehicle will automatically search for a parking space and park the vehicle. In some embodiments, the machine learning results are used to navigate the vehicle using a destination from a user's calendar. In various embodiments, the machine learning results are used to implement autonomous driving safety features such as collision avoidance and automatic emergency braking. For example, in some embodiments, the deep learning system detects an object that may collide with the vehicle, and the vehicle control module applies the brakes accordingly. In some embodiments, the vehicle control module uses deep learning analysis to implement side, front, and / or rear collision warnings that warn the user of the vehicle of a possible collision with an obstacle beside, in front of, or behind the vehicle. In various embodiments, the vehicle control module can activate warning systems such as collision alerts, audio alerts, visual alerts, and / or physical alerts (such as vibration alerts) etc. to notify the user of an emergency or a situation where driver attention is necessary. In some embodiments, the vehicle control module can initiate a communication response according to the situation, such as an emergency response call, a text message, a network update, and / or another communication response, to notify, for example, another party of an emergency. In some embodiments, the vehicle control module can adjust the lighting based on the deep learning analysis results, including high / low beam lights, brake lights, interior lights, emergency lights, etc. In some embodiments, the vehicle control module can also adjust the audio inside or around the vehicle based on the deep learning analysis results, including using a horn, modifying the audio played from the vehicle's sound system (e.g., music, phone calls, etc.), adjusting the volume of the sound system, playing an audio alert, enabling a microphone, and so on.

[0019] Figure 1 is a flowchart showing an embodiment of a process for performing machine learning processing using a deep learning pipeline. For example, Figure 1 the process can be used to implement autonomous driving and autonomous driving features of driver assistance vehicles, thereby improving safety and reducing the risk of accidents. In some embodiments, Figure 1 the process preprocesses data captured by sensors for deep learning analysis. By preprocessing the sensor data, the data provided for deep learning analysis can be enhanced, and more accurate results can be provided for controlling the vehicle. In some embodiments, the preprocessing addresses data mismatches between the data captured by the sensors and the data that the neural network expects for deep learning.

[0020] At 101, sensor data is received. For example, the sensor data is captured by one or more sensors fixed to the vehicle. In some embodiments, the sensors are fixed to the environment and / or other vehicles, and the data is received remotely. In various embodiments, the sensor data is image data, such as the RGB or YUV channels of an image. In some embodiments, the sensor data is captured using a high dynamic range camera. In some embodiments, the sensor data is radar, LiDAR, and / or ultrasonic data. In various embodiments, the LiDAR data is data captured using lasers and can include techniques known as light detection and ranging and laser imaging, detection, and ranging. In various embodiments, the bit depth of the sensor data exceeds the bit depth of the neural network used for deep learning analysis.

[0021] At 103, data preprocessing is performed on the sensor data. In some embodiments, the sensor data can be preprocessed one or more times. For example, the data can first be preprocessed to remove noise, correct alignment problems and / or blurring, and so on. In some embodiments, two or more different filter passes are performed on the data. For example, a high-pass filter can be performed on the data, and a low-pass filter can be performed on the data. In some embodiments, one or more band-pass filters can be performed. For example, in addition to high-pass and low-pass, one or more band-passes can be performed on the data. In various embodiments, the sensor data is divided into two or more data sets, such as a high-pass data set and a low-pass data set. In some embodiments, one or more band-pass data sets are also created. In various embodiments, the different data sets are different components of the sensor data.

[0022] In some embodiments, the different components created by preprocessing the data include feature and / or edge components and global data components. In various embodiments, the feature and / or edge components are created by performing a high-pass or band-pass filter on the sensor data, and the global data components are created by performing a low-pass or band-pass filter on the sensor data. In some embodiments, one or more different filtering techniques may be used to extract the feature / edge data and / or the global data.

[0023] In various embodiments, one or more components of the sensor data are processed. For example, the high-pass component may be processed by removing noise from the image data and / or enhancing the local contrast of the image data. In some embodiments, the low-pass component is compressed and / or downsampled. In various embodiments, the different components are compressed and / or downsampled. For example, the components may be appropriately compressed, resized, and / or downsampled to adjust the size and / or resolution of the data for inputting the data into a layer of a machine learning model. In some embodiments, the bit depth of the sensor data is adjusted. For example, the data channels of camera-captured data at 20 bits or other suitable bit depth are compressed or quantized to 8 bits to prepare the channels for an 8-bit machine learning model. In some embodiments, one or more sensors capture data at a bit depth of 12 bits, 16 bits, 20 bits, 32 bits, or another suitable bit depth greater than the bit depth used by the deep learning network.

[0024] In various embodiments, the preprocessing performed at 103 is executed by an image preprocessor. In some embodiments, the image preprocessor is a graphics processing unit (GPU), a central processing unit (CPU), an artificial intelligence (AI) processor, an image signal processor, a tone mapper processor, or other similar hardware processor. In various embodiments, different image preprocessors are used to extract and / or preprocess different data components in parallel.

[0025] At 105, a deep learning analysis is performed. For example, the deep learning analysis is performed using a machine learning model such as an artificial neural network. In various embodiments, the deep learning analysis receives the processed sensor data for 103 as input. In some embodiments, the processed sensor data is received at 105 as multiple different components such as a high-pass data component and a low-pass data component. In some embodiments, the different data components are received as inputs to different layers of the machine learning model. For example, the neural network receives the high-pass component as the initial input to the first layer of the network and receives the low-pass component as the input to a subsequent layer of the network.

[0026] At 107, the results of deep learning analysis are provided for vehicle control. For example, the results can be provided to a vehicle control module to adjust the speed and / or steering of the vehicle. In various embodiments, the results are provided to implement an autonomous driving function. For example, the results can indicate an object that should be avoided by steering the vehicle. As another example, the results can indicate a vehicle that should be avoided by braking and changing the position of the vehicle in the lane.

[0027] Figure 2 is a flowchart showing an embodiment of a process for performing machine learning processing using a deep learning pipeline. For example, Figure 2 the process can be used to preprocess sensor data, extract image components from the sensor data, preprocess the extracted image components, and then provide these components for deep learning analysis. The results of the deep learning analysis can be used to implement autonomous driving to improve safety and reduce the risk of accidents. In some embodiments, Figure 2 the process is used to perform Figure 1 the process. In some embodiments, step 201 is performed at 101 of Figure 1 ; steps 203, 205, 207, and / or 209 are performed at 103 of Figure 1 ; and / or step 211 is performed at 105 and / or 107 of Figure 1 . By processing the extracted components of the sensor data, the processed data provided to the machine learning model is enhanced, thereby obtaining excellent results from the deep learning analysis instead of using other non-enhanced data. In the example shown, the results of the deep learning analysis are used for vehicle control.

[0028] At 201, sensor data is received. In various embodiments, the sensor data is image data captured from a sensor such as a high dynamic range camera. In some embodiments, the sensor data is captured from one or more different sensors. In some embodiments, the image data is captured using a bit depth of 12 bits or higher to increase the data fidelity.

[0029] At 203, the data is preprocessed. In some embodiments, an image preprocessor such as an image signal processor, a graphics processing unit (GPU), a tone mapper processor, a central processing unit (CPU), an artificial intelligence (AI) processor, or other similar hardware processor is used to preprocess the data. In various embodiments, linearization, demosaicing, and / or another processing technique can be performed on the captured sensor data. In various embodiments, the high-resolution sensor data is preprocessed to enhance the fidelity of the captured data and / or reduce the introduction of errors through subsequent steps. In some embodiments, the preprocessing step is optional.

[0030] At 205, one or more image components are extracted. In some embodiments, two image components are extracted. For example, a feature / edge data component of the sensor data is extracted, and a global data component of the sensor data is extracted. In some embodiments, a high-pass component and a low-pass component of the sensor data are extracted. In some embodiments, one or more additional band-pass components are extracted from the sensor data. In various embodiments, high-pass, low-pass, and / or band-pass filters are used to extract different components of the sensor data. In some embodiments, a tone mapper is used to extract the image components. In some embodiments, the global data and / or low-pass component data are extracted by downsampling the sensor data using binning or a similar technique. In various embodiments, the extracted components hold and preserve target signal information as image data components, but do not actually detect or identify features related to the target information. For example, the extraction of an image component corresponding to edge data results in an image component having target signal information for accurately identifying edges, but the extraction performed at 205 does not detect the presence of edges in the sensor data.

[0031] In some embodiments, a process that preserves the response of the first layer of a deep learning analysis is used to extract the image data components extracted for the first layer of a machine learning network. For example, the relevant signal information for the first layer is preserved such that the result of the analysis performed on the image components after the analysis of the first layer is similar to the analysis performed on the corresponding sensor data before the extraction as image components. In various embodiments, the result is preserved for filters as small as a 5×5 matrix filter.

[0032] In some embodiments, the extracted data components are created by combining multiple channels of the captured image into one or more channels. For example, the red, green, and blue channels can be averaged to create a new channel for the data component. In various embodiments, the extracted data components can be constructed from one or more different channels of source captured data and / or different captured images of different sensors. For example, data from multiple sensors can be combined into a single data component.

[0033] In some embodiments, an image pre-processor, such as the pre-processor of step 203, is used to extract different components. In some embodiments, an image signal processor may be used to extract different components. In various embodiments, a graphics processing unit (GPU) may be used to extract different components. In some embodiments, different pre-processors are used to extract different components so that multiple components can be extracted in parallel. For example, an image signal processor may be used to extract a high-pass component, and a GPU may be used to extract a low-pass component. As another example, an image signal processor may be used to extract a low-pass component, and a GPU may be used to extract a high-pass component. In some embodiments, a tone mapper processor is used to extract an image component (such as a high-pass component), and a GPU is used to extract a separate image component (such as a low-pass component) in parallel. In some embodiments, the tone mapper is part of the image signal processor. In some embodiments, there are multiple instances of similar pre-processors for performing extraction in parallel.

[0034] At 207, component pre-processing is performed. In some embodiments, an image pre-processor, such as the pre-processor of step 203 and / or 205, is used to pre-process one or more components. In some embodiments, different pre-processors are used to pre-process different components so that pre-processing can be performed on different components in parallel. For example, an image signal processor may be used to process a high-pass component, and a graphics processing unit (GPU) may be used to process a low-pass component. In some embodiments, a tone mapper processor is used to process one image component, and a GPU is used to process a separate image component in parallel. In some embodiments, there are multiple instances of similar pre-processors for pre-processing different components in parallel.

[0035] In various embodiments, pre-processing includes downsampling and / or compressing the image component data. In some embodiments, pre-processing includes removing noise from the component data. In some embodiments, pre-processing includes compressing or quantizing the captured data from 20 bits down to an 8-bit data field. In some embodiments, pre-processing includes converting the size of the image component to a lower resolution. For example, the image component may be half, a quarter, an eighth, a sixteenth, a thirty-second, a sixty-fourth, or other suitable scaling ratio of the original sensor image size. In various embodiments, the image component is reduced to a size suitable for the input layer of a machine learning model.

[0036] At 209, the components are provided to appropriate network layers of a deep learning network. For example, different components may be provided to different layers of the network. In some embodiments, the network is a neural network having multiple layers. For example, the first layer of the neural network receives high-pass component data as input. One of the subsequent network layers receives low-pass component data corresponding to global illumination data as input. In various embodiments, different components extracted at 205 and preprocessed at 207 are received at different layers of the neural network. As another example, feature and / or edge data components are provided as input to the first layer of a deep learning network such as an artificial neural network. The global data component is provided to subsequent layers and may be provided as a compressed and / or downsampled version of the data since the global data does not require as high of a precision as the feature and / or edge component data. In various embodiments, the global data is more easily compressed without loss of information and may be provided at a later layer of the network.

[0037] In some embodiments, the machine learning model consists of multiple successive layers, where one or more subsequent layers receive input data that has a dimensionality attribute that is less than the previous layer. For example, the first layer of the network may receive an image size similar to the size of the captured image. Subsequent layers may receive input data that is half or a quarter of the size of the captured image. The reduction in the input data dimensionality reduces the computation of the subsequent layers and improves the efficiency of the deep learning analysis. By providing the sensor input data as different components and at different layers, the computational efficiency can be improved. The earlier layers of the network require increased computation, especially since the amount and size of the data are greater than the subsequent layers. Since the input data has been compressed by the previous layer of the network and / or the preprocessing at 207, the subsequent layers can compute more efficiently.

[0038] At 211, the results of the deep learning analysis are provided for vehicle control. For example, the machine learning results using the processed image components can be used to control the movement of the vehicle. In some embodiments, the results correspond to vehicle control actions. For example, the results may correspond to the speed and steering of the vehicle. In some embodiments, the results are received by a vehicle control module for assisting in maneuvering the vehicle. In some embodiments, the results are used to improve the safety of the vehicle. In various embodiments, the results provided at 211 are determined by performing a deep learning analysis on the components provided at 209.

[0039] Figure 3 is a flowchart showing an embodiment of a process for performing machine learning processing using component data. In the illustrated example, Figure 3The process is used to extract feature and edge data separated from global data from sensor data. Then, these two data sets are fed into a deep learning network at different stages to infer vehicle control results. By separating the two components and providing them at different stages, the initial layers of the network can dedicate computational resources to initial edge and feature detection. In some embodiments, the initial stage dedicates resources to the initial identification of objects such as roads, lane markings, obstacles, vehicles, pedestrians, traffic signs, etc. Subsequent layers can utilize the global data in a more efficient computational manner as the global data is less resource-intensive. Since machine learning can be computationally and data-intensive, a data pipeline using different image components at different stages is used to improve the efficiency of deep learning computations and reduce the data resource requirements for analysis. In some embodiments, Figure 3 The process is used to perform Figure 1 and / or Figure 2 The process. In some embodiments, step 301 is performed at Figure 1 101 of Figure 2 and / or Figure 1 201 of Figure 2 ; step 303 is performed at Figure 1 103 of Figure 2 and / or Figure 1 203 of Figure 2 ; steps 311 and / or 321 are performed at Figure 1 103 of Figure 2 and / or Figure 1 205 of Figure 2 ; steps 313, 323 and / or 325 are performed at Figure 1 103 of Figure 2 and / or Figure 1 207 and 209 of Figure 2 ; steps 315 and / or 335 are performed at Figure 1 105 of Figure 2 and / or

[0040] At 301, sensor data is received. In various embodiments, the sensor data is data captured by one or more sensors of a vehicle. In some embodiments, the sensor data is received as described in step 101 of Figure 1 and / or Figure 2 step 201 of

[0041] At 303, data preprocessing is performed. For example, the sensor data is enhanced by preprocessing the data. In some embodiments, the data is cleaned by performing denoising, alignment, or other appropriate filters. In various embodiments, as described in step 103 of Figure 1 and / or Figure 2Preprocess the data as described in step 203. In the example shown, the process continues to steps 311 and 321. In some embodiments, the processing at 311 and 321 runs in parallel to extract and process different components of the sensor data. In some embodiments, each branch of the processing (e.g., the branch starting at 311 and the branch starting at 321) runs sequentially or pipelined. For example, the processing starts at step 311 to prepare data for the initial layer of the network. In some embodiments, the preprocessing step is optional.

[0042] At 311, features and / or edge data are extracted from the sensor data. For example, the feature data and / or edge data are extracted from the captured sensor data as component data. In some embodiments, the component data retains the relevant signal information from the sensor data for identifying features and / or edges. In various embodiments, the extraction process maintains the signal information crucial for identifying and detecting features and / or edges and does not actually identify or detect features or edges from the sensor data. In various embodiments, the features and / or edges are detected during one or more analysis steps at 315 and / or 335. In some embodiments, the extracted features and / or edge data have the same bit depth as the original captured data. In some embodiments, the extracted data is feature data, edge data, or a combination of feature and edge data. In some embodiments, a high-pass filter is used to extract features and / or edge data from the sensor data. In various embodiments, a tone mapper processor is calibrated to extract features and / or edge data from the sensor data.

[0043] At 313, preprocessing is performed on the features and / or edge data. For example, a denoising filter can be applied to the data to improve the signal quality. As another example, different preprocessing techniques such as local contrast enhancement, gain adjustment, thresholding, noise filtering, etc. can be applied before deep learning analysis to enhance the feature and edge data. In various embodiments, the preprocessing is customized to enhance the feature and edge characteristics of the data rather than applying more general preprocessing techniques to the overall sensor data. In some embodiments, the preprocessing includes performing compression and / or downsampling on the extracted data. In some embodiments, the preprocessing step at 313 is optional.

[0044] At 315, an initial analysis is performed using feature and / or edge data. In some embodiments, the initial analysis is a deep learning analysis using a machine learning model such as a neural network. In various embodiments, the initial analysis receives the feature and edge data as input to the first layer of the network. In some embodiments, the initial layer of the network is configured to detect features and / or edges in the captured image preferentially. In various embodiments, the deep learning analysis is performed using an artificial neural network such as a convolutional neural network. In some embodiments, the analysis runs on an artificial intelligence (AI) processor.

[0045] At 321, global data is extracted from the sensor data. For example, the global data is extracted from the captured sensor data as component data. In some embodiments, the global data corresponds to global illumination data. In some embodiments, the extracted global data has the same bit depth as the original captured data. In some embodiments, a low-pass filter is used to extract the global data from the sensor data. In various embodiments, a tone mapper processor is calibrated to extract the global data from the sensor data. Other techniques such as merging, resampling, and downsampling can also be used to extract the global data. In various embodiments, the extraction process preserves data that may be globally relevant and does not identify and detect global features from the sensor data. In various embodiments, global features are detected by the analysis performed at 335.

[0046] At 323, preprocessing is performed on the global data. For example, a denoising filter can be applied to the data to improve the signal quality. As another example, different preprocessing techniques such as local contrast enhancement, gain adjustment, thresholding, and noise filtering can be applied to enhance the global data before the deep learning analysis. In various embodiments, the preprocessing is customized to enhance the characteristics of the global data rather than applying more general preprocessing techniques to the sensor data as a whole. In some embodiments, the preprocessing of the global data includes compressing the data. In some embodiments, the preprocessing step at 323 is optional.

[0047] At 325, the global data is downsampled. For example, the resolution of the global data is reduced. In some embodiments, the size of the global data is reduced to improve the computational efficiency of analyzing the data and to configure the global data as input to a later layer of the deep learning network. In some embodiments, the global data is downsampled by merging, resampling, or another suitable technique. In some embodiments, a graphics processing unit (GPU) or an image signal processor is used to perform the downsampling. In various embodiments, downsampling is applied to the global data because the global data does not have the same resolution requirements as the feature and / or edge data. In some embodiments, the downsampling performed at 325 is carried out when the global data is extracted at 321.

[0048] At 335, the results of deep learning analysis of feature and / or edge data and global data are used as inputs to perform additional deep learning analysis. In various embodiments, the deep learning analysis receives global data as an input at a later layer of the deep learning network. In various embodiments, the expected input data size at the layer receiving the global data is less than the expected input data size of the initial input layer. For example, the input size of the global data input layer can be half or a quarter of the input size of the initial layer of the deep learning network. In some embodiments, a later layer of the network utilizes the global data to enhance the results of the initial layer. In various embodiments, deep learning analysis is performed and a vehicle control result is determined. For example, a convolutional neural network is used to determine the vehicle control result. In some embodiments, the analysis runs on an artificial intelligence (AI) processor.

[0049] At 337, the results of the deep learning analysis are provided for vehicle control. For example, the machine learning results of the extracted and processed image components are used to control the movement of the vehicle. In some embodiments, the results correspond to vehicle control actions. In some embodiments, the results are provided as described in step 107 of Figure 1 and / or Figure 2 step 211 of

[0050] Figure 4 is a flowchart showing an embodiment of a process for performing machine learning processing using high-pass and low-pass component data. In the example shown, Figure 4 the process is for extracting two data components from sensor data and providing these components to different layers of a deep learning network such as an artificial neural network. The two components are extracted using high-pass and low-pass filters. In various embodiments, the results are used to achieve autonomous driving with improved accuracy, safety, and / or comfort results. In some embodiments, Figure 4 the process is for performing Figure 1 and / or 2 and / or the process of 3. In some embodiments, step 401 is performed at 103 of Figure 1 and / or 203 of Figure 2 and / or 303 of Figure 3 ; step 403 is performed at 103 of Figure 1 and / or 205 of Figure 2 and / or 311 of Figure 3 ; step 413 is performed at 103 of Figure 1 and / or 205 of Figure 2 and / or 321 of Figure 3 ; step 405 is performed at 103 of Figure 1 and / or 207 and 209 of Figure 2 and / or 313 of Figure 3 ; steps 415 and 417 are atFigure 1 at 103 of Figure 2 207 and 209 of, and / or Figure 3 execute at 323 and 325 of; Step 407 is at Figure 1 at 105 of Figure 2 211 of and / or Figure 3 execute at 315 of; and / or Step 421 is at Figure 1 at 105 of Figure 2 211 of and / or Figure 3 execute at 335 of.

[0051] At 401, preprocess the data. In some embodiments, the data is sensor data captured from one or more sensors such as a high-dynamic range camera, radar, ultrasound, and / or LiDAR sensor. In various embodiments, as described with respect to Figure 1 , Figure 2 203 of and / or Figure 3 303 of to preprocess the data. Once the data is preprocessed, the process continues to 403 and 413. In some embodiments, steps 403 and 413 run in parallel.

[0052] At 403, perform a high-pass filter on the data. For example, perform a high-pass filter on the captured sensor data to extract high-pass component data. In some embodiments, use a graphics processing unit (GPU), tone mapper processor, image signal processor, or other image preprocessor to perform the high-pass filter. In some embodiments, the high-pass data component represents the features and / or edges of the captured sensor data. In various embodiments, the high-pass filter is configured to maintain the response of the first layer of the deep learning process. For example, construct the high-pass filter to maintain the response to a small filter at the top of the machine learning network. Maintain the relevant signal information for the first layer of the network such that the analysis results performed on the high-pass component data after the first layer are similar to the analysis results performed on the unfiltered data after the first layer. In various embodiments, the results are maintained for filters as small as a 5×5 matrix filter.

[0053] At 413, perform a low-pass filter on the data. For example, perform a low-pass filter on the captured sensor data to extract low-pass component data. In some embodiments, use a graphics processing unit (GPU), tone mapper processor, image signal processor, or other image preprocessor to perform the low-pass filter. In some embodiments, the low-pass data component represents the global data of the captured sensor data, such as global illumination data.

[0054] In various embodiments, the filtering performed at 403 and 413 may use the same or different image preprocessors. For example, a tone mapper processor is used to extract the high-pass data component, while a graphics processing unit (GPU) is used to extract the low-pass data component. In some embodiments, the high-pass or low-pass data is extracted by subtracting one of the data components from the original captured data.

[0055] At 405 and 415, post-processing is performed on the respective high-pass and low-pass data components. In various embodiments, different post-processing techniques are utilized to enhance the signal quality and / or reduce the amount of data required to represent the data. For example, denoising, demosaicking, local contrast enhancement, gain adjustment, and / or threshold processing procedures, etc., may be performed on the respective high-pass and / or low-pass data components. In some embodiments, the data components are compressed and / or downsampled. For example, once the high-pass and / or low-pass data is extracted, the respective data components can be compressed to more effectively utilize the full bit depth range. In some embodiments, the respective data components are compressed or quantized from a higher bit depth captured by the sensor to a lower bit depth compatible with the deep learning network. For example, sensor data captured at 12 bits, 16 bits, 20 bits, 32 bits, or other suitable bit depths per channel can be compressed or quantized to a lower bit depth such as 8 bits per channel. In some embodiments, the post-processing steps at 405 and / or 415 are optional.

[0056] At 417, the low-pass data component is downsampled. In various embodiments, the low-pass data component is fed into the network at a later stage of the network and can be downsampled to a more efficient resource size. For example, the low-pass data component can be extracted at the full sensor size and reduced to half or a quarter of the original size. Other reduction percentages are possible. In various embodiments, the low-pass data is downsampled but the relevant signal information is retained. In many cases, the low-pass data can be easily downsampled without losing signal information. By downsampling the data, the data can be more easily and quickly analyzed in later layers of the deep learning network.

[0057] At 407, a deep learning analysis is performed on the high-pass data component. In some embodiments, the high-pass data component is fed into the initial layer of the deep learning network and represents the most important data for feature and edge detection. In various embodiments, the results of the deep learning analysis performed on the high-pass data component in the first layer are fed into subsequent layers of the network. For example, a neural network may include multiple layers, such as five or more layers. The first layer receives the high-pass data component as input, and the second layer receives the results of the deep learning analysis performed by the first layer. In various embodiments, the second layer or later layers receive the low-pass data component as an additional input to perform additional deep learning analysis.

[0058] At 421, additional deep learning analysis is performed using the analysis results executed at 407 and the downsampled low-pass data component at 417. In various embodiments, the deep learning analysis infers vehicle control results. For example, the results of the deep learning analysis at 407 and 421 are used to control the vehicle for autonomous driving.

[0059] Figure 5 is a flowchart of an embodiment showing a process for performing machine learning processing using high-pass, band-pass, and low-pass component data. In the illustrated example, Figure 5 The process is used to extract three or more data components from sensor data and provide them at different layers of a deep learning network such as an artificial neural network. Similar to Figure 4 The process, high-pass and low-pass components are extracted. Additionally, Figure 5 The process extracts one or more band-pass data components. In various embodiments, the sensor data is decomposed into multiple components provided to different layers of the deep learning network such that the deep learning analysis can emphasize different data sets at different layers of the network.

[0060] In some embodiments, Figure 5 The process is used to perform Figure 1 , 2 , 3, and / or 4 of the processes. In some embodiments, step 501 is performed at 103 of Figure 1 , 203 of Figure 2 , 303 of Figure 3 , and / or 401 of Figure 4 . In some embodiments, step 503 is performed at 103 of Figure 1 , 205 of Figure 2 , 311 of Figure 3 , and / or 403 of Figure 4 ; step 513 is performed at 103 of Figure 1 , 205 of Figure 2 , and / or 311 or 321 of Figure 3 ; and / or step 523 is performed at 103 of Figure 1 , 205 of Figure 2 , 321 of Figure 3 , and / or 413 of Figure 4 . In some embodiments, step 505 is performed at 103 of Figure 1 , 207 and 209 of Figure 2 , 313 of Figure 3 , and / or step 405 of Figure 4 ; step 515 is performed at 103 of Figure 1 , 207 and 209 of Figure 2 , 313, 323, and / or 325 of Figure 3 , and / or Figure 4performed at 405, 415, and / or 417; and / or step 525 is performed at Figure 1 103 of Figure 2 207 and 209 of Figure 3 323 and 325 of, and / or Figure 4 415 and 417 of. In some embodiments, step 537 is performed at Figure 1 105 of Figure 2 211 of Figure 3 315 and 335 of, and / or Figure 4 407 and 421 of.

[0061] At 501, the data is preprocessed. In some embodiments, the data is sensor data captured from one or more sensors such as a high dynamic range camera, radar, ultrasound, and / or LiDAR sensor. In various embodiments, the data is preprocessed as described with respect to Figure 1 103 of Figure 2 203 of Figure 3 303 of and / or Figure 4 401. Once the data is preprocessed, the processing continues to 503, 513, and 523. In some embodiments, steps 503, 513, and 523 run in parallel.

[0062] At 503, a high-pass filter is performed on the data. For example, a high-pass filter is performed on the captured sensor data to extract high-pass component data. In some embodiments, a graphics processing unit (GPU), a tone mapper processor, an image signal processor, or another image preprocessor is used to perform the high-pass filter. In some embodiments, the high-pass data component represents the features and / or edges of the captured sensor data.

[0063] At 513, one or more band-pass filters are performed on the data to extract one or more band-pass data components. For example, a band-pass filter is performed on the captured sensor data to extract component data including a mixture of features, edges, intermediate, and / or global data. In various embodiments, one or more band-pass components can be extracted. In some embodiments, a graphics processing unit (GPU), a tone mapper processor, an image signal processor, or another image preprocessor is used to perform the low-pass filter. In some embodiments, the band-pass data component represents data that is neither primarily edge / feature data nor primarily global data of the captured sensor data. In some embodiments, the band-pass data is used to maintain data fidelity that might be lost using only the high-pass data component and the low-pass data component.

[0064] At 523, a low-pass filter is applied to the data. For example, a low-pass filter is applied to the captured sensor data to extract low-pass component data. In some embodiments, a graphics processing unit (GPU), a tone mapper processor, an image signal processor, or other image pre-processor is used to perform the low-pass filter. In some embodiments, the low-pass data component represents global data of the captured sensor data, such as global illumination data.

[0065] In various embodiments, the filtering performed at 503, 513, and 523 may use the same or different image pre-processors. For example, a tone mapper processor is used to extract high-pass data components, while a graphics processing unit (GPU) is used to extract band-pass and / or low-pass data components. In some embodiments, data components are extracted by subtracting one or more data components from the original captured data.

[0066] At 505, 515, and 525, post-processing is performed on the respective high-pass, band-pass, and low-pass data components. In various embodiments, different post-processing techniques are utilized to enhance signal quality and / or reduce the amount of data required to represent the data. In some embodiments, for the network layer receiving the data components, the different components are compressed and / or downsampled to an appropriate size. In various embodiments, the high-pass data will have a higher resolution than the band-pass data, and the band-pass data will have a higher resolution than the low-pass data. In some embodiments, different band-pass data components will also have different resolutions suitable for the network layer, and each data component is provided as an input. In some embodiments, the respective data components are compressed or quantized from a higher bit depth captured by the sensor to a lower bit depth compatible with the deep learning network. For example, sensor data captured at 12 bits per channel can be compressed or quantized to 8 bits per channel. In various embodiments, as described with respect to Figure 2 207 of and / or Figure 4 405, 415, and / or 417 of are applied preprocessing filters.

[0067] At 537, deep learning analysis is performed using the data component results of 505, 515, and 525. In some embodiments, the high-pass data components are fed into the initial layer of the deep learning network and represent the most important data for feature and edge detection. One or more band-pass data components are fed into the (multiple) intermediate layers of the network and include additional data for identifying features / edges and / or useful intermediate or global information. The low-pass data components are fed into a later layer of the network and include global information to improve the analysis results of the deep learning network. When performing deep learning analysis, as the analysis progresses, other data components representing different sensor data are fed into different layers to improve the accuracy of the results. In various embodiments, the deep learning analysis infers vehicle control results. For example, the results of the deep learning analysis are used to control the vehicle for autonomous driving. In some embodiments, the machine learning results are provided to the vehicle control module to at least partially autonomously operate the vehicle.

[0068] Figure 6 is a block diagram showing an embodiment of a deep learning system for autonomous driving. In some embodiments, Figure 6 the deep learning system can be used to implement autonomous driving features for autonomous driving and driver assistance vehicles. For example, using sensors fixed to the vehicle, sensor data can be captured, processed as different input components, and then fed into different stages of the deep learning network. The results of the deep learning analysis are used by the vehicle control module to assist in the operation of the vehicle. In some embodiments, the vehicle control module is used for autonomous driving or driver assistance operation of the vehicle. In various embodiments, Figures 1 - 5 the process utilizes a deep learning system such as Figure 6 described in

[0069] In the example shown, the deep learning system 600 is a deep learning network that includes a sensor 601, an image preprocessor 603, a deep learning network 605, an artificial intelligence (AI) processor 607, a vehicle control module 609, and a network interface 611. In various embodiments, the different components are communicatively connected. For example, sensor data from the sensor 601 is fed into the image preprocessor 603. The processed sensor data components of the image preprocessor 603 are fed into the deep learning network 605 running on the AI processor 607. The output of the deep learning network 605 running on the AI processor 607 is fed into the vehicle control module 609. In various embodiments, the network interface 611 is used to communicate with a remote server based on the autonomous operation of the vehicle, make calls, send and / or receive text messages, and so on.

[0070] In some embodiments, sensor 601 includes one or more sensors. In various embodiments, sensor 601 can be fixed to the vehicle at different positions of the vehicle and / or oriented in one or more different directions. For example, sensor 601 can be fixed to the front, side, rear, and / or roof of the vehicle in the forward, backward, sideward, etc. directions. In some embodiments, sensor 610 can be an image sensor such as a high-dynamic range camera. In some embodiments, sensor 601 includes non-visual sensors. In some embodiments, sensor 601 includes radar, LiDAR, and / or ultrasonic sensors, etc. In some embodiments, sensor 601 is not mounted on a vehicle having vehicle control module 609. For example, sensor 601 can be mounted on an adjacent vehicle and / or fixed to the road or environment and be included as part of a deep learning system for capturing sensor data.

[0071] In some embodiments, image pre-processor 603 is used to preprocess the sensor data of sensor 601. For example, image pre-processor 603 can be used to preprocess sensor data, divide the sensor data into one or more components, and / or post-process one or more components. In some embodiments, image pre-processor 603 is a graphics processing unit (GPU), a central processing unit (CPU), an image signal processor, or a dedicated image processor. In various embodiments, image pre-processor 603 is a tone mapping processor for processing high-dynamic range data. In some embodiments, image pre-processor 603 is implemented as part of artificial intelligence (AI) processor 607. For example, image pre-processor 603 can be a component of AI processor 607.

[0072] In some embodiments, deep learning network 605 is a deep learning network for implementing autonomous vehicle control. For example, deep learning network 605 can be an artificial neural network such as a convolutional neural network (CNN), which is trained using sensor data and is used to output vehicle control results to vehicle control module 609. In various embodiments, deep learning network 605 is a multi-stage learning network and can receive input data at two or more different stages of the network. For example, deep learning network 605 can receive feature and / or edge data at the first layer of deep learning network 605 and receive global data at a later layer of deep learning network 605 (e.g., the second or third layer, etc.). In various embodiments, deep learning network 605 receives data at two or more different layers of the network and can compress and / or reduce the size of the data when processing the data through different layers. For example, the data size at layer 1 is higher in resolution than the data at subsequent stages. In some embodiments, the data size at layer 1 is the full resolution of the captured image data, while the data at subsequent layers is a lower resolution of the captured image data (e.g., one quarter of the size). In various embodiments, the input data received from image preprocessor 603 at the (multiple) subsequent layers of deep learning network 605 matches the (multiple) internal data resolutions of the data processed through one or more previous layers.

[0073] In some embodiments, artificial intelligence (AI) processor 607 is a hardware processor for running deep learning network 605. In some embodiments, AI processor 607 is a dedicated AI processor for performing inference on sensor data using a convolutional neural network (CNN). In some embodiments, AI processor 607 is optimized for the bit depth of the sensor data. In some embodiments, AI processor 607 is optimized for deep learning operations such as neural network operations, including convolution, dot product, vector, and / or matrix operations, etc. In some embodiments, AI processor 607 is implemented using a graphics processing unit (GPU). In various embodiments, AI processor 607 is coupled to a memory that is configured to provide instructions to AI processor 607 that, when executed, cause AI processor 607 to perform a deep learning analysis on the received input sensor data and determine the results of machine learning for at least partially autonomous operation of the vehicle.

[0074] In some embodiments, the vehicle control module 609 is configured to process the output of the artificial intelligence (AI) processor 607 and convert the output into vehicle control operations. In some embodiments, the vehicle control module 609 is configured to control the vehicle for autonomous driving. In some embodiments, the vehicle control module 609 may adjust the speed and / or steering of the vehicle. For example, the vehicle control module 609 may be configured to control the vehicle by braking, steering, changing lanes, accelerating, and merging into another lane, etc. In some embodiments, the vehicle control module 609 is configured to control vehicle lighting, such as brake lights, turn signals, headlights, etc. In some embodiments, the vehicle control module 609 is configured to control the vehicle audio conditions, such as the vehicle's sound system, playing audio alerts, enabling the microphone, enabling the horn, etc. In some embodiments, the vehicle control module 609 is configured to control the notification system, which includes a warning system, to notify the driver and / or passengers of driving events, such as potential collisions or the approach of a predetermined destination. In some embodiments, the vehicle control module 609 is configured to adjust sensors, such as the sensors 601 of the vehicle. For example, the vehicle control module 609 may be configured to change the parameters of one or more sensors, such as modifying the orientation, changing the output resolution and / or format type, increasing or decreasing the capture rate, adjusting the captured dynamic range, adjusting the focus of the camera, enabling and / or disabling the sensors, etc. In some embodiments, the vehicle control module 609 may be configured to change the parameters of the image pre-processor 603, such as modifying the frequency range of the filter, adjusting the feature and / or edge detection parameters, adjusting the channels and bit depth, etc. In various embodiments, the vehicle control module 609 is configured to implement autonomous driving and / or driver assistance control of the vehicle.

[0075] In some embodiments, network interface 611 is a communication interface for sending and / or receiving data including voice data. In various embodiments, network interface 611 includes a cellular or wireless interface for interfacing with a remote server, connecting and making voice calls, sending and / or receiving text messages, etc. For example, network interface 611 may be used to receive instructions and / or updates to operating parameters for sensor 601, image pre-processor 603, deep learning network 605, AI processor 607, and / or vehicle control module 609. For example, the machine learning model of deep learning network 605 may be updated using network interface 611. As another example, network interface 611 may be used to update the firmware of sensor 601 and / or the operating parameters of image pre-processor 603, such as image processing parameters. In some embodiments, network interface 611 is used to make an emergency contact with emergency services in the event of an accident or near-accident. For example, in the event of a collision, network interface 611 may be used to contact emergency services for assistance and may inform the emergency services of the vehicle's location and collision details. In various embodiments, network interface 611 is used to implement autonomous driving features, such as accessing calendar information to retrieve and / or update destination location and / or expected arrival time.

[0076] Although the foregoing embodiments have been described in detail for purposes of clarity of understanding, the present invention is not limited to the details provided. There are many alternative ways to implement the present invention. The disclosed embodiments are illustrative and not restrictive.

Claims

1. A method, comprising: receiving an image captured using a sensor with an autonomous operating system; extracting from the image a global data component and a feature data component that form input data, wherein the global data component is associated with global illumination data and the feature data component is associated with edge data; providing the input data to a convolutional neural network including a plurality of layers, wherein the plurality of layers are consecutive and form respective parts of the convolutional neural network, wherein the feature data component is provided as an input to a first layer of the plurality of layers, wherein the global data component and an intermediate result output from a previous layer are provided as inputs to a second layer of the plurality of layers, and wherein the second layer is after the first layer; and obtaining information indicating a system control result based on a result of the convolutional neural network, the system control result informing an autonomous operation of the autonomous operating system.

2. The method according to claim 1, wherein after extraction, the global data component is downsampled and the downsampled global data component is provided as an input to the second layer.

3. The method according to claim 1, wherein a denoising filter is applied to the extracted global data component.

4. The method according to claim 1, wherein the global data component is extracted via a low-pass filter.

5. The method according to claim 1, wherein the feature data component is extracted via a high-pass filter.

6. The method according to claim 1, wherein one or more of the following are performed on at least a portion of the input data: denoising, demosaicing, local contrast enhancement, gain adjustment, and / or threshold processing.

7. The method according to claim 1, wherein the first layer is an initial layer of the convolutional neural network.

8. The method according to claim 1, wherein a third data component is extracted from the image via a band-pass filter and the third data component forms part of the input data.

9. The method according to claim 8, wherein the third data component is provided to a third layer of the convolutional neural network and the third layer is after the first layer and before the second layer.

10. The method according to claim 1, wherein the system control result is associated with one or more of the following: braking, steering, changing lanes, accelerating, and merging into a different lane.

11. A computer program product embodied in a non-transitory computer-readable storage medium and including computer instructions for: receiving an image captured using a sensor; extracting from the image a global data component and a feature data component that form input data, wherein the global data component is associated with global illumination data and the feature data component is associated with edge data; Providing the input data to a convolutional neural network including a plurality of layers, wherein the plurality of layers are consecutive and form respective parts of the convolutional neural network, wherein the feature data component is provided as an input to the first layer of the plurality of layers, wherein the global data component and an intermediate result output from a previous layer are provided as inputs to the second layer of the plurality of layers, and wherein the second layer is after the first layer; and Obtaining information indicating a system control result based on the result of the convolutional neural network, the system control result informing an autonomous operation of an autonomous operating system.

12. The computer program product according to claim 11, wherein after extraction, the global data component is downsampled, and wherein the downsampled global data component is provided as an input to the second layer.

13. The computer program product according to claim 11, wherein the global data component is extracted via a low-pass filter.

14. The computer program product according to claim 11, wherein the feature data component is extracted via a high-pass filter.

15. The computer program product according to claim 11, wherein a third data component is extracted from the image via a band-pass filter, and wherein the third data component forms part of the input data.

16. A system comprising: a plurality of sensors; one or more processors and a computer storage medium storing instructions which, when executed by the one or more processors, cause the one or more processors to: receive at least one image from at least one of the plurality of sensors; extract a global data component and a feature data component forming input data from the at least one image, wherein the global data component is associated with global illumination data and the feature data component is associated with edge data; provide the input data to a convolutional neural network including a plurality of layers, wherein the plurality of layers are consecutive and form respective parts of the convolutional neural network, wherein the feature data component is provided as an input to the first layer of the plurality of layers, wherein the global data component and an intermediate result output from a previous layer are provided as inputs to the second layer of the plurality of layers, and wherein the second layer is after the first layer; and obtain information indicating a system control result based on the result of the convolutional neural network, the system control result informing an autonomous operation of an autonomous operating system.

17. The system according to claim 16, wherein after extraction, the global data component is downsampled, and wherein the downsampled global data component is provided as an input to the second layer.

18. The system according to claim 16, wherein the global data component is extracted via a low-pass filter.

19. The system according to claim 16, wherein the feature data component is extracted via a high-pass filter.

20. The system according to claim 16, wherein the third data component is extracted from the at least one image via a band-pass filter, and wherein the third data component forms part of the input data.