Deep learning-based real-time detection and correction of faulty sensors in autonomous machines
By adopting a deep learning-based automatic inspection mechanism in autonomous machines, the problem of sensors providing low-quality data is solved, and data accuracy and system reliability are improved.
Patent Information
- Application Number
- CN201811275743.4
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Priority Date
- 2017-11-28
- Filing Date
- 2018-10-30
- Publication Date
- 2025-05-16
- Estimated Expiration
- 2038-10-30
AI Technical Summary
The prior art is difficult to effectively detect and process problematic sensors that provide low quality or misleading data in autonomous machines, resulting in reduced data accuracy.
The automatic inspection mechanism based on deep learning is adopted to detect and capture components such as logic, serial logic, training and reasoning logic to detect the abnormal state of the sensor in real time, and repair it in real time through classification and prediction logic.
Real-time detection and repair of problematic sensors in autonomous machines is realized, data quality and system accuracy are ensured, and the reliability and safety of autonomous machines are improved.
Smart Images

Figure CN109840586B_ABST
Abstract
Description
Technical Field
[0001] Embodiments described herein relate generally to data processing and, more particularly, to facilitating deep learning-based real-time detection and correction of compromised sensors in autonomous machines. Background Art
[0002] Autonomous machines are expected to grow exponentially in the future, which in turn will likely require sensors (e.g., cameras) to lead the growth in facilitating various tasks (e.g., autonomous driving).
[0003] Conventional techniques use multiple sensors to attempt to apply data / sensor fusion for providing some redundancy to ensure accuracy; however, these conventional techniques are severely limited because they cannot cope with or avoid those sensors that provide low quality or misleading data. BRIEF DESCRIPTION OF THE DRAWINGS
[0004] Embodiments are illustrated by way of example, and not by way of limitation, in the figures of the accompanying drawings in which like references refer to similar elements.
[0005] Figure 1 A computing device employing a sensor automated inspection mechanism is shown according to one embodiment.
[0006] Figure 2 According to one embodiment, Figure 1 The sensor automatically checks the mechanism.
[0007] Figure 3A Static inputs from multiple sensors are shown according to one embodiment.
[0008] Figure 3B Dynamic input from a single sensor is shown according to one embodiment.
[0009] Figure 3C Dynamic input from a single sensor is shown according to one embodiment.
[0010] Figure 4A An architectural setup is shown that provides a sequence of transactions for real-time detection and correction of problematic sensors using deep learning according to one embodiment.
[0011] Figure 4B A method for real-time detection and correction of problematic sensors using deep learning is shown according to one embodiment.
[0012] Figure 5 A computer device capable of supporting and implementing one or more embodiments according to one embodiment is shown.
[0013] Figure 6An embodiment of a computing environment capable of supporting and implementing one or more embodiments according to one embodiment is shown. DETAILED DESCRIPTION
[0014] In the following description, numerous specific details are set forth. However, the embodiments described herein may be practiced without these specific details. In other instances, well-known circuits, structures, and techniques are not shown in detail to avoid obscuring the understanding of the description.
[0015] Embodiments provide novel techniques for deep learning-based detection, notification, and correction of problematic sensors in autonomous machines. In one embodiment, the automated inspection may include one or more of: detecting problematic sensors, issuing alarms to warn of problematic sensors, providing real-time repair of any distortions of problematic sensors, etc.
[0016] It is contemplated that embodiments are not limited to any number or type of sensors; however, for simplicity, clarity, and ease of understanding, one or more cameras may be used as exemplary sensors throughout this document, but embodiments are not limited thereto.
[0017] It is contemplated that terms such as "request," "query," "job," "work," "work item," and "workload" may be used interchangeably throughout this document. Similarly, an "application" or "agent" may refer to or include an application program interface (API) (e.g., a free rendering API (e.g., Open Graphics Library)). 11. 12, etc.)) provided by a computer program, software application, game, workstation application, etc., where "dispatch" can be interchangeably referred to as a "work unit" or "rendering", and similarly, an "application" can be interchangeably referred to as a "workflow", or simply, an "agent". For example, a workload (e.g., a workload for a three-dimensional (3D) game) can include and issue any number and type of "frames", where each frame can represent an image (e.g., a sailboat, a human face). In addition, each frame can include and provide any number and type of work units, where each work unit can represent a portion (e.g., a mast of a sailboat, a forehead of a human face) of the image (e.g., a sailboat, a human face) represented by its corresponding frame. However, for consistency, throughout this document, each item can be referred to by a single term (e.g., "dispatch", "agent", etc.).
[0018] In some embodiments, terms such as "display screen" and "display surface" may be used interchangeably to refer to the visible portion of the display device, while the remainder of the display device may be embedded in a computing device (e.g., a smart phone, a wearable device, etc.). It is contemplated and noted that the embodiments are not limited to any particular computing device, software application, hardware component, display device, display screen or surface, protocol, standard, etc. For example, the embodiments may be applied to and used for any number and type of real-time applications on any number and type of computers (e.g., desktop computers, laptop computers, tablet computers, smart phones, head-mounted displays and other wearable devices, etc.). In addition, for example, the range of scenes rendered using the novel technology for efficient performance may range from simple scenes (e.g., desktop compositing) to complex scenes (e.g., 3D games, augmented reality applications, etc.).
[0019] It is noted that throughout this document, reference may be made interchangeably to terms or abbreviations such as convolutional neural network (CNN), CNN, neural network (NN), NN, deep neural network (DNN), DNN, recurrent neural network (RNN), RNN, etc. Also, reference may be made interchangeably to terms such as “autonomous machine” or simply “machine”, “autonomous vehicle” or simply “vehicle”, “autonomous agent” or simply “agent”, “autonomous device” or “computing device”, “robot”, etc. throughout this document.
[0020] Figure 1 A computing device 100 is shown that employs a sensor automatic inspection mechanism ("automatic inspection mechanism") 110 according to one embodiment. The computing device 100 represents or represents any number and type of intelligent devices (such as (but not limited to) intelligent command devices or intelligent personal digital assistants, home / office automation systems, home appliances (such as washing machines, televisions, etc.), mobile devices (such as smart phones, tablet computers, etc.), gaming devices, handheld devices, wearable devices (such as smart watches, smart bracelets, etc.), virtual reality (VR) devices, head-mounted displays (HMDs), Internet of Things (IoT) devices, laptop computers, desktop computers, server computers, set-top boxes (such as Internet-based cable TV set-top boxes, etc.), devices based on the Global Positioning System (GPS), etc.) communication and data processing devices.
[0021] In some embodiments, the computing device 100 may include (but is not limited to) autonomous machines or artificial intelligence agents (e.g., mechanical agents or machines, electronic agents or machines, virtual agents or machines, motor agents or machines, etc.). Examples of autonomous machines or artificial intelligence agents may include (but are not limited to) robots, autonomous vehicles (e.g., self-driving cars, self-flying aircraft, self-navigating ships, etc.), autonomous equipment (self-operating construction vehicles, self-operating medical equipment, etc.), etc. In addition, "autonomous vehicles" are not limited to cars, but they may include any number and type of autonomous machines (e.g., robots, autonomous equipment, home autonomous equipment, etc.), and any one or more tasks or operations related to these autonomous machines may be interchangeably referenced with autonomous driving.
[0022] Furthermore, for example, computing device 100 may include a computer platform hosting an integrated circuit (“IC”) that integrates various hardware and / or software components of computing device 100 on a single chip (eg, a system on a chip (“SoC” or “SOC”)).
[0023] As shown, in one embodiment, computing device 100 may include any number and types of hardware and / or software components, such as (but not limited to) a graphics processing unit (“GPU” or simply “graphics processor”) 114, a graphics driver (also known as a “GPU driver”, “graphics driver logic”, “driver logic”, user mode driver (UMD), UMD, user mode driver framework (UMDF), UMDF or simply “driver”) 116, a central processing unit (“CPU” or simply “application processor”) 112, memory 104, network devices, drivers, etc., and input / output (I / O) sources 108 (e.g., a touch screen, a touchpad, a touch disk, a virtual or conventional keyboard, a virtual or conventional mouse, ports, connectors, etc.). Computing device 100 may include an operating system (OS) 106, which acts as an interface between the hardware and / or physical resources of computer device 100 and a user.
[0024] It should be understood that for a particular implementation, a system with fewer or more configurations than the above examples may be preferred. Thus, depending on a variety of factors (e.g., price constraints, performance requirements, technological improvements, or other circumstances), the configuration of computing device 100 may vary from implementation to implementation.
[0025] Embodiments may be implemented as any of the following or a combination thereof: one or more microchips or integrated circuits interconnected using a motherboard, hardwired logic, software stored in a memory device and executed by a microprocessor, firmware, an application specific integrated circuit (ASIC), and / or a field programmable gate array (FPGA). In addition, the terms "logic", "module", "component", "engine", and "mechanism" may include software or hardware and / or a combination thereof (e.g., firmware) by way of example.
[0026] In one embodiment, as shown, the automatic inspection mechanism 110 may be hosted by an operating system 106 that communicates with an I / O source 108 of the computing device 100. In another embodiment, the automatic inspection mechanism 110 may be hosted or facilitated by a graphics driver 116. In yet another embodiment, the automatic inspection mechanism 110 may be hosted or part of a graphics processing unit (“GPU” or simply “graphics processor”) 114 or firmware of the graphics processor 114. For example, the automatic inspection mechanism 110 may be embedded in or implemented as part of the processing hardware of the graphics processor 114. Similarly, in yet another embodiment, the automatic inspection mechanism 110 may be hosted or part of a central processing unit (“CPU” or simply “application processor”) 112. For example, the automatic inspection mechanism 110 may be embedded in or implemented as part of the processing hardware of the application processor 112.
[0027] In yet another embodiment, the automatic inspection mechanism 110 may be hosted or part of any number or type of components of the computing device 100, for example, a portion of the automatic inspection mechanism 110 may be hosted or part of the operating system 116, another portion may be hosted or part of the graphics processor 114, another portion may be hosted or part of the application processor 112, and one or more portions of the automatic inspection mechanism 110 may be hosted or part of the operating system 116 and / or any number and type of devices of the computing device 100. It is contemplated that embodiments are not limited to any particular implementation or hosting of the automatic inspection mechanism 110, and one or more portions or components of the automatic inspection mechanism 110 may be implemented or implemented as hardware, software, or any combination thereof (e.g., firmware).
[0028] The computing device 100 may host a network interface to provide access to a network (e.g., a LAN, a wide area network (WAN), a metropolitan area network (MAN), a personal area network (PAN), Bluetooth, a cloud network, a mobile network (e.g., 3rd generation (3G), 4th generation (4G), etc.), an intranet, the Internet, etc.). The network interface may include, for example, a wireless network interface having an antenna (which may represent one or more antennas). The network interface may also include, for example, a wired network interface to communicate with a remote device via a network cable (which may be, for example, an Ethernet cable, a coaxial cable, a fiber optic cable, a serial cable, or a parallel cable).
[0029] Embodiments may be provided, for example, as a computer program product, which may include one or more machine-readable media having machine-executable instructions stored thereon, which when executed by one or more machines (e.g., a computer, a network of computers, or other electronic devices) may cause one or more machines to perform operations according to the embodiments described herein. The machine-readable medium may include, but is not limited to, a floppy disk, an optical disk, a CD-ROM (compact disk-read only memory) and a magnetic optical disk, ROM, RAM, EPROM (erasable programmable read-only memory), EEPROM (electrically erasable programmable read-only memory), a magnetic or optical card, a flash memory, or other types of media / machine-readable media suitable for storing machine-executable instructions.
[0030] In addition, embodiments may be downloaded as a computer program product, wherein the program may be transmitted from a remote computer (e.g., a server) to a requesting computer (e.g., a client) via a communication link (e.g., a modem and / or a network connection) by way of one or more data signals embodied and / or modulated in a carrier wave or other propagation medium.
[0031] Throughout this document, the term "user" may be interchangeably referred to as "viewer", "observer", "speaker", "person", "individual", "end-user", etc. Note that throughout this document, terms such as "graphics domain" may be interchangeably referred to as "graphics processing unit", "graphics processor" or simply "GPU", and similarly, "CPU domain" or "host domain" may be interchangeably referred to as "computer processing unit", "application processor" or simply "CPU".
[0032] It is noted that throughout this document, terms such as "node", "computer node", "server", "server device", "cloud computer", "cloud server", "cloud server computer", "machine", "host machine", "device", "computing device", "computer", "computing system", etc. can be used interchangeably. It is also noted that throughout this document, terms such as "application", "software application", "program", "software program", "package", "software software package", etc. can be used interchangeably. Furthermore, throughout this document, terms such as "job", "input", "request", "message", etc. can be used interchangeably.
[0033] Figure 2 According to one embodiment, Figure 1 For the sake of simplicity, the following does not repeat or discuss the referenced Figure 1 In one embodiment, the automated inspection mechanism 110 may include any number and type of components, such as (but not limited to): detection and capture logic 201; concatenation logic 203; training and reasoning logic 205; communication / compatibility logic 209; and classification and prediction logic 207.
[0034] The computing device 100 (also interchangeably referred to as an “autonomous machine” throughout this document) is also shown to include a user interface 219 (e.g., a graphical user interface (GUI)-based user interface, a web browser, a cloud-based platform user interface, a software application-based user interface, other user or application programming interface (API), etc.). The computing device 100 may also include an I / O source 108 having a capture / sensing component 231 (e.g., cameras A 242A, B 242B, C 242C, D 242D (e.g., RealSense TM cameras), sensors, microphones 241, etc.) and output components 233 (e.g., a display device or simply a display 244 (e.g., an integrated display, a tensor display, a projection screen, a display screen, etc.), a speaker device or simply a speaker 243, etc.).
[0035] Computing device 100 is also shown as being able to access and / or communicate with one or more databases 225 and / or one or more other computing devices via one or more communication mediums 230 (e.g., a network (e.g., a cloud network, a proximity network, the Internet), etc.).
[0036] In some embodiments, database 225 may include one or more of storage media or devices, libraries, data sources, etc., having any amount or type of information (e.g., data, metadata, etc.) related to any number and type of applications (e.g., data and / or metadata related to one or more users, physical locations or regions, applicable laws, policies and / or regulations, user preferences and / or profiles, security and / or authentication data, historical and / or preference details, etc.).
[0037] As previously mentioned, the computing device 100 may host an I / O source 108 including a capture / sensing component 231 and an output component 233. In one embodiment, the capture / sensing component 231 may include a sensor array including, but not limited to, a microphone 241 (e.g., an ultrasonic microphone), cameras 242A-242D (e.g., a two-dimensional (2D) camera, a three-dimensional (3D) camera, an infrared (IR) camera, a depth sensing camera, etc.), capacitors, radio components, radar components, etc., scanners, and / or accelerometers, etc. Similarly, the output component 233 may include any number and type of speakers 243, display devices 244 (e.g., a screen, a projector, a light emitting diode (LED), and / or a vibration motor, etc.).
[0038] For example, as shown, the capture / sensing component 231 may include any number and type of microphones 241 (e.g., multiple microphones or microphone arrays (e.g., ultrasonic microphones, dynamic microphones, fiber optic microphones, laser microphones, etc.)). It is contemplated that one or more of the microphones 241 act as one or more input devices for accepting or receiving audio input (e.g., human voice) into the computing device 100 and converting the audio or sound into electrical signals. Similarly, it is contemplated that one or more of the cameras 242A-242D act as one or more input devices for detecting and capturing images and / or video of scenes, objects, etc., and providing the captured data as video input to the computing device 100.
[0039] The contemplated embodiments are not limited to any number or type of microphones 241, cameras 242A-242D, speakers 243, displays 244, etc. For example, as facilitated by detection and capture logic 201, one or more of microphones 241 may be used to simultaneously detect speech or sound from multiple users or speakers (e.g., speakers 250). Similarly, as facilitated by detection and capture logic 201, one or more of cameras 242A-242D may be used to capture images or video of a geographic location (e.g., a room) and its contents (e.g., furniture, electronic devices, people, animals, ground, etc.) and form a collection of image or video streams from the captured data for further processing by automated inspection mechanism 110 at computing device 100.
[0040] Similarly, as shown, output component 233 may include any number and type of speakers 243 to act as output devices for outputting or emitting audio from computing device 100 for any number or type of reasons (e.g., human listening or consumption). For example, speaker 243 works in reverse of microphone 241, where speaker 243 converts electrical signals into sound.
[0041] As described above, embodiments are not limited to any number or type of sensors as part of, embedded in, or coupled to capture / sensing component 231 (e.g., microphone 241, cameras 242A-242D, etc.). In other words, embodiments are applicable to and compatible with any number of types of sensors; however, cameras 242A-242D are used as examples throughout this document for the purpose of simplicity and clarity of discussion. Similarly, embodiments are applicable to all types and styles of cameras, and therefore, cameras 242A-242D do not have to be of a specific type.
[0042] As described above, with the growth of autonomous machines (e.g., self-driving vehicles, drones, home appliances, etc.), sensors of all kinds are expected to impact and facilitate specific tasks necessary for the vitality of autonomous machines (e.g., sensors acting as eyes behind the wheels in the case of self-driving vehicles). Thus, data quality becomes a critical factor when autonomous machines are involved for any number of reasons (e.g., safety, security, trust, etc., specifically, in life-or-death situations, commercial environments, etc.).
[0043] It is expected that high-quality data can ensure that the artificial intelligence (AI) of the autonomous machine (e.g., computing device 100) receives high-quality input (e.g., images, videos, etc.) for outputting high-quality performance. It is also expected that even if one of the sensors (e.g., cameras 242A-242D) is defective or not performing to its full potential (e.g., due to mud or excessive fog on its lens, etc.), the overall performance of the computing device 100 may suffer because its accuracy is compromised.
[0044] For example, the automated inspection mechanism 110 provides novel techniques for filtering cameras 242A-242D to detect any anomalies in any one or more of the cameras 242A-242D that may be responsible for or potentially provide poor quality input, wherein the anomalies in or relating to these cameras 242A-242D may include (but are not limited to) dirt / mud on the lens, obstructions in front of the lens, masking (e.g., fog), physical damage technical issues, etc.
[0045] Conventional techniques are unable to detect these anomalies and therefore cannot guarantee the accuracy of the data being collected by the autonomous machine's sensors, which often results in low-quality data or even misleading data.
[0046] Embodiments provide novel techniques for detecting anomalies with respect to sensors (e.g., cameras 242A-242D) in real time, issuing warnings or alerts as needed, and repairing or fixing these anomalies. In one embodiment, the automated inspection mechanism 110 provides novel techniques for sensor automated inspection (SAC) to detect and inspect the status of each sensor (e.g., camera 242A-242D) in a system (e.g., autonomous machine 100) to ensure that all sensors are working well before undertaking any task (e.g., before driving an autonomous vehicle), and to continue to inspect the sensors to ensure that they continue to work, or in the event of any anomaly, to repair them in real time while they are performing any task (e.g., driving).
[0047] In one embodiment, the automated inspection mechanism 110 uses deep learning of the autonomous machine 100 to ensure in real time that the cameras 242A-242D and any other sensors are in working order, or that they are at least immediately noticed and repaired in the event of any problems. As will be further described later in this document, the automated inspection mechanism 110 can use a deep neural network (DNN) (e.g., a convolutional deep learning classifier of a convolutional neural network (CNN)) to continuously and accurately check the real-time status of the cameras 242A-242D and other sensors, and then use the training data to detect and predict which of the cameras 242A-242D or other sensors are likely to be damaged or out of service.
[0048] One of the major weaknesses with conventional technology is when the camera lens becomes covered with debris (eg, dirt, smudges, mud, etc.) because when this occurs, the camera has no ability to detect or capture regardless of the level of debris.
[0049] Embodiments provide for the use of deep learning on an autonomous machine (e.g., autonomous machine 100) to handle complex matters (e.g., in the case of a stain or mud on the lens of a camera (e.g., camera 242A), the obstruction can be continuously observed, including taking into account any movement or changes associated with camera 242A, the stain, and / or the scene). The situation is detected and observed in real time so that the defective or obstructed camera 242A can be repaired.
[0050] In one embodiment, the detection and capture logic 201 of the automated inspection mechanism 110 can be used to trigger one or more cameras 242A-242D located at various locations to capture one or more scenes in front of them. Figure 3A As shown, there may be multiple cameras 242A-242D, fixed in their positions, to capture static input (e.g., to capture a scene from different angles simultaneously). Figure 3BAs shown, in another embodiment, a single camera (e.g., camera 242A) can be used to capture dynamic input (e.g., capturing a scene from the same angle at different points in time). In yet another input, as with respect to Figure 3C As shown, the stain or debris itself may be dynamic or moving, and therefore, a camera (eg, camera 242B) may be used to capture the scene while capturing the movement of the debris.
[0051] As described above, the sensors of the capture / sensing component 231 are not limited to any number or type of cameras 242A-242D or microphone 241, and the sensors may also include other sensors (e.g., light detection and ranging (LiDAR) sensors, ultrasonic sensors, and any number or type of other sensors mentioned or described throughout this document), and any inputs from these sensors can be input into a neural network (e.g., a CNN's softmax) for classification purposes.
[0052] Referring back to the automated inspection mechanism 110, since one or more cameras 242A-242D are capturing the scene, any internal or external issues with any camera 242A-242D may also be detected, where internal issues include any physical flaws (e.g., a lens or portion of the camera is damaged) or technical issues (e.g., the camera stops working) and external issues are related to any form of obstruction (e.g., snow, trees, dirt, mud, debris, people, animals, etc.) that may be on the lens or in the field of view of the lens that blocks the view of the scene.
[0053] For example, if some mud is found on the lens of camera 242A, the detection and capture logic 201 can be triggered to detect the mud or at least detect that the view from camera 242A is blocked in some way. In the event of any movement associated with any camera 242A, the scene (e.g., people moving, waves, traffic moving, etc.), and / or the mud itself (e.g., flowing downward or in the direction of the wind, etc.), the detection and capture logic 201 can collect data that includes information related to the blockage of the view from camera 242A and any one or more of the above-mentioned movements.
[0054] Once the data is collected by the detection and capture logic 201, it is then forwarded to the concatenation logic 203 as an input. As described above, embodiments are not limited to camera inputs, and these inputs may come from other sensors and include LiDAR inputs, radar inputs, microphone inputs, etc., where there may be some degree of overlap in the detection of the same object from these sensors. In one embodiment, in the case of multiple inputs from multiple sensors (e.g., two or more of the cameras 242A-242D) at the same time or at different points in time and / or from the same or different angles and / or from the same sensor (e.g., camera 242A) at multiple points in time and / or from the same or different angles, the concatenation logic 203 may then be triggered to concatenate (or concatenate) these inputs into a single input.
[0055] In one embodiment, any concatenated input of the multiple inputs can then be forwarded to a deep learning neural network model (e.g., CNN) for training and reasoning by the training and reasoning logic 205. In one embodiment, the concatenation logic 203 performs the concatenation in addition to or before the data being processed by the deep learning model, so that there is better flexibility to set their order to further benefit the training process. It is contemplated that embodiments are not limited to any number and type of deep learning models, so that the CNN can be any kind or type of CNN (e.g., AlexNet, GoogLeNet, RESNET, etc.) that is commonly used.
[0056] It is contemplated that a deep learning neural network / model (e.g., CNN) refers to a combination of artificial neural networks for analyzing, training, and reasoning about any range of input data. For example, CNNs are much faster and may require relatively less data processing than conventional algorithms. It is also contemplated that once input data is received at a CNN, the data may then be processed by layers (e.g., convolutional layers, pooling layers, rectified linear unit (ReLU) layers, fully connected layers, loss / output layers, etc.), where each layer performs specific processing tasks for training and reasoning purposes.
[0057] For example, a convolutional layer can be viewed as a core layer with many learnable filters or kernels with receptive fields extending through the full depth of the input capacity. The convolutional layer can start processing any data from the input and move to another layer (e.g., a pooling layer) where a form of non-linear downsampling is performed, where, for example, these non-linear downsampling functions can implement pooling (e.g., max pooling). Similarly, data is further processed and trained at a ReLU layer, which applies a non-saturating activation function to increase the non-linear nature of the decision function and the network without affecting the receptive field of the convolutional layer.
[0058] Although embodiments are not limited to any number or type of layers of a neural network (e.g., CNN), the training process can continue with a fully connected layer, where high-level reasoning is provided after several convolutional and pooling layers. In other words, a CNN can receive input data and perform feature mapping, sampling, convolution, subsampling, and then output the result.
[0059] For example, the loss / output layer can indicate how the training is biased against deviations between the predicted labels and the true labels, wherein the loss / output layer can be viewed as the last layer in the CNN. For example, a softmax loss can be used to predict a single class among mutually exclusive classes. In addition, in one embodiment, the classification and prediction logic 207 (e.g., the softmax and classification layers of the loss / output layer) can be used for classification and prediction purposes, wherein the two layers are generated by the softmax layer and the classification layer functions, respectively.
[0060] In one embodiment, after all data from the inputs have been processed for training and inference, the classification and prediction logic 207 can then be used to identify which of the sensors (e.g., cameras 242A-242D) may have a problem. Once identified, the classification and prediction logic 207 can issue a notification about the bad camera in the cameras 242A-242D, for example, displaying a notification at the display device 244, listening to it through the speaker device 243, etc. In one embodiment, the notification can then be used, for example, by a user to get the defective camera in the cameras 242A-242D and fix the problem (e.g., remove mud from the lens, manually or automatically fix any technical errors with the lens, replace the defective camera in the cameras 242-242D with another one, etc.).
[0061] In one embodiment, a specific tag may be used for notification purposes, for example, tag: 0 may indicate that all sensors are good, while tag: 1 may indicate that the first sensor is damaged, tag: 2 may indicate that the second sensor is damaged, tag: 3 may indicate that the third sensor is damaged, tag: 4 may indicate that the fourth sensor is damaged, and so on. Similarly, tag: 1 may indicate that the first sensor is good, tag: 2 may indicate that the second sensor is good, and so on. It is contemplated that embodiments are not limited to any form of notification, and any or a combination of words, numbers, images, videos, audio, etc. may be used to convey the result of whether the sensors are working well.
[0062] In addition, for example, where a single input data layer is associated with each of the cameras 242A-242D, multiple channels (e.g., 12 (3*4=12) channels in the case of four cameras 242A-242D) can provide all the data for the four images corresponding to the four cameras 242A-242D, where the data can be randomly loaded by scrambling the order of the channels. Using this data, a deep learning model (e.g., a CNN) can calculate loss (during training) as well as accuracy (during validation), so when it comes to prediction, there is no need to use labels. In some embodiments, the training data can include a large sampling of images (e.g., thousands or tens of thousands of sampled images per camera 242A-242D), and the validation data can also include a large sampling of images (e.g., hundreds or thousands of sampled images per camera 242A-242D) to provide robust training / inference on the data facilitated by the training and inference data, and then obtain accurate results, including recognition, prediction, etc. facilitated by the classification and prediction logic 207.
[0063] The capture / sensing component 231 may also include any number and type of cameras 242A, 242B, 242C, 242D (e.g., depth sensing cameras or capture devices (e.g., RealSense TM Depth sensing cameras)). These images with depth information have been effectively used for various computer vision and computational photography effects (e.g., but not limited to, scene understanding, refocusing, compositing, cinema graphics, etc.). Similarly, for example, the display can include any number and type of displays (e.g., integral displays, tensor displays, stereoscopic displays, etc.), including but not limited to embedded or connected display screens, display devices, projectors, etc.
[0064] The capture / sensing component 231 may also include one or more of the following: a vibration component, a tactile component, a conductive element, a biometric sensor, a chemical detector, a signal detector, an electroencephalogram, functional near-infrared spectroscopy, a wave detector, a force sensor (e.g., an accelerometer), an illuminator, an eye tracking or gaze tracking system, head tracking, etc., which can be used to capture any amount or type of visual data (e.g., images (e.g., photos, videos, movies, audio / video streams, etc.)) and non-visual data (e.g., audio streams or signals (e.g., sounds, noise, vibrations, ultrasound, etc.), radio waves (e.g., wireless signals (e.g., wireless signals with data, metadata, symbols, etc.)), chemical changes or properties (e.g., humidity, body temperature, etc.), biometric readings (e.g., fingerprints, etc.), brain waves, cerebral blood circulation, environmental / meteorological conditions, maps, etc.). It is expected that "sensors" and "detectors" can be referenced interchangeably throughout this document. It is also contemplated that one or more capture / sensing components 231 may also include one or more of supporting or supplemental devices (e.g., illuminators (e.g., IR illuminators), light fixtures, generators, sound blockers, etc.) for capturing and / or sensing data.
[0065] It is also contemplated that in one embodiment, the capture / sensing component 231 may also include any number and type of context sensors (e.g., linear accelerometers) for sensing or detecting any number and type of context (e.g., estimating horizontality, linear acceleration, etc. associated with a mobile computing device, etc.). For example, the capture / sensing component 231 may include any number and type of sensors, such as (but not limited to): accelerometers (e.g., linear accelerometers for measuring linear acceleration, etc.); inertial devices (e.g., inertial accelerometers, inertial gyroscopes, microelectromechanical systems (MEMS) gyroscopes, inertial navigation units, etc.); and gravity gradiometers for studying and measuring changes in gravitational acceleration due to gravity, etc.
[0066] In addition, for example, the capture / sensing component 231 may include (but is not limited to): audio / visual devices (e.g., camera microphones, speakers, etc.); context-aware sensors (e.g., temperature sensors, facial expression and feature measurement sensors working with one or more cameras of the audio / visual device, environmental sensors (e.g., for sensing background color, light, etc.); biometric sensors (e.g., for detecting fingerprints, etc.), calendar maintenance and reading devices, etc.); global positioning system (GPS) sensors; resource requesters; and / or TEE logic. The TEE logic may be employed separately or be part of the resource requester and / or I / O subsystem, etc. The capture / sensing component 231 may also include voice recognition devices, photo recognition devices, facial and other body recognition components, voice to text conversion components, etc.
[0067] Similarly, the output component 233 may include a dynamic tactile touch screen with a tactile effector as an example of presenting a visualization of touch, wherein an embodiment thereof may be an ultrasonic generator that can send a signal in space that, when reaching, for example, a human finger, can produce a tactile sensation or similar sensation on the finger. In addition, for example, and in one embodiment, the output component 233 may include (but is not limited to) one or more of the following items: a light source, a display device and / or screen, an audio speaker, a tactile component, a conductive element, a bone conduction speaker, a smell or smell visual and / or non-visual presentation device, a touch or touch visual and / or non-visual presentation device, an animation display device, a biometric display device, an X-ray display device, a high-resolution display, a high dynamic range display, a multi-view display, and a head-mounted display (HMD) for at least one of virtual reality (VR) and augmented reality (AR), etc.
[0068] It is contemplated that the embodiments are not limited to any particular number or type of use case scenarios, architectural placements, or component arrangements; however, for simplicity and clarity, illustrations and descriptions are provided and discussed throughout this document for exemplary purposes, but the embodiments are not limited thereto. Additionally, throughout this document, a "user" may refer to someone who has access to one or more computing devices (e.g., computing device 100), and may be referenced interchangeably with a "person," "individual," "human being," "he," "she," "child," "adult," "viewer," "player," "player," "developer," "programmer," and the like.
[0069] The communication / compatibility logic 209 can be used to facilitate the communication of various components, networks, computing devices, databases 225, and / or communication media 230, etc. with any number and type of other computing devices (e.g., wearable computing devices, mobile computing devices, desktop computers, server computing devices, etc.), processing devices (e.g., central processing units (CPUs), graphics processing units (GPUs), etc.), capture / sensing components (e.g., non-visual data sensors / detectors (e.g., audio sensors, olfactory sensors, tactile sensors, signal sensors, vibration sensors, chemical detectors, radio wave detectors, force sensors, weather / temperature sensors, body / biometric sensors, scanners, etc.), etc.), and / or other devices. scanner, etc.) and visual data sensors / detectors (e.g., cameras, etc.), user / context awareness components and / or identification / authentication sensors / devices (e.g., biometric sensors / detectors, scanners, etc.), memory or storage devices, data sources and / or databases (e.g., data storage devices, hard drives, solid-state drives, hard disks, memory cards or devices, memory circuits, etc.), networks (e.g., cloud networks, the Internet, the Internet of Things, intranets, cellular networks, proximity networks (e.g., Bluetooth, Bluetooth Low Energy (BLE), Bluetooth Smart, Wi-Fi proximity, radio frequency identification, near field communication, body area networks, etc.)), wireless or wired communications and related protocols (e.g., Dynamic communication and compatibility between the wireless networks (e.g., WiMAX, Ethernet, etc.), connection and location management technologies, software applications / websites (e.g., social and / or business networking websites, business applications, games and other entertainment applications, etc.), programming languages, etc., while ensuring compatibility with changing technologies, parameters, protocols, standards, etc.
[0070] Throughout this document, terms such as "logic," "component," "module," "framework," "engine," "tool," "circuit," and the like may be used interchangeably and include, by way of example, software, hardware, and / or any combination of software and hardware (e.g., firmware). In one example, "logic" may refer to or include a software component capable of operating on one or more of an operating system, a graphics driver, and the like of a computing device (e.g., computing device 100). In another example, "logic" may refer to or include a hardware component capable of being physically installed along with or becoming a part of system hardware elements (e.g., an application processor, a graphics processor, and the like) of one or more computing devices (e.g., computing device 100). In yet another embodiment, "logic" may refer to or include a firmware component capable of becoming a part of system firmware (e.g., firmware of an application processor or a graphics processor, and the like) of a computing device (e.g., computing device 100).
[0071] In addition, specific brands, words, terms, phrases, names and / or abbreviations (e.g., “sensor,” “camera,” “autonomous machine,” “sensor automated inspection,” “deep learning,” “convolutional neural network,” “concatenation,” “training,” “inference,” “classification,” “prediction,” “RealSense TM Any use of the terms "camera", "real-time", "automated", "dynamic", "user interface", "camera", "sensor", "microphone", "display", "speaker", "verification", "authentication", "privacy", "user", "user profile", "user preference", "transmitter", "receiver", "personal device", "smart device", "mobile computer", "wearable device", "IoT device", "proximate network", "cloud network", "server computer", etc.) should not be construed as limiting the embodiments to software or devices carrying that label in a product or in a document outside of this document.
[0072] It is contemplated that any number and type of components may be added and / or removed from the automated inspection mechanism 110 to facilitate various embodiments including adding, removing, and / or enhancing specific features. For simplicity, clarity, and ease of understanding of the automated inspection mechanism 110, many standard and / or well-known components (e.g., components of a computing device) are not shown or discussed herein. It is contemplated that the embodiments described herein are not limited to any technology, topology, system, architecture, and / or standard, and are flexible enough to accept and adapt to any feature changes.
[0073] Figure 3A According to one embodiment and as previously referred to Figure 2 For the sake of brevity, the following may not discuss or repeat the previous references. Figure 1-Figure 2 A lot of details described.
[0074] In the illustrated embodiment, four images of a scene A 301, B 303, C 305, and D 307 are shown as being represented by Figure 2 242B, C 242C, and D 242D, wherein the plurality of images 301-307 are based on static data captured by the four cameras 242A-242D over a period of time. For example, sensors (e.g., cameras 242A-242D, radar, etc.) can be used to capture similar data for the same sensing purpose, such as for an automated driving vehicle that is to be aware of objects near or around it.
[0075] In this embodiment, for design and testing purposes only, the four cameras 242A-242D are shown as capturing four images 301-307 of the same scene simultaneously but from different angles and / or positions. In addition, as shown, one of the images (e.g., image 301) shows the corresponding camera 242A having clarity issues (e.g., due to some kind of stain 309 (e.g., mud, dirt, debris, etc.) on the lens of camera 242A). It is expected that these issues may cause a lot of problems when autonomous machines (e.g., self-driving vehicles) are involved.
[0076] In one embodiment, as shown in FIG. Figure 2 As discussed above, by collecting a large amount of data (e.g., several thousand data inputs) and using them as training data, validation data, etc. in a deep learning model (e.g., CNN), Figure 1 As facilitated by the automated inspection mechanism 110 of FIG. 1 , real-time detection of stain 309 is permitted. This real-time detection then permits real-time notification and real-time correction of stain 309 so that any imperfections with camera 242A can be repaired and all cameras 242A-242D can realize their potential and collect data to enable the use of autonomous machines (e.g., Figure 1 The autonomous machine 100) is safe, reliable and efficient.
[0077] Figure 3B According to one embodiment and as previously referred to Figure 2 The dynamic input from a single sensor described herein. For the sake of brevity, the previously referenced Figure 1-Figure 3A A lot of details described.
[0078] In the illustrated embodiment, a single sensor (e.g. Figure 2 Camera D 242D) can be used to capture four images A 311, B 313, C 315, D 317 of a single scene but with different time stamps (e.g., at different points in time). In the illustrated pattern, capturing the scene at different points in time shows the scene as moving (e.g., from right to left), while the stain 319 is shown as being located in one location (e.g., in one spot on the lens of camera D 242D).
[0079] As revealed in the four images 311-317, over time, the object 321 (e.g., a book) captured by the camera 242D appears to be moving (e.g., from right to left), while the stain 319 is stationary (or in a real sense, slowly moving in a different pattern, such as Figure 3C As shown in the embodiment of FIG. Figure 2 and Figure 3AAs described, thousands of images are collected and input into the deep learning model for training and validation purposes, which then enables testing of the deep learning model. Once tested, the deep learning model can be used to identify and correct sensor problems (e.g., stain 319 on camera 242D) in real time.
[0080] Figure 3C According to one embodiment and as previously referred to Figure 2 The dynamic input from a single sensor described herein. For the sake of brevity, the previously referenced Figure 1-Figure 3B A lot of details described.
[0081] In one embodiment, as with respect to having a dynamic input reference through a single sensor Figure 3B As described, in this illustrated embodiment, a single sensor (e.g., camera B 242B) captures four images A 331, B 333, C 335, D 337 of a single scene, wherein a spot 339 on the lens of camera 242B is shown moving with an object 341 (e.g., a book) in the background scene. For example, the spot 339 may be a piece of mud on the lens of camera 242B that moves downward over time due to gravity or sideways due to wind, movement of camera 242B, etc.
[0082] In one embodiment, as described above, this data regarding the stain 339 and its movement may be captured by one or more sensors (e.g., the camera 242B itself) and input into a trained deep learning neural network / model (e.g., a CNN), which then predicts and provides in real time the exact location of the stain 339, the sensors affected by the stain 339 (e.g., the camera 242B), and how to correct the problem (e.g., how to remove the stain 339 from the lens of the camera 242B).
[0083] In one embodiment, such training of the deep learning neural network / model is achieved by inputting a set of (several thousand) of these inputs as examples for training and validation data and testing the deep learning model. For example, the deep learning model may use a deep learning neural network (e.g., CNN) to first extract features of all sensors (e.g., camera 242B), then fuse the data and use a classifier to identify problematic sensors (e.g., camera 242B). For example, camera 242B may be assigned a label (e.g., label 2: second sensor is damaged, etc.).
[0084] Figure 4A An architectural arrangement 400 is shown for providing a sequence of transactions for real-time detection and correction of problematic sensors using deep learning according to one embodiment. Figure 1-3C Any process or transaction may be performed by processing logic, which may include hardware (e.g., circuitry, dedicated logic, programmable logic, etc.), software (e.g., instructions running on a processing device), or a combination thereof, such as Figure 1 As facilitated by the automatic inspection mechanism 110 of the embodiment. For simplicity and clarity in the presentation, any process or transaction associated with the description may be shown or stated in a linear order; however, it is contemplated that any number of processes or transactions may be performed in parallel, asynchronously, or in a different order.
[0085] In one embodiment, the transaction sequence at the architecture setup 400 begins at 401: inputting sequential multiple times of captured data from one or more sensors for concatenation and then inputting it into a deep learning model (e.g., deep learning model 421). The data may include multiple inputs based on data captured by any number and type of sensors (e.g., cameras, LiDAR, radar, etc.), with some degree of overlap in their detections (e.g., when detecting the same object in a scene). At 403, as referenced Figure 2 As described by the concatenation logic 203, data from these inputs are sent for concatenation so that these multiple inputs are then concatenated into a single input and sent to the data learning model 421 for training (405) and inference (407).
[0086] It is contemplated that in one embodiment, concatenation is performed before or outside of sending the data to the deep learning model 421 so that there is greater flexibility in setting the order to gain the most benefit from training (405). As shown, in one embodiment, reasoning (407) can be part of training (405), or in another embodiment, reasoning (407) and training (405) can be performed separately.
[0087] In one embodiment, once input into the deep learning model 421, it is then input into and processed by the CNN 409, where the data is processed through multiple layers, as described with reference to Figure 2 For example, at the classification layer 411, a common classification layer (e.g., a fully connected layer, a softmax layer, etc.) may be included, as described in reference Figure 2 As further described.
[0088] In one embodiment, the transaction sequence provided by the architecture setup 400 may continue with: obtaining results 413 through the output layer, which may identify or predict whether one or more sensors are technically defective or blocked by objects or debris, or are not working for any reason. Once the results 413 have been obtained, the various labels 415 are compared to determine the loss and the appropriate label for providing the user with respect to the one or more defective sensors. The transaction sequence may continue with: back propagation of data 417, and therefore, performing more weighted updates 419 all at the deep learning model 421.
[0089] Figure 4B A method 450 for real-time detection and correction of problematic sensors using deep learning is shown according to one embodiment. For the sake of brevity, the previously referenced Figure 1-4A Any process or transaction may be performed by processing logic, which may include hardware (e.g., circuitry, dedicated logic, programmable logic, etc.), software (e.g., instructions running on a processing device), or a combination thereof, such as Figure 1 As facilitated by the automatic inspection mechanism 110 of the embodiment. For simplicity and clarity in the presentation, any process or transaction associated with the description may be shown or stated in a linear order; however, it is contemplated that any number of processes or transactions may be performed in parallel, asynchronously, or in a different order.
[0090] The method 450 begins at box 451: detecting data including one or more images of a scene captured by one or more sensors (e.g., cameras) simultaneously or over a period of time, wherein, where the data is spread over multiple inputs, the multiple inputs are provided for concatenation. At box 453, the multiple inputs are concatenated into a single input of the data and provided to a deep learning model for further processing (e.g., training, inference, validation, etc.). At box 455, the data is received at a deep learning model for training and inference, wherein the deep learning model includes a neural network (e.g., CNN) having multiple processing layers.
[0091] As expected and as referenced Figure 2 As discussed, the data that has gone through the training and inference phases may be processed and modified at several levels, including at the CNN (which may include multiple processing layers of its own). In one embodiment, at block 457, the trained deep learning model classifies the data and predicts an outcome based on all the processing and classification. For example, the prediction of the outcome may indicate and identify in real time whether any one or more sensors are defective or blocked, so that the defective or blocked sensors may be repaired in real time.
[0092] Figure 5FIG. 5 shows a computing device 500 according to one implementation. The computing device 500 shown may be used with Figure 1 The computing device 500 may be the same or similar to the computing device 100 of the present invention. The computing device 500 houses a system board 502. The board 502 may include a plurality of components, including but not limited to a processor 504 and at least one communication package 506. The communication package is coupled to one or more antennas 516. The processor 504 is physically and electrically coupled to the board 502.
[0093] Depending on its application, the computing device 500 may include other components that may or may not be physically and electrically coupled to the board 502. These other components include, but are not limited to, volatile memory (e.g., DRAM) 508, non-volatile memory (e.g., ROM) 509, flash memory (not shown), graphics processor 512, digital signal processor (not shown), encryption processor (not shown), chipset 514, antenna 516, display 518 (e.g., touch screen display), touch screen controller 520, battery 522, audio codec (not shown), video codec (not shown), power amplifier 524, global positioning system (GPS) device 526, compass 528, accelerometer (not shown), gyroscope (not shown), speaker 530, camera 532, microphone array 534, and mass storage device (e.g., hard disk drive) 510, compact disk (CD) (not shown), digital versatile disk (DVD) (not shown), etc. These components may be connected to the system board 502, mounted to the system board, or combined with any other components.
[0094] The communication package 506 enables wireless and / or wired communication for transferring data to and from the computing device 500. The term "wireless" and its derivatives may be used to describe circuits, devices, systems, methods, techniques, communication channels, etc. that can transfer data via a non-solid medium using modulated electromagnetic radiation. Although the term does not imply that the associated devices do not contain any wires, in some embodiments they may not contain any wires. The communication package 506 may implement any number of wireless or wired standards or protocols, including but not limited to WiFi (IEEE 802.11 family), WiMAX (IEEE 802.16 family), IEEE 802.20, Long Term Evolution (LTE), Ev-DO, HSPA+, HSDPA+, HSUPA+, EDGE, GSM, GPRS, CDMA, TDMA, DECT, Bluetooth, Ethernet and its derivatives, and any other wireless and wired protocols referred to as 3G, 4G, 5G and above. The computing device 500 may include multiple communication packages 506. For example, the first communication packet 506 may be dedicated to shorter range wireless communications (e.g., Wi-Fi and Bluetooth), while the second communication packet 506 may be dedicated to longer range wireless communications (e.g., GPS, EDGE, GPRS, CDMA, WiMAX, LTE, Ev-DO, etc.).
[0095] The camera 532, including any depth sensor or proximity sensor, is coupled to an optional image processor 536 to perform conversions, analysis, noise reduction, comparisons, depth or distance analysis, image understanding, and other processing described herein. The processor 504 is coupled to the image processor to drive processing through interrupts, set parameters, and control operations of the image processor and camera. The image processing may instead be performed in the processor 504, the graphics CPU 512, the camera 532, or in any other device.
[0096] In various implementations, computing device 500 may be a laptop, a netbook, a notebook, an ultrabook, a smartphone, a tablet, a personal digital assistant (PDA), an ultra-mobile PC, a mobile phone, a desktop computer, a server, a set-top box, an entertainment control unit, a digital camera, a portable music player, or a digital video recorder. The computing device may be fixed, portable, or wearable. In other implementations, computing device 500 may be any other electronic device that processes data or records data for processing elsewhere.
[0097] Embodiments may be implemented using one or more memory chips, controllers, CPUs (central processing units), microchips or integrated circuits interconnected using a motherboard, application specific integrated circuits (ASICs), and / or field programmable gate arrays (FPGAs). The term "logic" may include, by way of example, software or hardware and / or a combination of software and hardware.
[0098] References to "one embodiment," "an embodiment," "example embodiment," "various embodiments," etc. indicate that the described embodiment may include a particular feature, structure, or characteristic, but not every embodiment must include the particular feature, structure, or characteristic. Additionally, some embodiments may have some, all, or none of the features described for other embodiments.
[0099] In the following description and claims, the term "coupled," along with its derivatives, may be used. "Coupled" is used to indicate that two or more elements co-operate or interact with each other, but they may or may not have intervening physical or electrical components between them.
[0100] As used in the claims, unless otherwise specified, the use of ordinal adjectives "first," "second," "third," etc., to describe common elements merely indicates that different instances of the same element are referred to and is not intended to imply that the elements described must be in a given order temporally, spatially, in ranking, or in any other manner.
[0101] The accompanying drawings and the foregoing description give examples of embodiments. Those skilled in the art will appreciate that one or more of the elements described may be well combined into a single functional element. Alternatively, certain elements may be separated into multiple functional elements. Elements from one embodiment may be added to another embodiment. For example, the order of the processing described herein may be changed and is not limited to the manner described herein. In addition, the actions of any flowchart do not need to be implemented in the order shown; nor do they necessarily need to perform all actions. In addition, those actions that do not depend on other actions may be performed in parallel with other actions. The scope of the embodiments is by no means limited by these specific examples. Whether or not explicitly given in the specification, many variations are possible, such as differences in structure, size, and use of materials. The scope of the embodiments is at least as wide as given in the following claims.
[0102] Embodiments may be provided, for example, as a computer program product, which may include one or more transient or non-transient machine-readable storage media having machine-executable instructions stored thereon, which when executed by one or more machines (e.g., a computer, a network of computers, or other electronic devices) may cause one or more machines to perform operations according to the embodiments described herein. The machine-readable medium may include, but is not limited to, a floppy disk, an optical disk, a CD-ROM (compact disk-read only memory) and a magnetic optical disk, ROM, RAM, EPROM (erasable programmable read-only memory), EEPROM (electrically erasable programmable read-only memory), a magnetic or optical card, a flash memory, or other types of media / machine-readable media suitable for storing machine-executable instructions.
[0103] Figure 6 An embodiment of a computing environment 600 capable of supporting the operations discussed above is shown. Figure 5 Various different hardware architectures and forms of implementing modules and systems are shown.
[0104] The command execution module 601 includes a central processing unit to cache and execute commands and distribute tasks between the other modules and the system shown. It can include an instruction stack, a cache memory for storing intermediate results and final results, and a mass memory for storing applications and operating systems. The command execution module can also serve as a central coordination and task distribution unit for the system.
[0105] The screen rendering module 621 draws objects on one or more screens for the user to view. It can be adapted to receive data from the virtual object behavior module 604, as described below, and render virtual objects and any other objects and forces on the appropriate one or more screens. Therefore, the data from the virtual object behavior module will determine, for example, the position and dynamics of the virtual object and the associated gestures, forces and objects, and the screen rendering module will depict the virtual object and the associated objects and environment on the screen accordingly. The screen rendering module can be further adapted to receive data from the adjacent screen perspective module 607, as described below, if the virtual object can be moved to the display of the device associated with the adjacent screen perspective module, then depict the target landing area for the virtual object. Therefore, for example, if the virtual object is moving from the main screen to the auxiliary screen, the adjacent screen perspective module 2 can send data to the screen rendering module to suggest one or more target landing areas for the virtual object on the track to the user's hand movement or eye movement, for example in the form of a shadow.
[0106] The object and gesture recognition module 622 can be adapted to recognize and track the user's hand and arm gestures. The module can be used to recognize hands, fingers, finger gestures, hand movements, and the position of hands relative to the display. For example, the object and gesture recognition module can, for example, determine that the user makes a body part gesture to throw or toss a virtual object onto one or another of the multiple screens, or the user makes a body part gesture to move a virtual object to the border of one or another of the multiple screens. The object and gesture recognition system can be coupled to a camera or camera array, a microphone or microphone array, a touch screen or touch surface, or a pointing device, or some combination of these items to detect gestures and commands from the user.
[0107] The touch screen or touch surface of the object and gesture recognition system may include a touch screen sensor. Data from the sensor can be fed to hardware, software, firmware, or a combination thereof to map the touch gestures of the user's hand on the screen or surface to corresponding dynamic behaviors of the virtual object. The sensor date can be used for momentum and inertia factors to allow various momentum behaviors for virtual objects based on input from the user's hand (e.g., the sweep rate of the user's finger relative to the screen). A pinch gesture can be interpreted as a command for lifting a virtual object from the display screen or for starting to generate a virtual binding associated with the virtual object or for zooming in or out on the display. The object and gesture recognition system can generate similar commands using one or more cameras without the aid of a touch surface.
[0108] The attention direction module 623 can be equipped with a camera or other sensor to track the position or orientation of the user's face or hands. When a gesture or voice command is issued, the system can determine the appropriate screen for the gesture. In one example, a camera is mounted near each display to detect whether the user is facing the display. If so, the attention direction module information is provided to the object and gesture recognition module 622 to ensure that the gesture or command is associated with the appropriate library for effective display. Similarly, if the user is not looking at all screens, the command can be omitted.
[0109] The device proximity detection module 625 can use proximity sensors, compasses, GPS (Global Positioning System) receivers, personal area network radios, and other types of sensors along with triangulation and other techniques to determine the proximity of other devices. Once a nearby device is detected, it can be registered with the system, and its type can be determined as an input device or a display device or both. For input devices, the received data can then be applied to the object gesture and recognition module 622. For display devices, it can be considered by the adjacent screen perspective module 607.
[0110] The virtual object behavior module 604 is adapted to receive input from the object speed and direction module and apply the input to the virtual object being displayed in the display. Thus, for example, the object and gesture recognition system will interpret the user gestures and by mapping the captured movements of the user's hand to the recognized movements, the virtual object tracker module will associate the position and movement of the virtual object with the movement recognized by the object and gesture recognition system, the object and speed and direction module will capture the dynamics of the movement of the virtual object, and the virtual object behavior module will receive input from the object and speed and direction module to generate data that will guide the movement of the virtual object to correspond to the input from the object and speed and direction module.
[0111] On the other hand, the virtual object tracker module 606 can be adapted to track where the virtual object should be located in the three-dimensional space near the display and which body part of the user is holding the virtual object based on the input from the object and gesture recognition module. The virtual object tracker module 606 can, for example, track the virtual object as it moves on and between screens, and track which body part of the user is holding the virtual object. Tracking the body part that is holding the virtual object allows continuous knowledge of the aerial movement of the body part, and therefore ultimately whether the virtual object has been released onto one or more screens.
[0112] The gesture to view and screen synchronization module 608 receives a selection of a view and a screen or both from the focus direction module 623 and, in some cases, receives a voice command for determining which view is the active view and which screen is the active screen. It then causes the loading of the relevant gesture library for use with the object and gesture recognition module 622. Various views of an application on one or more screens may be associated with an alternative gesture library or set of gesture templates for a given view.
[0113] The adjacent screen perspective module 607, which may include or be coupled to the device proximity detection module 625, may be adapted to determine the angle and position of one display relative to another display. Projection displays include, for example, images projected onto a wall or screen. The ability to detect the proximity of a nearby screen and the corresponding angle or orientation projected from it may be achieved, for example, by either infrared transmitters and receivers or electromagnetic or light detection sensing capabilities. For technologies that allow projection displays with touch input, incoming video may be analyzed to determine the position of the projection display and correct the distortion produced by displaying at an angle. Accelerometers, magnetometers, compasses, or cameras may be used to determine the angle at which the device is being held, while infrared transmitters and cameras may allow the orientation of the screen device to be determined in conjunction with sensors on adjacent devices. The adjacent screen perspective module 607 may determine the coordinates of adjacent screens relative to their own screen coordinates in this manner. Therefore, the adjacent screen perspective module may determine which devices are close to each other and other potential targets for moving one or more virtual objects on the screen. The adjacent screen perspective module may further allow the position of the screen to be associated with a model of a three-dimensional space representing all existing objects and virtual objects.
[0114] The object and speed and direction module 603 may be adapted to estimate the dynamics of the moving virtual object (e.g., its trajectory (linear or angular), momentum (linear or angular), etc.) by receiving input from the virtual object tracker module. The object and speed and direction module may be further adapted to estimate the dynamics of any physical forces by, for example, estimating the acceleration, deflection, extension, etc. of the virtual bindings and the dynamic behavior of the virtual object when the user's body part is released. The object and speed and direction module may also estimate the velocity of the object (e.g., the velocity of the hand and fingers) using image motion, size and angle changes.
[0115] The image motion, image size, and angular changes of objects in the image plane or in three-dimensional space can be used by the momentum and inertia module 602 to estimate the velocity and direction of objects in space or on a display. The momentum and inertia module is coupled to the object and gesture recognition module 622 to estimate the velocity of gestures performed by hands, fingers, and other body parts, and then apply those estimates to determine the momentum and velocity of virtual objects that will be affected by the gesture.
[0116] 3D image interaction and effect module 605 tracks the interaction of the user and the 3D image that appears to extend out one or more screens. The influence of the object in the z-axis (towards and leaving the plane of the screen) and the relative influence of these objects on each other can be calculated. For example, before the virtual object arrives at the plane of the screen, the object thrown by the user's gesture can be affected by the 3D object in the foreground. The object can change the direction or speed of the projection or completely destroy it. The 3D image interaction and effect module can render objects in the foreground on one or more displays. As shown, each component (e.g., components 601, 602, 603, 604, 605, 606, 607 and 608) is connected via interconnection or bus (e.g., bus 609).
[0117] The following clauses and / or examples pertain to further embodiments or examples. The details in the examples may be used anywhere in one or more embodiments. Various features of different embodiments or examples may be combined differently, with some features included and other features not included, to suit a variety of different applications. Examples may include subject matter such as a method, a module for performing the actions of the method, at least one machine-readable medium including instructions that, when executed by a machine, cause the machine to perform the actions of the method, or an apparatus or system for facilitating hybrid communications according to the embodiments and examples described herein.
[0118] Some embodiments pertain to Example 1, which includes an apparatus for facilitating deep learning-based real-time detection and correction of problematic sensors in an autonomous machine, the apparatus comprising: detection and capture logic for facilitating one or more sensors to capture one or more images of a scene, wherein an image in the one or more images is determined to be unclear, wherein the one or more sensors include one or more cameras; and classification and prediction logic for facilitating a deep learning model to identify a sensor associated with the image in real-time.
[0119] Example 2 includes the subject matter described in Example 1, further comprising: concatenation logic for receiving one or more data inputs associated with the one or more images to concatenate the one or more data inputs into a single data input to be processed by the deep learning model, wherein the device includes an autonomous machine, the autonomous machine including one or more of a self-driving vehicle, an autonomous flying vehicle, an autonomous sailing vehicle, and an autonomous home appliance.
[0120] Example 3 includes the subject matter described in Examples 1-2, further comprising: training and inference logic for facilitating the deep learning model to receive the single data input to perform one or more deep learning processes including a training process and an inference process to obtain real-time identification of a sensor associated with an unclear image, wherein the sensor includes a camera.
[0121] Example 4 includes the subject matter described in Examples 1-3, wherein the training and inference logic is further used to: facilitate the deep learning model to receive multiple data inputs and run the multiple data inputs through the training process and the inference process so that real-time identification of the sensor is accurate and timely.
[0122] Example 5 includes the subject matter of Examples 1-4, wherein the deep learning model comprises one or more neural networks, the one or more neural networks comprising one or more convolutional neural networks, wherein the image is obscured due to one or more of a technical flaw in the sensor or a physical obstruction to the sensor, wherein the physical obstruction is due to a person, plant, animal, or object blocking the sensor, or dirt, smudges, mud, or debris covering a portion of a lens of the sensor.
[0123] Example 6 includes the subject matter of Examples 1-5, wherein the classification and prediction logic is used to provide one or more of real-time notification of unclear images and real-time automatic correction of the sensor.
[0124] Example 7 includes the subject matter of Examples 1-6, wherein the apparatus comprises one or more processors having a graphics processor co-located on a common semiconductor package with an application processor.
[0125] Some embodiments pertain to Example 8, which includes a method for facilitating deep learning-based real-time detection and correction of problematic sensors in an autonomous machine, the method comprising: facilitating one or more sensors to capture one or more images of a scene, wherein an image in the one or more images is determined to be unclear, wherein the one or more sensors include one or more cameras of a computing device; and facilitating a deep learning model to identify, in real-time, a sensor associated with the image.
[0126] Example 9 includes the subject matter described in Example 8, further comprising: receiving one or more data inputs associated with the one or more images to concatenate the one or more data inputs into a single data input to be processed by the deep learning model, wherein the device includes an autonomous machine, the autonomous machine including one or more of a self-driving vehicle, an autonomous flying vehicle, an autonomous sailing vehicle, and an autonomous home appliance.
[0127] Example 10 includes the subject matter of Examples 8-9, further comprising: facilitating the deep learning model to receive the single data input to perform one or more deep learning processes including a training process and an inference process to obtain real-time recognition of a sensor associated with an unclear image, wherein the sensor includes a camera.
[0128] Example 11 includes the subject matter of Examples 8-10, wherein the deep learning model is further used to: receive multiple data inputs and run the multiple data inputs through the training process and the inference process so that real-time identification of the sensor is accurate and timely.
[0129] Example 12 includes the subject matter of Examples 8-11, wherein the deep learning model comprises one or more neural networks, the one or more neural networks comprising one or more convolutional neural networks, wherein the image is obscured due to one or more of a technical flaw in the sensor or a physical obstruction to the sensor, wherein the physical obstruction is due to a person, plant, animal, or object blocking the sensor, or dirt, smudges, mud, or debris covering a portion of a lens of the sensor.
[0130] Example 13 includes the subject matter of Examples 8-12, further comprising providing one or more of real-time notification of unclear images and real-time automatic correction of the sensor.
[0131] Example 14 includes the subject matter of Examples 8-13, wherein the computing device comprises one or more processors having a graphics processor co-located on a common semiconductor package with an application processor.
[0132] Some embodiments pertain to Example 15, which includes a data processing system comprising a computing device having a memory coupled to a processing device, the processing device being configured to: facilitate one or more sensors to capture one or more images of a scene, wherein an image in the one or more images is determined to be unclear, wherein the one or more sensors include one or more cameras; and facilitate a deep learning model to identify, in real time, a sensor associated with the image.
[0133] Example 16 includes the subject matter of Example 15, wherein the processing device is further used to: receive one or more data inputs associated with the one or more images to concatenate the one or more data inputs into a single data input to be processed by the deep learning model, wherein the device includes an autonomous machine, the autonomous machine including one or more of a self-driving vehicle, an autonomous flying vehicle, an autonomous sailing vehicle, and an autonomous home appliance.
[0134] Example 17 includes the subject matter of Examples 15-16, wherein the processing device is further used to: facilitate the deep learning model to receive the single data input to perform one or more deep learning processes including a training process and an inference process to obtain real-time identification of a sensor associated with an unclear image, wherein the sensor includes a camera.
[0135] Example 18 includes the subject matter of Examples 15-17, wherein the deep learning model is further used to: receive multiple data inputs and run the multiple data inputs through the training process and the inference process so that real-time identification of the sensor is accurate and timely.
[0136] Example 19 includes the subject matter of Examples 15-18, wherein the deep learning model comprises one or more neural networks, the one or more neural networks comprising one or more convolutional neural networks, wherein the image is obscured due to one or more of a technical flaw in the sensor or a physical obstruction to the sensor, wherein the physical obstruction is due to a person, plant, animal, or object blocking the sensor, or dirt, smudges, mud, or debris covering a portion of a lens of the sensor.
[0137] Example 20 includes the subject matter of Examples 15-19, wherein the processing device is further configured to provide one or more of real-time notification of unclear images and real-time automatic correction of the sensor.
[0138] Example 21 includes the subject matter of Examples 15-20, wherein the computing device comprises one or more processors having a graphics processor co-located on a common semiconductor package with an application processor.
[0139] Some embodiments belong to Example 22, which includes a device for facilitating simultaneous recognition and processing of multiple voices from multiple users, the device comprising: a module for facilitating one or more sensors to capture one or more images of a scene, wherein an image in the one or more images is determined to be unclear, wherein the one or more sensors include one or more cameras; and a module for facilitating a deep learning model to identify a sensor associated with the image in real time.
[0140] Example 23 includes the subject matter described in Example 22, further comprising: a module for receiving one or more data inputs associated with the one or more images to concatenate the one or more data inputs into a single data input to be processed by the deep learning model, wherein the device includes an autonomous machine, the autonomous machine including one or more of a self-driving vehicle, an autonomous flying vehicle, an autonomous sailing vehicle, and an autonomous home appliance.
[0141] Example 24 includes the subject matter described in Examples 22-23, further comprising: a module for facilitating the deep learning model to receive the single data input to perform one or more deep learning processes including a training process and an inference process to obtain real-time identification of a sensor associated with an unclear image, wherein the sensor includes a camera.
[0142] Example 25 includes the subject matter of Examples 22-24, wherein the deep learning model is further used to: receive multiple data inputs and run the multiple data inputs through the training process and the inference process so that real-time identification of the sensor is accurate and timely.
[0143] Example 26 includes the subject matter of Examples 22-25, wherein the deep learning model comprises one or more neural networks, the one or more neural networks comprising one or more convolutional neural networks, wherein the image is obscured due to one or more of a technical flaw in the sensor or a physical obstruction to the sensor, wherein the physical obstruction is due to a person, plant, animal, or object blocking the sensor, or dirt, smudges, mud, or debris covering a portion of a lens of the sensor.
[0144] Example 27 includes the subject matter of Examples 22-26, further comprising: means for providing one or more of real-time notification of unclear images and real-time automatic correction of the sensor.
[0145] Example 28 includes the subject matter of Examples 22-27, wherein the apparatus comprises one or more processors having a graphics processor co-located on a common semiconductor package with an application processor.
[0146] Example 29 includes at least one non-transitory or tangible machine-readable medium including a plurality of instructions that, when executed on a computing device, are used to implement or perform the method of any of Examples 8-14.
[0147] Example 30 includes at least one machine-readable medium comprising a plurality of instructions that, when executed on a computing device, are used to implement or perform the method of any of Examples 8-14.
[0148] Example 31 includes a system comprising a mechanism for implementing or executing a method as described in any of Examples 8-14.
[0149] Example 32 includes an apparatus comprising a module for performing a method as described in any of Examples 8-14.
[0150] Example 33 includes a computing device arranged to implement or perform a method as described in any of Examples 8-14.
[0151] Example 34 includes a communication device arranged to implement or perform a method as described in any of Examples 8-14.
[0152] Example 35 includes at least one machine-readable medium comprising a plurality of instructions, which, when executed on a computing device, are used to implement or perform a method as described in any of the preceding examples or implement an apparatus as described in any of the preceding examples.
[0153] Example 36 includes at least one non-transitory or tangible machine-readable medium, including a plurality of instructions, which when executed on a computing device are used to implement or perform a method as described in any of the preceding examples or implement an apparatus as described in any of the preceding examples.
[0154] Example 37 includes a system comprising a mechanism for implementing or executing a method as described in any of the preceding examples or implementing an apparatus as described in any of the preceding examples.
[0155] Example 38 includes an apparatus comprising a module for performing a method as described in any of the preceding examples.
[0156] Example 39 includes a computing device arranged to implement or execute a method as described in any of the preceding examples or to implement an apparatus as described in any of the preceding examples.
[0157] Example 40 includes a communication device arranged to implement or execute a method as described in any of the preceding examples or to implement an apparatus as described in any of the preceding examples.
[0158] The accompanying drawings and the foregoing description give examples of embodiments. Those skilled in the art will appreciate that one or more of the elements described may be well combined into a single functional element. Alternatively, certain elements may be separated into multiple functional elements. Elements from one embodiment may be added to another embodiment. For example, the order of the processing described herein may be changed and is not limited to the manner described herein. In addition, the actions of any flowchart do not need to be implemented in the order shown; nor do they necessarily need to perform all actions. In addition, those actions that do not depend on other actions may be performed in parallel with other actions. The scope of the embodiments is by no means limited by these specific examples. Whether or not explicitly given in the specification, many variations are possible, such as differences in structure, size, and material use. The scope of the embodiments is at least as wide as given in the following claims.
Claims
1. An apparatus for facilitating deep learning-based real-time detection and correction of problematic sensors in an autonomous machine, the apparatus comprising: One or more processors for: capturing one or more images of a scene via one or more sensors, wherein an image in the one or more images is determined to be an unclear image due to one or more of a technical flaw associated with the sensor or a physical obstruction of the sensor; receiving one or more data inputs associated with the one or more images to concatenate the one or more data inputs into a single data input to be processed by a deep learning model, wherein the one or more data inputs are collected from the one or more sensors and concatenated into the single data input before being processed by the deep learning model; classifying data associated with the processed single data input to generate a result, the result comprising one or more of identifying the sensor as defective or blocked, notifying the sensor as being identified as defective or blocked, wherein the sensor identified as defective or blocked is further classified as a problematic sensor; Based on identifying or notifying that the sensor is defective or blocked, automatically correcting the defective or blocked sensor associated with the unclear image so that the sensor is automatically corrected to overcome the technical defect or physical blockage, wherein the automatic correction includes: issuing one or more alarms for warning in real time and repairing the problematic sensor in real time; and The single data input is received to perform one or more deep learning processes including a training process and an inference process to obtain real-time identification of sensors associated with the unclear image, wherein the results including the real-time identification of sensors being imperfect or blocked are propagated with weighted updates to predict one or more technical imperfections associated with the one or more sensors.
2. The device according to claim 1, wherein: The deep learning model includes one or more neural networks, including one or more convolutional neural networks, wherein the physical obstruction is attributed to a person, plant, animal, or object blocking the sensor, or dirt, smudges, mud, or debris covering a portion of a lens of the sensor.
3. The device according to claim 1, wherein: The apparatus comprises an autonomous machine comprising one or more of an autonomous driving vehicle, an autonomous flying vehicle, an autonomous sailing vehicle, and an autonomous home appliance, wherein the sensor comprises a camera, wherein the one or more processors comprise one or more of a graphics processor and an application processor, wherein the graphics processor is co-located with the application processor on a common semiconductor package.
4. A method for facilitating deep learning-based real-time detection and correction of problematic sensors in an autonomous machine, the method comprising: capturing one or more images of a scene via one or more sensors associated with a computing device, wherein an image in the one or more images is determined to be an unclear image due to one or more of a technical flaw associated with the sensor or a physical obstruction of the sensor; receiving one or more data inputs associated with the one or more images to concatenate the one or more data inputs into a single data input to be processed by a deep learning model, wherein the one or more data inputs are collected from the one or more sensors and concatenated into the single data input before being processed by the deep learning model; classifying data associated with the processed single data input to generate a result, the result comprising one or more of identifying the sensor as defective or blocked, notifying the sensor as being identified as defective or blocked, wherein the sensor identified as defective or blocked is further classified as a problematic sensor; Based on identifying or notifying that the sensor is defective or blocked, automatically correcting the defective or blocked sensor associated with the unclear image so that the sensor is automatically corrected to overcome the technical defect or physical blockage, wherein the automatic correction includes: issuing one or more alarms for warning in real time and repairing the problematic sensor in real time; and The single data input is received to perform one or more deep learning processes including a training process and an inference process to obtain real-time identification of sensors associated with the unclear image, wherein the results including the real-time identification of sensors being imperfect or blocked are propagated with weighted updates to predict one or more technical imperfections associated with the one or more sensors.
5. The method of claim 4, wherein: The computing device comprises an autonomous machine, the autonomous machine comprising one or more of an autonomous vehicle, an autonomous flying vehicle, an autonomous sailing vehicle, and an autonomous home device, wherein the sensor comprises a camera, wherein the deep learning model comprises one or more neural networks, the one or more neural networks comprising one or more convolutional neural networks, wherein the physical obstruction is due to a person, plant, animal, or object blocking the sensor, or dirt, smudges, mud, or debris covering a portion of a lens of the sensor.
6. The method of claim 4, wherein: The computing device includes one or more processors including one or more of a graphics processor and an application processor, wherein the graphics processor is co-located with the application processor on a common semiconductor package.
7. At least one machine-readable medium comprising a plurality of instructions for implementing or performing the method of any one of claims 4 to 6 when the instructions are executed on a computing device.
8. A system comprising a mechanism for implementing or executing the method according to any one of claims 4 to 6.
9. An apparatus comprising means for executing the method according to any one of claims 4 to 6.
10. A computing device arranged to implement or perform the method according to any one of claims 4 to 6.
11. A communication device, arranged to implement or execute the method according to any one of claims 4 to 6.