Fast adaptation for deep learning applications by back propagation
By combining the calculation of inference configuration gradient and the input configuration gradient, dynamically updating the configuration settings is solved, and the problems of waste of resources and insufficient accuracy in traditional technologies are achieved, and rapid adaptation and efficient resource utilization are achieved in multi-access edge computing environments.
Patent Information
- Application Number
- CN202480005691.6
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Priority Date
- 2023-01-31
- Filing Date
- 2024-01-10
- Publication Date
- 2025-07-25
AI Technical Summary
Traditional configuration management technology has limited effectiveness in timely adjustment of configuration settings, making it difficult to quickly adapt to content changes and lead to waste of computing resources, especially in multi-access edge computing environments that affect the accuracy of inference data.
By combining calculation of inference configuration gradients and input configuration gradients, dynamically update configuration settings, and using the backpropagation technology of deep neural networks, optimize resource allocation to improve the accuracy of inference data.
It realizes rapid adaptation of optimal configuration settings in a multi-access edge computing environment, reducing waste of computing resources, while maintaining or improving the accuracy of inference data.
Smart Images

Figure CN120380481A_ABST
Abstract
Description
BACKGROUND OF THE INVENTION
[0001] With the advent of 5G and multi-access edge computing (MEC), data analysis pipelines have become crucial for improving the accuracy of sensor data captured by sensor devices deployed on-site. In MEC, there is a hierarchy of devices and servers. The hierarchy of devices and servers can jointly form a data analysis pipeline for artificial intelligence (AI) applications to interpret and respond to input data. For example, Internet of Things (IoT) devices (such as cameras in personal or commercial security systems, municipal traffic cameras, dashboard cameras on vehicles, etc.) capture streaming data (such as video data) and send the streaming data to a cellular tower. The cellular tower relays the streaming data as uplink data traffic to an edge server in the local (i.e., "locally deployed") edge. Some IoT devices perform preprocessing on the input data and / or process it as part of the inference of the input data. The local edge server can generate inferences based on the streaming data and send the inference data to a network server at the network edge of the cloud infrastructure. The network server can also generate inference data and transmit the data to a cloud server for further processing. In some aspects, the combination of IoT devices and servers in the MEC hierarchy can form a data analysis pipeline. The data analysis pipeline can include a series of operations at various positions in the pipeline for analyzing data, identifying targets in the data, and making decisions through various stages of generating inference data.
[0002] Accordingly, while traditional configuration management techniques have limited effectiveness in adjusting configuration settings in a timely manner, there is a need to determine and dynamically update configuration settings to minimize the overhead use of computing resources while adapting in a timely manner to sudden changes in content. The aspects disclosed herein have been made based on these and other general considerations. Additionally, although relatively specific problems may be discussed, it should be understood that the examples should not be limited to solving the specific problems identified elsewhere in the background or in this disclosure. SUMMARY OF THE INVENTION
[0003] Aspects of the present disclosure relate to dynamically determining and updating configuration settings associated with capturing input data for inference in a data analysis pipeline under a multi-access edge computing (MEC) system. As described above, MEC involves a hierarchy of servers and data centers with different levels of resource availability and geographical locations. IoT devices (such as cameras, sensing devices, etc.) capture data according to configuration settings. The IoT devices and / or the local edge server use machine learning models (such as deep neural networks) to perform data analysis (such as video analysis of video streaming data, such as inference determination). The present disclosure dynamically updates the configuration settings based on changes or gradients associated with the results of inference determination and changes in the configuration settings.
[0004] The disclosed technology relates to techniques for dynamically determining values of configuration settings associated with content captured by a sensing device and updating configuration setting data to continuously maintain and improve the accuracy of inferring the captured input data. Inferring the captured input data may include using a deep neural network as part of a data processing pipeline. The disclosed technology determines an inference configuration gradient to further determine and update the configuration settings for capturing subsequent data.
[0005] In some aspects, the term "gradient" may refer to the ratio of the difference between values of one parameter type to the difference between values of another parameter type. The term "gradient" may further refer to the derivative of one parameter type with respect to another parameter type. An inference configuration gradient may refer to the derivative of inference data with respect to configuration settings. That is, the inference configuration gradient describes how much a small change in the current value of each parameter type of the configuration settings for the input data will cause the confidence score associated with the inference to change from below (or above) a confidence threshold to above (or below) the confidence threshold. An input inference gradient may refer to the derivative of input data with respect to inference data. The present disclosure uses inference configuration gradient data to determine parameter types and values of the parameter types to update the configuration settings. The IoT device captures subsequent input data according to the updated configuration settings, thereby improving or at least maintaining the accuracy level of the inference input data stream.
[0006] In some aspects, the term "data analysis pipeline" may refer to a series of devices and servers for capturing and analyzing data. For example, streaming data (e.g., video stream data) may be captured by an IoT device (e.g., a camera) and transmitted via a cellular tower to a server tier within a MEC system with resource constraints for handling changes in the streaming data. The term "local edge" may refer to a data center at a remote location at the far edge of a private cloud, which may be close to one or more cellular towers. The RAN combined with the core network of a cloud service provider represents the backbone network for mobile radio telecommunications. For example, a cellular tower may receive and send radio signals to communicate with an IoT device (e.g., a camera) via the RAN (e.g., 5G). Various service applications may perform different functions, such as network monitoring or video streaming, and may be responsible for evaluating data associated with the data traffic. For example, a service application (e.g., an AI application) may perform data analysis on a video stream, such as object recognition (e.g., object counting, face recognition, or human recognition).
[0007] The present invention content is provided to introduce a selection of concepts in a simplified form, which will be further described in the following detailed implementation. The present invention content is not intended to identify the key features or essential features of the claimed subject matter, nor is it intended to limit the scope of the claimed subject matter. Additional aspects, features, and / or advantages of the examples will be partially described in the following description, and partially will be obvious from the description, or can be learned through the practice of the present disclosure. BRIEF DESCRIPTION OF THE DRAWINGS
[0008] The following non-limiting and non-exhaustive examples are described with reference to the accompanying drawings.
[0009] Figure 1 An overview of an example system for reducing streaming data based on a data stream protocol in a data analysis pipeline across MEC tiers according to some aspects of the present disclosure is shown.
[0010] Figure 2 An example system for dynamically updating configuration settings for capturing input data according to some aspects of the present disclosure is shown.
[0011] Figure 3 An example device for dynamically updating configuration settings for capturing input data according to some aspects of the present disclosure is shown.
[0012] Figure 4 An example of calculating a gradient value according to some aspects of the present disclosure is shown.
[0013] Figure 5 An example of data associated with dynamically adapting configuration settings to capture input data according to some aspects of the present disclosure is shown.
[0014] Figure 6 An example of a method for dynamically updating configuration settings to capture input data according to some aspects of the present disclosure is shown.
[0015] Figure 7 A block diagram showing an example of the physical components of a computing device that can practice some aspects of the present disclosure.
[0016] Figure 8 A simplified block diagram of a computing device that can be used to practice some aspects of the present disclosure. DETAILED DESCRIPTION
[0017] Aspects of the present disclosure are described more fully hereinafter with reference to the accompanying drawings, which show specific example aspects from a part of this document. However, the different aspects of the present disclosure may be implemented in many different ways and should not be construed as limited to the aspects set forth herein; rather, these aspects are provided so that this disclosure will be thorough and complete and will fully convey the scope of these aspects to those skilled in the art. The practical aspects may be in the form of a method, system, or apparatus. Accordingly, the aspects may take the form of a hardware implementation, a fully software implementation, or an implementation combining software and hardware aspects. Therefore, the following detailed description should not be taken in a limiting sense.
[0018] The use of artificial intelligence in inference applications has become increasingly popular in automatically analyzing sensor data. Inference applications perform inferences on various data streams captured by sensing devices. For example, some systems use deep neural networks (DNNs) to analyze video content depicting automotive traffic, monitor the speed of vehicles, and identify traffic congestion on streets. Devices are connected to the network as "Internet of Things" (IoT) devices and as sensing devices (e.g., cameras, infrared cameras, optical detection and ranging (e.g., LiDAR) sensors, thermal sensors, etc.).
[0019] IoT devices and associated systems can capture input data and determine the characteristics of that data by processing the data in a data analysis pipeline. Devices and servers in the data analysis pipeline can process the data serially or in parallel to generate inferences (e.g., object recognition) and perform actions based on the inferences. For example, the data analysis performed on video stream data can include identifying regions of interest, identifying the type of object (e.g., face, apple, car, etc.), generating inference data based on the regions of interest and / or object type (e.g., identity of a person), and processing the inference data to determine an action (e.g., making a phone call to a person). Accordingly, the earlier stages of the data analysis pipeline can include processing video frames, while the later stages may not include processing video frames. Some inference processes may occur in IoT devices, while other parts of the inference process can be handled by devices and servers higher in the MEC hierarchy.
[0020] Capturing input data based on configuration settings that match the inference focus improves the accuracy level of the inference data. For example, when inference needs to distinguish objects based on the detailed differences of objects in a video data frame, capturing the video data frame at a high image resolution enables the accuracy level of the video data captured by the inference to be improved. In addition, when the content of the video frame changes substantially, updating the value of the configuration settings in a timely manner becomes important for maintaining and improving the accuracy level of the inference. For example, when the content of the video data captured by a vehicle's dash cam changes significantly as the vehicle exits a tunnel into a brighter location, the brightness setting for capturing image data needs to be adjusted quickly. Some traditional systems can automatically adjust the configuration settings associated with the sensitivity to brightness based on the captured images in response to changes in the sensed brightness.
[0021] In some aspects, traditional systems analyze different configuration settings and switch from using one profile to using another to maintain the accuracy level of the inference. The traditional system will select a profile that includes configuration setting data that produces high inference accuracy without exceeding a predetermined resource budget. However, each profile needs to be determined based on performing additional inferences to analyze multiple configurations (some configurations are expensive in terms of resource usage). Therefore, when the resources available on the graphics processing unit (GPU) are limited, the frequency of traditional system profile updates decreases. When the content of the captured data changes and the optimal configuration has deviated, the decrease in the profile update frequency generally results in using an outdated configuration. Additionally or alternatively, the traditional system can significantly reduce the search space for updating the profile, thereby generating a profile with a suboptimal configuration.
[0022] Some other traditional systems adapt the configuration settings based on the most recent DNN outputs (including intermediate outputs). Although no additional inferences are required to obtain the DNN outputs, this approach tends to adapt slowly to changes in the captured content. Each DNN output reveals limited information about the highly complex relationship between the configuration and the DNN inference output as well as the accuracy level of the inference. The limited information from the DNN output is generally caused by the complexity of the DNN and the preprocessing modules (e.g., processing video codecs) used before executing the DNN. For example, traditional systems for video object detection adjust the video encoding quality by collecting regions of interest from the DNN that may contain objects of interest. The traditional system encodes these regions at a higher encoding quality on subsequent video frames to be captured. Therefore, objects within these regions are more likely to be detected.
[0023] When a traditional system updates configuration settings based on the output of an inference from a DNN, a problem of misalignment of the region of interest occurs. When misalignment occurs, the object detection DNN may fail to detect some objects, so the traditional system needs to further adjust the corresponding region. Since the region generated by the DNN may not be precisely aligned with the region that actually contains the object, the object detection DNN may fail to detect some objects, so the traditional system needs to further adjust the corresponding region.
[0024] In addition, traditional methods that use inference data output from a DNN are typically designed to update specific types of parameters in the configuration settings. For example, a traditional system for video object detection extracts regions that may contain objects and encodes them at a higher encoding quality by updating the values of parameters associated with encoding quality and region selection in the configuration settings. However, such information about videos containing objects may be less useful in other types of parameters in the configuration settings. For example, different types of parameters for the configuration settings are used to discard some frames of a video without degrading the inference accuracy. The process for deciding which frames to discard needs to be based on information about whether the inference results of the discarded frames are similar to the inference results of the retained / kept frames.
[0025] The disclosed techniques address the problem of servicing configuration settings (e.g., knobs) for various parameter types associated with IoT devices. The present disclosure also enables AI applications to quickly adapt and maintain optimal configuration settings when there are a large number of types in the configuration settings without causing a significant increase in the use of computational resources for inference data. The disclosed techniques can adapt the optimal configuration by allocating values to various parameter types of the configuration settings in a way that increases the use of resources associated with configuration settings having a high inference configuration gradient. The present disclosure also reduces the use of resources associated with parameter types of configuration settings having a low inference configuration gradient. In this way, dynamic adaptation improves accuracy by using more resources associated with one configuration setting while reducing the use of resources associated with some other configuration settings without degrading accuracy.
[0026] In some aspects, the disclosed techniques decouple the calculation of the inference configuration gradient (i.e., how a change in the configuration settings changes the output of the inference) into 1) how a change in the configuration settings changes the input data used for inference, and 2) how a change in the input data used for inference changes the inference (i.e., the output of the DNN). The former can be calculated parameter by parameter in the configuration settings without using the DNN to perform inference. The latter can be determined via backpropagation of the DNN using a GPU.
[0027] Figure 1An overview of an example system 100 for identifying techniques for reducing streaming data transmitted across edge tiers under multi-access edge computing in accordance with aspects of the present disclosure is shown. Cellular towers 102A-C communicate wirelessly with IoT devices (e.g., cameras, health monitors, watches, appliances, etc.) via a telecommunications network. Input devices 104A-C represent examples of IoT devices (e.g., cameras) that communicate with the cellular towers 102A-C at a site. In some aspects, input devices 104A-C are capable of capturing video images, processing the video images according to a data stream protocol to reduce the amount of data, and transmitting the reduced data to one or more of the cellular towers 102A-C via a wireless network (e.g., a 5G cellular wireless network). For example, the corresponding input devices 104A-C may capture scenes for video surveillance, such as traffic surveillance or security surveillance. The example system 100 also includes local edges 110A-B (including local edge servers 116A-B), a network edge 130 (including a core network server), and a cloud 150 (including cloud servers responsible for providing cloud services). In some aspects, the example system 100 corresponds to a cloud RAN infrastructure of a mobile wireless telecommunications network.
[0028] In some aspects, input devices 104A-C may filter the captured video data as a preprocessing of the input data stream according to configuration settings. In some other aspects, input devices 104A-C may use a model to identify targets in the captured video data. Input devices 104A-C may include accelerators (e.g., GPUs) to process the captured video stream data. Input devices 104A-C may identify regions of interest and track moving targets (e.g., cars) in the captured video frames. Input devices 104A-C may also determine the types and numbers of targets identified in the video frames. In some aspects, the data describing the types and numbers of targets may be in text format (e.g., in a format using extensible markup language). In some other aspects, the data may include one or more portions of the video images of the captured video frames. One or more portions of the video images may correspond to regions of interest for further processing in a data analysis pipeline. The techniques for processing the video stream data may depend on the computing and memory resources available in the corresponding input devices 104A-C. Input devices 104A-C may transmit the generated inference data as simplified stream data in the format of a data stream protocol. The transmitted data may be further processed in a data analysis pipeline at the local edges 110A-B, the network edge 130, and / or the cloud 150.
[0029] As shown in the figure, the local edge 110A - B is a data center that is part of the cloud RAN. In some aspects, the local edge 110A - B enables cloud integration with the radio access network (RAN). The local edge 110A - B includes local edge servers 116A - B that process incoming and outgoing data traffic. The local edge servers 116A - B can execute service applications 120A - B. In some aspects, the local edge 110A - B is typically geographically distant from the data centers associated with the core network and cloud services. The remote sites are geographically close to the corresponding cellular towers. For example, the proximity can be within a few kilometers. As shown in the figure, local edge 110A is close to cellular tower 102A, and local edge 110B is close to cellular towers 102B - C. In some aspects, inference generators 122A and 122B respectively generate inferences based on applying machine learning models (e.g., deep learning models 124A and 124B (e.g., deep neural networks)) to input data streams (e.g., video streams) captured by IoT devices (e.g., input devices 104A - C). In some aspects, the deep learning models 124A - B represent pre - trained models. The inference generators 122A - B send inference data to an upstream server (e.g., the network edge server 134 of network edge 130).
[0030] In some aspects, configuration updaters 126A - B determine and update the configuration settings associated with the corresponding input devices 104A - C. Example types of configuration settings can include, but are not limited to, image resolution, quantization parameters, video bitrate, frame rate, frame filtering threshold, fine - grained video compression, and fine - grained feature compression. In some aspects, higher values of the configuration settings mean allocating and consuming more computing resources, thus enabling the inference data to reach a higher level of accuracy. For example, increasing the data resolution (e.g., image resolution) consumes more computing resources in terms of memory and processing the image data and further improves the accuracy of the inference input data.
[0031] The configuration updaters 126A - B can determine an inference configuration gradient based on a combination of an input configuration gradient and an inference input gradient. In some aspects, the configuration updaters 126A - B determine the input configuration gradient based on a data stream sequence, and the configuration setting data associated with the corresponding input devices 104A - C is received by the local edge 110A. The input configuration gradient data indicates how changes in the configuration settings alter the content of the input data. Additionally, the configuration updaters 126A - B generate inference input gradient data based on backpropagation as performed by the deep learning models 124A - B.
[0032] In some aspects, the local edge 110A can aggregate the data it receives from the input devices 104A-B and generate inference data based on the aggregated data from the respective input devices 104A-B for transmission to the network edge 130. For example, the input devices 104A-B can capture videos of the same location from different perspectives (e.g., street scenes from two opposite directions). The local edge 110A can aggregate the two data streams to generate inference data.
[0033] In other aspects, as data centers get closer to the cloud 150, server resources (including processing units and memory) become more robust and powerful. As an example, the server 154 can be more powerful than the network edge server 134, and the network edge server 134 can be more powerful than the local edge servers 116A-B.
[0034] In some aspects, the network edge 130 is at a regional data center of a private cloud service. For example, the regional data center can be located several tens of kilometers away from the cellular towers 102A-C. The network edge 130 includes a service application 140 that performs data analysis when executed. For example, the service application 140 includes a video ML model or an inference generator 142, and the service application 140 uses machine learning techniques (such as neural networks) to perform and manage video analysis to train the analysis model. The network edge 130 can include a wider range of memory resources than the memory resources available to the local edge servers 116A-B of the local edge 110A-B.
[0035] The cloud 150 (services) includes cloud servers for performing resource-intensive non-real-time service operations. In some aspects, one or more servers in the cloud 150 can be located at a central position in the cloud RAN infrastructure. In this case, the central position can be located several hundred kilometers away from the cellular towers 102A-C. In some aspects, the cloud 150 includes a service application 160 for performing data analysis. The service application 160 can perform processing tasks similar to those of the service application 140 in the network edge 130.
[0036] In some aspects, the local edge 110A-B, which is closer to the cellular towers 102A-C and the input devices 104A-C (or IoT devices) than the cloud 150, can provide real-time processing. In contrast, the cloud 150, which is the farthest from the cellular towers 102A-C and the input devices 104A-C in the cloud RAN infrastructure, can provide processing in a non-real-time manner.
[0037] The service applications 120A-B include program instructions for processing data according to a predetermined data analysis scenario on the local edge servers 116A-B. The predetermined analysis can include, for example, an inference generator 122A-B for generating inferences based on the captured data. In some aspects, the inference generator 122A-B performs video analysis and generates inference data by extracting and identifying objects from video stream data according to a trained model. For example, the inference generator 122A-B can rely on multiple trained models to identify different types of objects (e.g., trees, animals, people, cars, etc.), generate a count of the objects (e.g., the number of people in a video frame), and / or identify a specific object (e.g., a specific person based on facial recognition). In some aspects, each model can be trained to identify different types of objects.
[0038] The incoming video stream can include background data and object data captured by IoT devices (e.g., the input devices 104A-C) and transmitted to the cellular towers 102A-C. For example, the service applications 120A-B can analyze the video stream and extract a portion of the video stream as a region of interest, which can include object data rather than background data. Once extracted, the region of interest can be evaluated to identify an object (e.g., a person's face), as described above, or the service applications 120A-B can send the extracted region of interest (instead of the full video stream) to the cloud for further processing (e.g., identifying a person by performing facial recognition on the person's face). In some aspects, the edge servers 116 include limited computing and memory resources, while the network edge servers 134 at the network edge 130 include robust enough resources to perform facial recognition on the video stream to identify a person's name.
[0039] In some other aspects, the corresponding input devices 104A-C can be integrated with the service applications 120A-B. Thus, the corresponding input devices 104A-C can include an inference generator 122A-B, a deep learning model 124A-B, and a configuration updater 126A-B. For example, the input device 104A can include a central processing unit (CPU) to process the capture of input data (e.g., video data) and the preprocessing of the input data for the deep learning model 124A. The input device 104A can also include a graphics processing unit (GPU) to perform the inference generator 122A using the deep learning model 124A. (See also Figure 3 )
[0040] It should be understood that with respect to Figure 1The various methods, devices, applications, features, etc. described are not intended to limit system 100 implemented by the specific applications and features described. Thus, additional controller configurations may be used to practice the methods and systems herein and / or the described features and applications may be excluded without departing from the methods and systems disclosed herein.
[0041] Figure 2 FIG. shows an example system for dynamically updating configuration settings for capturing input data in accordance with some aspects of the present disclosure. System 200 includes a sensor 202, a DNN-based data analyzer 204, and an inference data transmitter 206. In some other aspects, system 200 corresponds to a combination of input device 104A and local edge 110A. Sensor 202 (e.g., input device 104A as Figure 1 shown) captures images and sends input stream data (e.g., sensor stream data, video stream data, etc.) to the DNN-based data analyzer 204 (e.g., via local edge 110A of a cellular tower (e.g., cellular tower 102A as Figure 1 shown)). In some other aspects, when input device 104A includes service application 120A, system 200 corresponds to input device 104A.
[0042] Sensor 202 includes one or more types of sensors. For example, range detector 210 detects range or distance through optical detection and / or ranging. Range detector 210 includes configuration settings 240 for setting various parameter values for detecting range. Camera 212 captures video stream data as input using configuration settings 242. Image sensor 214 captures image data as input using configuration settings 244. Sensor 202 includes a configuration settings updater 216. Configuration settings updater 216 updates the configuration settings associated with the respective sensors (e.g., configuration settings 240 associated with range detector 210, configuration settings 242 associated with camera 212, configuration settings 244 associated with image sensor 214, etc.).
[0043] The DNN-based data analyzer 204 generates inference data based on input data using a trained deep neural network (DNN). In some aspects, CPU 250 performs input processing 220 and gradient processing 224. GPU 252 performs inference 222 using a trained deep neural network (e.g., deep learning model 124A as Figure 1 shown). In some aspects, CPU 250 and GPU 252 are different processors. Input processing 220 receives the captured data stream from sensor 202 and generates input data for inference 222 (deep neural network). In some aspects, the input data includes a multi-dimensional vector representing image data in pixels. Each dimension of the multi-dimensional vector corresponds to a pixel of the image data.
[0044] Gradient processing 224 generates a pair of input configuration gradients and inference input gradients based on a combination of the streaming data, configuration settings associated with sensor 202 that captures the data stream as input, and the result of inference 222 based on the streaming data. Based on the pair of input configuration gradients and inference input gradients, gradient processing 224 generates an inference configuration gradient by multiplying the values of the pair.
[0045] The inference configuration gradient indicates how a change in the configuration settings changes the inference data. The input configuration gradient indicates how a change in the configuration settings changes the input data. The inference-input gradient indicates how a change in the input data changes the inference. Given the inference configuration gradient, gradient processing 224 determines values to update one or more types of configuration settings associated with sensor 202 for capturing subsequent inputs in the streaming data. Configuration settings updater 216 updates the corresponding configuration settings of sensor 202.
[0046] In some aspects, the value of the inference can be based on a confidence value associated with the inference data. In some aspects, a change in the inference can be based on the difference between the confidence value and a predetermined threshold, thereby eliminating the need to process the inference multiple times and reducing GPU resource consumption.
[0047] It should be understood that the various methods, devices, applications, features, etc. described with respect to Figure 2 are not intended to limit system 200 implemented by the specific applications and features described. Thus, additional controller configurations can be used to practice the methods and systems described herein, and / or the described features and applications can be excluded without departing from the methods and systems disclosed herein.
[0048] Figure 3 An example device for dynamically updating configuration settings for capturing input data in accordance with aspects of the present disclosure is shown. Figure 3 An example of input device 302 is shown. Input device 302 includes input processing 304 (of the CPU), inference 306 (GPU), and inference data transmitter 314. In some aspects, input processing 304 includes input data receiver 310 (or data capture device), inference configuration gradient generator 320, and configuration updater 322. Inference 306 (GPU) includes inference generator 312 (deep neural network) and inference input gradient generator 316 (backpropagation).
[0049] The input data receiver 310 captures the input data stream 340. Examples of the input data stream 340 include a video data stream, a range (e.g., distance) data stream, and the like. The input configuration gradient generator 318 generates input configuration gradient data 346 based on the input data stream 340 and the configuration setting data 350 used to capture the input data stream 340. In some aspects, the configuration setting data 350 changes over time when the configuration updater 322 updates the configuration setting data 350.
[0050] The inference generator 312 (deep neural network) generates inference data 342 based on the input data stream 340. In some aspects, the inference data 342 may include confidence values. The higher the confidence value, the more likely the inference data 342 is to be inferred more accurately from the input data stream 340. The inference input gradient generator 316 (backpropagation) generates inference input gradient data 344.
[0051] The inference configuration gradient generator 320 receives the input configuration gradient data 346 and the inference input gradient data 344. By multiplying the input configuration gradient data 346 by the inference input gradient data 344, the inference configuration gradient generator 320 generates inference configuration gradient data 348.
[0052] The configuration updater 322 updates the value associated with the type of configuration setting for capturing subsequent data for the input data stream 340 based on the inference configuration gradient data 348.
[0053] Figure 4 An example of calculating a gradient value is shown in accordance with some aspects of the present disclosure. Hereinafter, the calculation 400 will be explained with reference to the systems, components, devices, modules, software, data structures, data characteristic representations, signaling diagrams, methods, etc. described in conjunction with Figure 1 、 Figure 2 、 Figure 3 、 Figure 5 、 Figure 6 、 Figure 7 、and Figure 8 The calculation 400 involves calculating the change in the inference of the configuration setting as the inference configuration gradient. The inference configuration gradient is generated based on the following equation:
[0054] In some aspects, the calculation 400 involves calculating the change in the inference of the configuration setting as the inference configuration gradient. The inference configuration gradient is generated based on the following equation: Inference configuration gradient = Inference input gradient x Input configuration gradient (1)
[0055] The terms for calculating the gradient include partial derivatives.
[0056] The inference configuration gradient indicates the degree of change in the inference result based on the degree of change in the value(s) of the configuration setting(s). The input configuration gradient indicates the degree of change in the input data based on the degree of change in the value(s) of the configuration setting(s). For example, the change in the input data can be described by the difference in pixel values between two frames of video data in an input data stream.
[0057] The inference input gradient indicates the degree of change in the inference result based on the change in the input data. In some aspects, the inference input gradient is equivalent to the result of backpropagation (i.e., saliency) of a deep neural network that performs inference on the input data. The backpropagation can be executed by a GPU reserved for processing inferences. In contrast, the processing related to capturing and preprocessing the input data can be executed by a CPU. Thus, determining the input configuration gradient and the inference configuration gradient consumes CPU resources, while determining the inference input gradient based on backpropagation consumes GPU resources. The processing of backpropagation incurs a substantially constant computational overhead on the GPU resources without the need to perform multiple inferences on the same input data. The GPU performs a combination of inference and backpropagation on the input data using the deep neural network.
[0058] Figure 5 An example of data associated with dynamically adapting configuration settings to capture input data is shown. In the following, example data 500 will be explained with reference to the systems, components, devices, modules, software, data structures, data characteristic representations, signaling diagrams, methods, etc. described in Figure 1 、 Figure 2 、 Figure 3 、 Figure 4 、 Figure 6 、 Figure 7 、and Figure 8 The example data 500 includes configuration setting data 502, input data #1 504a, input data #2 504b, inference data #1 506a, inference data #2 506b, input difference 530, configuration setting data difference 532, inference data difference 534, inference configuration gradient 536, and update data 538.
[0059]
[0060] In some aspects, the configuration setting data 502 includes an image resolution 510 and a frame rate 516. The image resolution 510 includes seven different values, each value associated with a different configuration setting on the image resolution. For example, the setting value 512 one (1) corresponds to a pixel size 514 of 172 pixels horizontally and 120 pixels vertically. Value two corresponds to 352x240 pixels, value three corresponds to 720x480 pixels, value four corresponds to 1280x1024 pixels, value five corresponds to 1920x1080 pixels, value six corresponds to 3840x2160 pixels, and value seven corresponds to 4000x3000 pixels. In some aspects, the higher the setting value, the higher the image resolution, resulting in a higher accuracy of the inference input image data. A camera (e.g., the input device 104A (IoT device) as shown in Figure 1 captures image data frames as video data according to the image resolution 510 specified in the configuration settings.
[0061] The frame rate 516 indicates the values of the setting value 518 and the number of frames per second 520 used to capture video data. For example, the setting value 518 one (1) corresponds to one frame per second. Two corresponds to two frames per second. Three corresponds to six frames per second. Four corresponds to 12 frames per second. Five corresponds to 30 frames per second. Six corresponds to 60 frames per second. Seven corresponds to 120 frames per second, etc. A camera (e.g., the input device 104A (IoT device) as shown in Figure 1 captures image data frames as video data according to the frame rate 516 specified in the configuration settings.
[0062] In some aspects, the input data #1 504a represents the first image data as input using the value seven (i.e., 4000x3000 pixels) of the image resolution specified by the configuration setting data 502. Similarly, the input data #2 504b represents the second image data as input using the image resolution value three (i.e., 720x480). According to this example, the image resolution of the input image data is reduced by four.
[0063] Given the input data #1 504a, the disclosed technique uses a deep neural network (e.g., the deep learning model 124A as shown in Figure 1 the inference 222 (deep neural network) as shown in Figure 2 and the Figure 3The inference generator 312 (deep neural network) shown is used to perform inference and generate inference data #1 506a. In some aspects, the inference data #1 506a indicates a confidence level of +8 that exceeds a predetermined threshold. The confidence level indicates the confidence level at which the inference data #1 506a accurately infers the input data #1 504a. Similarly, the inference data #2 506b indicates a confidence level of -4 that exceeds a predetermined threshold. Therefore, the confidence level associated with the inference input data is reduced by 12 on two examples of the input data. In an example, the reduction in the confidence value may be caused by a reduction in the image resolution on the input video data.
[0064] The degree of change in pixels between the input data #1 504a and the input data #2 504b is six. The difference in the configuration setting data (i.e., image resolution) is 4 (i.e., 7 minus 4). The difference in the inference data 534 is 12 (8 - (-4) = 12). Therefore, the inference configuration gradient 536 can be determined by multiplying the inference input gradient (12 / 6) by the input configuration gradient (6 / 4). The resulting value of the inference configuration gradient 536 is 3.
[0065] Based on the value of the inference configuration gradient 536 being 3, the configuration setting of the image resolution is increased by 3. Accordingly, the configuration setting of the image resolution changes from 3 (720x480) to 6 (i.e., 3840x2160).
[0066] Figure 6 is an example of a method for dynamically updating configuration settings to capture input data according to some aspects of the present disclosure. Figure 6 The general order of operations of the method 600 is shown. Generally, the method 600 begins with the start operation 602 and proceeds to the update configuration setting data operation 606 after the determine configuration setting data operation 618. The method 600 may include more or fewer steps, or the order of arrangement of the steps may be different from Figure 6 that shown. The method 600 can be executed as a set of computer-executable instructions executed by a computer system and encoded or stored on a computer-readable medium. Additionally, the method 600 can be executed by gates or circuits associated with a processor, ASIC, FPGA, SOC, or other hardware device. Hereinafter, reference will be made to Figure 1 、 Figure 2 、 Figure 3 、 Figure 4 、 Figure 5 、 Figure 7 、and Figure 8 described systems, components, devices, modules, software, data structures, data characteristic representations, signaling diagrams, methods, etc. to explain the method 600.
[0067] After starting operation 602, method 600 begins with capturing initial input data operation 604, which captures initial input data using configuration setting data. In some aspects, the configuration setting data used for capturing initial input data can be predetermined. The initial input data can be a frame of image data, video data, sensor data, etc. The capturing initial input data operation 604 can be performed by an input device (IoT device; such as Figure 1 the input devices 104A-C shown). For example, the initial input data can represent image data at a predetermined image resolution (e.g., level 7 of 4000x3000 pixels as shown in Figure 5 ).
[0068] The updating configuration setting data operation 606 updates the configuration setting data. In some aspects, the configuration settings configure the conditions for capturing input data. The types of configuration settings include one or more of the following: image resolution, frame rate, quantization parameter, video bit rate, frame filtering threshold, fine-grained video compression, fine-grained feature compression, etc. The configuration settings can include preprocessing the input data to generate an input for inference. In some aspects, the configuration setting data can be based on a change in the inference data associated with the previously captured input data. Additionally or alternatively, the configuration setting data can be predetermined.
[0069] The capturing input data operation 608 captures input data based on the updated configuration setting data of the input device. In some aspects, the input data can be image data captured at a reduced image resolution based on the updated configuration setting data. (For example, level 3 of 720x480 pixels as shown in Figure 5 ).
[0070] The generating input configuration gradient data operation 610 generates input configuration gradient data. In some aspects, the input configuration gradient data indicates the ratio of the change in the input data to the change in the configuration setting data. For example, when the degree of change in the pixel content of the input image data due to a reduction in image resolution is 6 and the change in image resolution is 4, the value of the input configuration gradient data is 1.5 (i.e., 6 / 4). In some aspects, the change in the pixel content of two input image data can be expressed by comparing the pixel values in the two-dimensional coordinates associated with the two input image data.
[0071] In some aspects, the configuration settings can include multiple types (e.g., image resolution, frame rate, bit rate, etc.). Determining the input configuration gradient data for each type of configuration setting can increase the encoding overhead proportional to the number of types. To reduce the overhead, the disclosed techniques identify those types of configuration settings that primarily trigger input changes on different non-overlapping portions of the input data. For example, some types of configuration settings affect a specific portion of the input data (e.g., the left half of the image data), while other types of configuration settings affect different portions of the input data. In some aspects, the generating input configuration gradient data operation 610 divides the input data (e.g., a frame of image data) into different portions corresponding to the respective types of configuration settings and generates the input configuration gradient data associated with the respective portions of the input data in parallel.
[0072] The generating inference data operation 612 generates inference data associated with the input data. In some aspects, a deep neural network can be used to perform the inference. The deep neural network uses an encoded multi-dimensional vector of the captured data as input and predicts the inference data as output. For example, the deep neural network can receive an encoded multi-dimensional vector representing a frame of video data and predict the target type (e.g., vehicle) and the number of targets (e.g., the number of vehicles) in the frame of the video data as the inference data. In some aspects, the deep neural network generates a confidence value associated with the inference data. The higher the confidence value, the higher the accuracy of the inference input data. In some aspects, the generating inference data operation 612 can be performed by a graphics processing unit (GPU) and / or a process different from the CPU and dedicated to processing inference operations.
[0073] In an example, the inference data for the initial input data (or previous input data) indicates a confidence level 8 scores higher than a predetermined threshold. In contrast, the confidence level of the inference data for the current input data is 4 scores lower than the predetermined threshold. The 12-point decrease in the confidence level of the inference may be caused by a decrease in the image resolution used to capture the input data. In some other aspects, a change in the content of the input data (e.g., a change in the scene, including a change in the content of a frame of video data captured by a dashcam of an IoT device as a vehicle exits from inside a tunnel to outside the tunnel) can cause a change in the configuration settings of the dashcam in response to the change in content.
[0074] The generating inference input gradient data operation 614 generates inference input gradient data (e.g., saliency) based on the inference input data. In some aspects, the backpropagation of the deep neural network generates the inference input gradient data. The inference-input gradient data describes how a change in the input data changes the result of the inference. In an example, when the change in the inference data is 12 and the change in the input data is 6, the inference input gradient data becomes 2 (i.e., 12 / 6).
[0075] Performing backpropagation operations can result in high GPU computational overhead and high network bandwidth requirements for streaming saliency. To reduce the propagation cost, the operation of generating inference input gradient data 614 can be performed periodically according to a predetermined period, and the generated data can be reused until the next generation of inference input gradient data. The periodic backpropagation operation can consume GPU resources in a predictable manner according to a periodic schedule. For example, the predetermined period can be ten frames of a video data stream. The reuse of saliency and the periodic generation of inference input gradient data reduce GPU overhead without sacrificing the accuracy of inference.
[0076] In some aspects, the difference can be represented by an absolute value, based on the assumption that the configuration setting has a "direction" such that as the resource consumption level increases, the accuracy level also increases (e.g., a higher image resolution of image data results in a higher accuracy level for inferring the image data), and changes along this "direction" only increase or maintain the accuracy level. The value i corresponds to one of the types of configuration settings among multiple types in the configuration setting:
[0077] The operation of generating inference configuration gradient data 616 generates inference configuration gradient data. In some aspects, the inference configuration gradient data indicates how changes in the configuration setting data change the inference data. In some aspects, the inference configuration gradient data is based on the product of the inference input gradient data and the input configuration gradient data.
[0078] The operation of determining configuration setting data 618 determines the configuration setting data based on the inference configuration gradient data. In some aspects, the inference configuration gradient data indicates the degree of change for updating the configuration setting data (e.g., an increase in computational resource consumption, which also indicates the degree of improvement in the accuracy of the inference input data). For example, inference configuration gradient data having a value of 3 on an image resolution scale indicates that the configuration setting data associated with the image resolution needs to be increased by 3. Accordingly, the set value of 3 (i.e., 720x480) is increased to 6 (i.e., 3840x2160, as Figure 5 shown).
[0079] In some aspects, the configuration settings can include multiple types and are not limited to image resolution and / or frame rate. The configuration setting data can be multi-dimensional. Accordingly, the inference configuration gradient data can be multi-dimensional, representing the partial derivatives of the multi-dimensional data. By using multi-dimensional gradient data, the disclosed techniques identify the types of configuration setting data to increase the values of specific types of configuration settings to improve accuracy and reduce the values of other specific types of configuration settings without reducing accuracy.
[0080] In some aspects, when a sensor (e.g., input device 104A as shown in Figure 1 ) and a deep neural network (e.g., inference generator 122A with deep learning model 124A as shown in Figure 1 ) are connected via a network, bandwidth consumption can become an issue. The raw sensor data (i.e., the raw input data stream) may need to be sent over the network to obtain the changes in the input data for the deep neural network before and after changing the configuration settings. To reduce communication overhead, the inference configuration gradient (i.e., backpropagation data or saliency) can be sent to the sensor so that the sensor can determine the input configuration gradient data within the sensor.
[0081] After determining the configure settings data operation 618, the steps of method 600 proceed to the update configure settings data operation 606. Method 600 also proceeds to the capture input data operation 608 using the updated configure settings data.
[0082] It should be understood that operations 602 - 618 are described for purposes of illustrating the method and system and are not intended to limit the present disclosure to a particular sequence of steps. For example, the steps may be performed in a different order without departing from the present disclosure, additional steps may be performed, and the disclosed steps may be excluded.
[0083] Figure 7 is a block diagram showing the physical components (e.g., hardware) of a computing device 700 that can practice aspects of the present disclosure. The computing device components described below can be applicable to the computing device described above. In a basic configuration, computing device 700 can include at least one processing unit 702 and a system memory 704. Depending on the configuration and type of the computing device, system memory 704 can include, but is not limited to, volatile storage (e.g., random access memory), non-volatile storage (e.g., read-only memory), flash memory, or any combination of these memories. System memory 704 can include an operating system 705 and one or more program tools 706 that are adapted to perform various aspects disclosed herein. The operating system 705 can be suitable for controlling the operation of computing device 700, for example. Additionally, aspects of the present disclosure can be practiced in conjunction with a graphics library, other operating systems, or any other application programs and are not limited to any particular application or system. This basic configuration is shown by those components within the dashed line 708 in Figure 7 . Computing device 700 can have additional features or functionality. For example, computing device 700 can also include additional (removable and / or non-removable) data storage devices, such as magnetic disks, optical disks, or magnetic tapes. Such additional storage is shown by removable storage device 709 and non-removable storage device 710 in Figure 7 .
[0084] As described above, multiple program tools and data files can be stored in the system memory 704. When executed on at least one processing unit 702, the program tool 706 (e.g., the application 720) can perform processes including but not limited to the aspects described herein. As described with respect to Figure 2 in more detail, the application 720 includes a model receiver 722, a model updater 724, a data receiver 726, an inference data generator 728, and a data transmitter 730. Other program tools that can be used in accordance with aspects of the present disclosure can include email and contact applications, word processing applications, spreadsheet applications, database applications, slide presentation applications, drawing or computer-aided applications, and the like.
[0085] In addition, aspects of the present disclosure can be practiced in a circuit including discrete electronic components, a packaged or integrated electronic chip containing logic gates, a circuit utilizing a microprocessor, or a single chip containing electronic components or a microprocessor. For example, aspects of the present disclosure can be practiced via a system-on-chip (SOC), where Figure 7 each or many of the components shown therein can be integrated onto a single integrated circuit. Such an SOC device can include one or more processing units, graphics units, communication units, system virtualization units, and various application functions, all of which are integrated (or "burned") onto a chip substrate as a single integrated circuit. When operating via an SOC, the functions described herein regarding the capabilities of the client switching protocol can be run via application-specific logic that is integrated with other components of the computing device 700 on a single integrated circuit (chip). Aspects of the present disclosure can also be practiced using other technologies capable of performing logical operations, such as, for example, AND, OR, and NOT, including but not limited to mechanical, optical, fluidic, and quantum technologies. Additionally, aspects of the present disclosure can be practiced within a general-purpose computer or in any other circuit or system.
[0086] The computing device 700 can also have one or more input devices 712, such as a keyboard, mouse, pen, voice or speech input device, touch or swipe input device, and the like. One or more output devices 714, such as a display, speaker, printer, and the like, can also be included. The above devices are examples, and other devices can be used. The computing device 700 can include one or more communication connections 716 that allow communication with other computing devices 750. Examples of communication connections 716 include but are not limited to radio frequency (RF) transmitter, receiver, and / or transceiver circuitry; universal serial bus (USB), parallel, and / or serial ports.
[0087] As used herein, the term computer-readable medium may include computer storage media. Computer storage media may include volatile and non-volatile, removable and non-removable media implemented in any method or technology for storage of information such as computer-readable instructions, data structures, or program modules. System memory 704, removable storage device 709, and non-removable storage device 710 are all examples of computer storage media (e.g., memory storage). Computer storage media may include RAM, ROM, electrically erasable programmable read-only memory (EEPROM), flash memory or other memory technology, CD-ROM, digital versatile disks (DVD) or other optical storage, magnetic cassettes, magnetic tape, magnetic disk storage or other magnetic storage devices, or any other article of manufacture that can be used to store the information and that can be accessed by computing device 700. Any such computer storage media may be part of computing device 700. Computer storage media does not include carrier waves or other propagated or modulated data signals.
[0088] Communication media may be embodied by computer-readable instructions, data structures, program modules, or other data in a modulated data signal such as a carrier wave or other transport mechanism and includes any information delivery media. The term "modulated data signal" may describe a signal that has one or more of its characteristics set or changed in such a manner as to encode information in the signal. By way of example, and not limitation, communication media may include wired media such as a wired network or direct-wired connection, and wireless media such as acoustic, radio frequency (RF), infrared and other wireless media.
[0089] Figure 8 A computing device 800 is shown that may utilize aspects of the present disclosure, e.g., a mobile phone, a smartphone, a wearable computer (such as a smartwatch), a tablet computer, a laptop computer, etc. In some aspects, a client utilized by a user (e.g., as an Figure 1 operator of a server in local edge 110A-B in Figure 8, one aspect of a computing device 800 for implementing these aspects is shown. In a basic configuration, the computing device 800 is a handheld computer with both input elements and output elements. The computing device 800 typically includes a display 805 and one or more input buttons 810 that allow a user to input information into the computing device 800. The display 805 of the computing device 800 can also be used as an input device (e.g., a touch screen display). If included as an optional input element, the side input element 815 allows further user input. The side input element 815 can be a rotary switch, a button, or any other type of manual input element. In alternative aspects, the computing device 800 can include more or less input elements. For example, in some aspects, the display 805 may not be a touch screen. In yet another alternative aspect, the computing device 800 is a portable telephone system, such as a cellular phone. The computing device 800 can also include an optional keyboard 835. The optional keyboard 835 can be a physical keyboard or a "soft" keyboard generated on a touch screen display. In various aspects, output elements include a display 805 for showing a graphical user interface (GUI), a visual indicator 820 (e.g., a light emitting diode), and / or an audio transducer 825 (e.g., a speaker). In some aspects, the computing device 800 includes a vibration transducer for providing tactile feedback to the user. In yet another aspect, the computing device 800 includes input and / or output ports for sending or receiving signals to or from external devices, such as an audio input (e.g., a microphone jack), an audio output (e.g., a headphone jack), and a video output (e.g., an HDMI port).
[0090] Figure 8 Further shown are computing devices, servers (e.g., Figure 1 The system 802 may be implemented as a "smart phone" capable of running one or more applications (e.g., browsers, email, calendars, contact managers, messaging clients, games, and media clients / players). In some aspects, the system 802 is integrated as a computing device, such as an integrated digital assistant (PDA) and a wireless phone.
[0091] One or more applications 866 may be loaded into the memory 862 and run on or associated with the operating system 864. Examples of applications include a telephone dialer, an email program, a personal information management (PIM) program, a word processing program, a spreadsheet program, an Internet browser program, a messaging program, and the like. The system 802 also includes a non-volatile storage area 868 within the memory 862. The non-volatile storage area 868 may be used to store persistent information that should not be lost even if the system 802 is powered off. The applications 866 may use and store information in the non-volatile storage area 868, such as emails, or other messages used by the email application. A synchronization application (not shown) also resides on the system 802 and is programmed to interact with a corresponding synchronization application residing on a host computer to keep the information stored in the non-volatile storage area 868 synchronized with the corresponding information stored at the host computer. It should be understood that other applications may be loaded into the memory 862 and run on the computing device 800 described herein.
[0092] The system 802 has a power source 870, which may be implemented as one or more batteries. The power source 870 may also include an external power source, such as an AC adapter or a charging docking station that supplements or recharges the battery.
[0093] The system 802 may also include a wireless interface layer 872 that performs the functions of sending and receiving radio frequency communications. The radio interface layer 872 supports a wireless connection between the system 802 and the "outside world" via a communication carrier or service provider. Transmissions to and from the radio interface layer 872 are under the control of the operating system 864. In other words, communications received by the radio interface layer 872 may be propagated to the applications 866 via the operating system 864, and vice versa.
[0094] A visual indicator 820 (e.g., an LED) can be used to provide visual notifications, and / or an audio interface 874 can be used to generate audible notifications via an audio transducer 825. In the illustrated configuration, the visual indicator 820 is a light-emitting diode (LED), and the audio transducer 825 is a speaker. These devices can be directly coupled to a power source 870 such that when activated, they remain on for a duration indicated by the notification mechanism even if the processor 860 and other components may be turned off to conserve battery power. The LED can be programmed to remain on indefinitely until the user takes an action to indicate the powered-on state of the device. The audio interface 874 is used to provide audible signals to the user and receive audible signals from the user. For example, in addition to being coupled to the audio transducer 825, the audio interface 874 can also be coupled to a microphone to receive audible input, such as to support a phone conversation. According to aspects of the present disclosure, as will be described below, the microphone can also be used as an audio sensor to support control of notifications. The system 802 can also include a video interface 876 that enables operation of the on-board camera 830 to record still images, video streams, etc.
[0095] The computing device 800 implementing the system 802 can have additional features or functionality. For example, the computing device 800 can also include additional data storage devices (removable and / or non-removable), such as magnetic disks, optical disks, or magnetic tapes. Such additional storage is Figure 8 illustrated by the non-volatile storage area 868 in B.
[0096] As described above, data / information generated or captured by the computing device 800 and stored via the system 802 can be stored locally on the computing device 800. Alternatively, the data can be stored on any number of storage media that can be accessed by the device via the radio interface layer 872 or via a wired connection between the computing device 800 and a separate computing device associated with the computing device 800 (e.g., a server computer in a distributed computing network such as the Internet). It should be understood that such data / information can be accessed via the radio interface layer 872 or via the distributed computing network via the computing device 800. Similarly, such data / information can be easily transmitted between computing devices for storage and use according to well-known data / information transmission and storage means, including email and collaborative data / information sharing systems.
[0097] The description and illustration of one or more aspects provided in this application are not intended to limit or restrict the scope of the present disclosure in any way. The claimed disclosure should not be construed as limited to, for example, any aspect or detail provided in this application. Various features (structures and methods), whether shown and described in combination or separately, are intended to be selectively included or omitted to produce embodiments having a particular set of features. Having provided the description and illustration of this application, those skilled in the art can envision variations, modifications, and alternative aspects that fall within the spirit of the broader aspects of the general inventive concept embodied in this application, and these variations, modifications, and alternative aspects do not depart from the broader scope of the claimed disclosure.
[0098] The present disclosure relates to determining and updating configuration settings for capturing input data in multi-access edge computing at least according to the examples provided in the following sections. A method for updating configuration settings is associated with capturing content using an Internet of Things (IoT) device, which is one of a plurality of IoT devices that are at least a part of an edge computing system. The method includes a first processor. The method includes: receiving input data based on a configuration value associated with the configuration settings of the IoT device; determining, by the first processor, a first gradient value, wherein the first gradient value represents a change in the input data relative to previously received input data based on a change in the configuration settings for capturing content using the IoT device; causing an edge server of the edge computing system to determine, by a second processor using a neural network, a second gradient value, wherein the second gradient value indicates a change in inference data based on the change in the input data, wherein the first processor and the second processor are different; updating, at least based on a combination of the first gradient value and the second gradient value, the configuration value associated with the configuration settings for adjusting the input operations of the IoT device, wherein the combination of the first gradient value and the second gradient value represents an expected change in the inference data relative to the change in the configuration value associated with the configuration settings of the IoT device, and wherein the updating results in improved inference of the input data by adjusting the configuration settings of the IoT device; and receiving subsequent input data using the IoT device based on the updated configuration value. The second processor includes a graphics processing unit associated with a data analysis pipeline for multi-access edge computing in a 5G telecommunications network, and wherein the second processor is different from the first processor. The first gradient value includes an input configuration gradient, wherein the input configuration gradient indicates the degree of change in the input data based on the updated configuration value. The second gradient value includes an inference input gradient, wherein the inference input gradient indicates the degree of change in a confidence score associated with the inference input data when the input data changes, and wherein the inference input gradient is based on the saliency associated with the neural network. The combination of the first gradient value and the second gradient value represents an inference configuration gradient, wherein the inference configuration gradient indicates the degree of change in the confidence score associated with inferring the input data based on the updated configuration value. The configuration value is associated with an image resolution for capturing content by the IoT device. The method further includes: receiving subsequent input data based on the updated configuration value associated with the configuration settings for operating the IoT device; and updating, at least based on a combination of a subsequent change in the configuration value including the configuration value and a subsequent change in inferring the subsequent input data, the configuration value of the configuration settings for further operating the IoT device.
[0099] Another aspect of the present technology relates to a system for capturing content using an IoT device among multiple Internet of Things (IoT) devices in an edge computing system. The IoT device includes a first processor, and the system includes: a memory; and a first processor configured to execute a method that includes: receiving input data based on a configuration value associated with a configuration setting of the IoT device for capturing content; determining, by the first processor, a first gradient value, where the first gradient value represents a change in the input data relative to previously received input data based on a change in the configuration setting for using the IoT device to capture content; causing an edge server of the edge computing system to determine, by a second processor using a neural network, a second gradient value, where the second gradient value indicates a change in inference data relative to previously generated inference data based on a change in the input data, where the first processor and the second processor are different; updating, at least based on a combination of the first gradient value and the second gradient value, a configuration value associated with a configuration setting for adjusting an input operation of the IoT device, where the combination of the first gradient value and the second gradient value represents an expected change in the inference data relative to a change in the configuration value associated with the configuration setting of the IoT device, and where the update results in improved inference of the input data due to adjusting the configuration setting of the IoT device; and receiving subsequent input data using the IoT device based on the updated configuration value. The second processor includes a graphics processing unit associated with a data analysis pipeline for multi-access edge computing in a 5G telecommunications network, and where the second processor is different from the first processor. The first gradient value includes an input configuration gradient, where the input configuration gradient indicates the degree of change in the input data based on the updated configuration value. The second gradient value includes an inference input gradient, where the inference input gradient indicates the degree of change in a confidence score associated with inference input data when the input data changes, and where the inference input gradient is based on the saliency associated with the neural network. The combination of the first gradient value and the second gradient value represents an inference configuration gradient, which is used to adjust the configuration value of the configuration setting for the input operation of the IoT device, where the inference configuration gradient indicates the degree of change in the confidence score associated with inferring the input data based on the updated configuration value. The configuration value is associated with an image resolution for capturing content. The first processor is further configured to execute a method that includes: receiving subsequent input data based on the updated configuration value associated with the configuration setting for operating the IoT device; and updating, at least based on a combination of a subsequent change in the configuration value including the configuration value and a subsequent change in inferring the subsequent input data, the configuration value of the configuration setting for further operating the IoT device.
[0100] On the other hand, the technology relates to an IoT device among a plurality of IoT devices in edge computing connected to an edge server. The IoT device includes a memory; and a first processor configured to execute a method that includes: receiving input data based on a configuration value associated with a configuration setting of the IoT device; determining, by the first processor, a first gradient value, where the first gradient value represents a change in the input data relative to previously received input data based on a change in the configuration setting for using the IoT device to capture content; causing the edge server to determine, by a second processor using a neural network, a second gradient value, where the second gradient value indicates a change in inference data based on the change in the input data, where the first processor and the second processor are different; updating, at least based on a combination of the first gradient value and the second gradient value, a configuration value associated with a configuration setting for adjusting an input operation of the IoT device, where the combination of the first gradient value and the second gradient value represents an expected change in the inference data relative to a change in the configuration value associated with the configuration setting of the IoT device, and where the update results in improved inference of the input data by adjusting the configuration setting of the IoT device; and receiving subsequent input data using the IoT device based on the updated configuration value. The second processor includes a graphics processing unit associated with a data analysis pipeline of multi-access edge computing in a 5G telecommunications network, and where the second processor is different from the first processor. The first gradient value includes an input configuration gradient, where the input configuration gradient indicates the degree of change in the input data based on the updated configuration value. The second gradient value includes an inference input gradient, where the inference input gradient indicates the degree of change in a confidence score associated with inference input data when the input data changes, and where the inference input gradient is based on the saliency associated with the neural network. The combination of the first gradient value and the second gradient value represents an inference configuration gradient, where the inference configuration gradient indicates the degree of change in a confidence score associated with inferring the input data based on the updated configuration value. The configuration value is associated with an image resolution for using the IoT device to capture content.
[0101] Any one of the above one or more aspects can be combined with any other one of the above one or more aspects. Any one of the one or more aspects as described herein.
Claims
1. A method for updating configuration settings associated with content captured using an Internet of Things (IoT) device, the IoT device being one of a plurality of IoT devices that are at least part of an edge computing system, the IoT device including a first processor, the method comprising: Receiving input data based on a configuration value associated with the configuration settings of the IoT device; Determining, by the first processor, a first gradient value, wherein the first gradient value represents a change in the input data relative to previously received input data based on a change in the configuration settings for using the IoT device to capture content; Causing an edge server of the edge computing system to determine, by a second processor using a neural network, a second gradient value, wherein the second gradient value indicates a change in inference data based on the change in the input data, wherein the first processor and the second processor are different; Updating the configuration value associated with the configuration settings for adjusting an input operation of the IoT device based at least on a combination of the first gradient value and the second gradient value, wherein the combination of the first gradient value and the second gradient value represents an expected change in the inference data relative to the change in the configuration value associated with the configuration settings of the IoT device, and wherein the update is caused by adjusting the configuration settings of the IoT device to improve the inference of the input data; And Receiving subsequent input data using the IoT device based on the updated configuration value.
2. The method according to claim 1, wherein the second processor includes a graphics processing unit associated with a data analysis pipeline for multi-access edge computing in a 5G telecommunications network, and wherein the second processor is different from the first processor.
3. The method according to claim 1, wherein the first gradient value includes an input configuration gradient, wherein the input configuration gradient indicates the degree of change in the input data based on updating the configuration value.
4. The method according to claim 1, wherein the second gradient value includes an inference input gradient, wherein the inference input gradient indicates the degree of change in a confidence score associated with inferring the input data when the input data changes, and wherein the inference input gradient is based on the saliency associated with the neural network.
5. The method according to claim 1, wherein the combination of the first gradient value and the second gradient value represents an inference configuration gradient, wherein the inference configuration gradient indicates the degree of change in a confidence score associated with inferring the input data based on updating the configuration value.
6. The method according to claim 1, wherein the configuration value is associated with an image resolution for capturing content by the IoT device.
7. The method according to claim 1, further comprising: Receiving the subsequent input data based on the updated configuration value associated with the configuration settings for operating the IoT device; And Update the configuration value of the configuration setting for further operation of the IoT device based at least on a combination of subsequent changes including the configuration value and subsequent changes in inferring the subsequent input data.
8. A system for capturing content using an IoT device among a plurality of Internet of Things (IoT) devices in an edge computing system, the IoT device including a first processor, the system including: A memory; And The first processor configured to execute a method, the method including: Receiving input data based on a configuration value associated with a configuration setting of the IoT device for capturing content; Determining, by the first processor, a first gradient value, where the first gradient value represents a change in the input data relative to previously received input data based on a change in the configuration setting for using the IoT device to capture content; Causing an edge server of the edge computing system to determine, by a second processor using a neural network, a second gradient value, where the second gradient value indicates a change in inferred data relative to previously generated inferred data based on the change in the input data, where the first processor and the second processor are different; Updating, at least based on a combination of the first gradient value and the second gradient value, the configuration value associated with the configuration setting for adjusting an input operation of the IoT device, where the combination of the first gradient value and the second gradient value represents an expected change in the inferred data relative to a change in the configuration value associated with the configuration setting of the IoT device, and where the update is caused by adjusting the configuration setting of the IoT device to improve the inference of the input data; and Receiving subsequent input data using the IoT device based on the updated configuration value.
9. The system according to claim 8, where the second processor includes a graphics processing unit associated with a data analysis pipeline for multi-access edge computing in a 5G telecommunications network, and where the second processor is different from the first processor.
10. The system according to claim 8, where the first gradient value includes an input configuration gradient, where the input configuration gradient indicates the degree of change in the input data based on updating the configuration value.
11. The system according to claim 8, where the second gradient value includes an inference input gradient, where the inference input gradient indicates the degree of change in a confidence score associated with inferring the input data when the input data changes, and where the inference input gradient is based on the significance associated with the neural network.
12. The system according to claim 8, where the combination of the first gradient value and the second gradient value represents an inference configuration gradient for adjusting the configuration value of the configuration setting for an input operation of the IoT device, where the inference configuration gradient indicates the degree of change in a confidence score associated with inferring the input data based on updating the configuration value.
13. The system according to claim 8, wherein the configuration value is associated with an image resolution for capturing content.
14. An IoT device among a plurality of IoT devices in edge computing connected to an edge server, the IoT device comprising: a memory; and a first processor configured to execute a method, the method comprising: receiving input data based on a configuration value associated with a configuration setting of the IoT device; determining, by the first processor, a first gradient value, wherein the first gradient value represents a change in the input data relative to previously received input data based on a change in the configuration setting for using the IoT device to capture content; causing the edge server to determine, by a second processor using a neural network, a second gradient value, wherein the second gradient value indicates a change in inference data based on the change in the input data, wherein the first processor and the second processor are different; updating, at least based on a combination of the first gradient value and the second gradient value, the configuration value associated with the configuration setting for adjusting an input operation of the IoT device, wherein the combination of the first gradient value and the second gradient value represents an expected change in the inference data relative to a change in the configuration value associated with the configuration setting of the IoT device, and wherein the update is caused by adjusting the configuration setting of the IoT device to improve the inference of the input data; and receiving subsequent input data using the IoT device based on the updated configuration value.
15. The IoT device according to claim 14, wherein the second processor comprises a graphics processing unit associated with a data analysis pipeline of multi-access edge computing in a 5G telecommunications network, and wherein the second processor is different from the first processor.