An autonomous deep learning image sensing data acquisition method
Patent Information
- Application Number
- CN202410286763.0
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Filing Date
- 2024-03-13
- Publication Date
- 2026-10-09
- Estimated Expiration
- 2044-03-13
AI Technical Summary
综上所述,PCA算法和AE算法在图像压缩中存在信息丢失、重构误差、统计特征依赖性和计算成本等方面的不足之处;2)Hermit恢复出的特征曲线在相邻两个采样时刻点组成的时间窗口内具有二阶连续可导性,但在时间窗口分割点处只满足一阶连续性,导致在采样周期内恢复出的整条曲线由于存在二阶间断点而平滑性较差
[0045] This invention first improves the Xception network by combining it with the generative model VAE to design a novel feature extraction model, X-VAE, which can more effectively extract features from high-dimensional images. Furthermore, we divide data processing into two parts: a cloud-based system and a local sensor network S. For the sensor network, we propose a sampling algorithm that allows sampling only at the next sampling moment calculated by the algorithm, thus reducing the sampling frequency of the sensor network, saving energy, and extending its lifespan. Simultaneously, we design an effective data filtering mechanism to effectively reduce network bandwidth consumption during data transmission without significantly affecting data recovery accuracy.
Smart Images

Figure CN118154832B_ABST
Abstract
Description
Technical Field
[0001] This invention relates to the field of deep learning technology, and more particularly to an autonomous deep learning method for acquiring image sensing data. Background Technology
[0002] Image sensor networks play a crucial role in real-world applications such as marine debris detection, smart homes, and wildlife conservation. An image sensor network typically consists of multiple smart cameras deployed in the monitoring environment. Like traditional sensor networks, the image data acquisition process involves two phases: data acquisition and data transmission. In the data acquisition phase, each image sensor node samples data from the real physical world at a specific sampling frequency. In the data transmission phase, each sensor node transmits the acquired discrete image data to the cloud for further processing or permanent storage. Furthermore, during this data acquisition process, image sensor networks face the same energy and bandwidth constraints as traditional sensor networks.
[0003] To reduce energy consumption at sensor nodes, adaptive frequency conversion sampling methods are typically employed during data acquisition. However, existing frequency conversion sampling algorithms primarily handle univariate data and cannot be directly applied to high-dimensional image data. Furthermore, the main objective of most existing algorithms is to conserve network energy, and the loss of accuracy is primarily measured based on equal-frequency sampling algorithms rather than the changes in the real physical world. This can lead to the loss of some key data points (such as extreme points) or even provide incorrect data to upper-layer applications. To address these issues, some literature has proposed a physical world-aware frequency conversion sampling algorithm. Based on a given error threshold, this algorithm adaptively calculates the sampling time points using Hermit interpolation and Spline interpolation, respectively, and reconstructs the collected discrete data points into curves to describe the continuously and smoothly changing real physical world. However, this method is only proposed for unidimensional data and cannot be directly applied to high-dimensional image data.
[0004] Currently, existing algorithms still have the following shortcomings in the image sensing data acquisition process: 1) In the data compression stage, the widely used data compression algorithms are PCA and CAE models, but they have some deficiencies. First, the PCA algorithm suffers from information loss in image compression. The PCA algorithm achieves compression by transforming the image into a lower-dimensional space, but in this process, some detailed information of the image may be lost. This may lead to a decrease in the quality of the compressed image, especially for images containing complex textures or subtle changes. Second, the AE algorithm may have a large reconstruction error problem in image compression. The AE algorithm achieves compression by learning the encoding and decoding process of the input image, but this process may lead to reconstruction errors. Reconstruction error refers to the difference between the decoded image after compression and the original image. For some images, the AE algorithm may not be able to completely and accurately reconstruct the original image, resulting in unsatisfactory quality of the compressed image. In addition, both the PCA and AE algorithms are somewhat dependent on the statistical features of the image. This means that if the statistical features of the input image do not match those used in the training data, the quality of the compressed image may be affected. For example, if the images in the training data are mainly indoor scenes, while the input images are outdoor landscapes, the quality of the compressed image may decrease. In summary, PCA and AE algorithms suffer from shortcomings in image compression, including information loss, reconstruction errors, dependence on statistical features, and computational costs. 2) The feature curve recovered by Hermit exhibits second-order continuous differentiability within a time window formed by two adjacent sampling points, but only first-order continuity at the time window segmentation points. This results in poor smoothness of the recovered curve due to second-order discontinuities within the sampling period. However, changes in the real physical world are often smooth and continuous, making curve smoothness crucial for image restoration quality. In practical applications, this is mainly reflected in the utilization of curve inflection points (i.e., points where the second derivative is zero). 3) Traditional data transmission algorithms acquire and process data through node autonomy, assuming each sensing node possesses a static data compression network. They do not consider incremental network updates, energy consumption of sensing nodes during transmission, or network bandwidth load. However, in reality, due to the complex structure and high dimensionality of image data, even with multiple nodes performing frequency-modulated sampling and simultaneously transmitting image data to the cloud, network bandwidth remains heavily burdened, increasing the risk of network congestion and data transmission failures. Summary of the Invention
[0005] In view of this, the purpose of this invention is to propose an autonomous deep learning method for acquiring image sensing data. This method simultaneously considers the node energy constraints and network bandwidth limitations faced in both the data acquisition and data transmission stages, as well as the changes in the real physical world. The data acquisition process is divided into two parts: the cloud and the sensor nodes. The cloud is mainly responsible for training and incrementally updating the data compression network, as well as approximate recovery of image data from non-sampling time points. The sensor nodes are mainly responsible for acquiring image data from the real physical world using a smoothness-sensitive frequency conversion sampling method and transmitting the data based on a data filtering mechanism.
[0006] The technical means employed in this invention are as follows:
[0007] A method for acquiring image sensing data using autonomous deep learning includes the following steps:
[0008] S1. Cloud-based compressed network pre-training: Given historical data collected by the sensor network, the historical data is input into the X-VAE network for training, resulting in a converged pre-trained X-VAE model, which is then distributed to the local sensor node network.
[0009] The Xception network, incorporating LayerScaler and Droppath regularization, is used as the encoder in X-VAE to extract features from the input image and output the corresponding low-dimensional features. These low-dimensional features are then input into the decoder, which incorporates a deep supervision mechanism. The decoder performs upscaling and recovery on the low-dimensional features to obtain the pre-trained X-VAE network.
[0010] S2. The local image sensor network S performs adaptive sampling: The image is sampled using an image compression model to obtain the image data at the current sampling time. Based on the image data at the current sampling time, the image data at the previous sampling time, and the X-VAE model transmitted from the cloud, the current and previous sampling time data are used as inputs into the compression network to obtain two low-dimensional features. The two low-dimensional features are then input into the sampling algorithm to calculate the next sampling time, and sampling is only performed at the sampling time to obtain the next sampled data.
[0011] S3, Local Image Sensor Network S filters out some data and uploads it to the cloud: Based on the X-VAE model obtained by S1 and the next sampling data collected by S2, the first sampling data is used as input into the X-VAE model to obtain the reconstruction error. The reconstruction error of each image is recorded and judged according to the error threshold ε specified by the user. If the reconstruction error is greater than the error threshold ε specified by the user, it is marked as filterable data and the filterable data is transmitted to the cloud. The filterable data is used as input to replace the historical data of S1 and incrementally update the X-VAE model. Thus, the X-VAE model is continuously iterated in the order of S1, S2, and S3, and the iterated X-VAE model is continuously output to ensure continuous operation and acquisition of image sensing data.
[0012] S4. Perform incremental updates of X-VAE based on the partially uploaded data, and restore all data using X-VAE: Data restoration is performed periodically. In the cloud, after each period when the X-VAE model completes the incremental update, the non-sampling time features are calculated based on the uploaded sampling time features. The decoder in the updated X-VAE model is then used to restore the data, and the restored data is stored in the cloud to obtain complete image sensing data.
[0013] Furthermore, S1 specifically includes the following steps:
[0014] S11: Use the historical data stored at time zero as the initial training dataset, and pre-train the network based on the initial training dataset.
[0015] S12: Perform preprocessing operations such as standardization and normalization on the initial training dataset;
[0016] S13: Input the dataset output from S12 into the X-VAE model. Improve feature extraction capability by repeating the block structure, residual connections, and separable convolution techniques in Xception to capture rich features. Accelerate training using LayerScaler and DropPath regularization techniques to improve the model's flexibility and robustness. At the same time, add a deep supervision mechanism to the decoder to process the output of each intermediate layer. Add an auxiliary reconstruction branch to the output, which runs in parallel with the main decoding path to directly generate image descriptions. The combination of complex encoder and decoder structures improves the model's accuracy. Finally, the pre-trained network X-VAE is obtained.
[0017] S14: Transfer the pre-trained X-VAE model to the local sensor node network.
[0018] Furthermore, in S13:
[0019] The separable convolution consists of Depthwise Convolution and Pointwise Convolution. Depthwise Convolution is used to learn the spatial relationships in the input feature map, using a 3×3 convolution kernel to handle one channel of the input feature map. Pointwise Convolution is used to learn the channel relationships in the input feature map, using a 1×1 convolution kernel, meaning that each point is traversed, and the depth of the convolution kernel is equal to the number of channels in the input image.
[0020] The block structure consists of residual connections and separable convolutions, which gradually reduce the spatial resolution of the feature map to obtain a more abstract feature representation.
[0021] Deep supervision mechanisms achieve feature representation at multiple scales by directly transforming the output of intermediate layers into output scales.
[0022] LayerScaler scales the input feature map using learnable parameters, dynamically adjusting the scale of the feature map according to the input data and task requirements;
[0023] DropPath randomly masks the output of convolutional blocks with a certain probability to perform model regularization and avoid overfitting. In X-VAE, LayerScaler is combined with DropPath to scale the output of each layer, thereby avoiding overfitting.
[0024] X-VAE consists of an Encoder, latent variables, and a Decoder. The X-VAE Encoder is composed of separable convolutions and blocks, and the Decoder is composed of separable inverse convolutions and inverse convolutions. It also incorporates a deep supervision mechanism.
[0025] Furthermore, S2 specifically includes:
[0026] S21: Use the Prid dataset to train an image compression model to obtain a pre-trained image compression model, making the extracted features more consistent with the main features of the image, and then distribute the pre-trained image compression model to the local sensor network.
[0027] S21: Train an image compression model using the Prid dataset to obtain a pre-trained image compression model. This makes the extracted features more consistent with the main features of the image. The pre-trained image compression model is then distributed to the local image sensor network S, which consists of several sensor nodes {s1, s2, ... s...}. n}composition;
[0028] S22: The local image sensor S uses the pre-trained image compression model X-VAE to extract features from the historical images acquired by S at time zero, and then uses the current sampling node s i Sampling time t i,c The obtained low-dimensional image features F(t) i,c ) as input to the sampling algorithm;
[0029] S23: Execute the sampling algorithm, based on the current sampling node s. i Sampling time t i,c The previous sampling time t i,c-1 , and its corresponding eigenvector F(t) i,c ), F(t) i,c-1 ), calculate the previous sampling interval [t] of the current sample. i,c-1 ,t i,c The maximum fourth derivative F (4) (ξ i,k );
[0030] S24: Calculate the value in the previous sampling interval [t] using the maximum fourth derivative. i,c-1 ,t i,c The curve of the i-th eigenvector changing with time is obtained to determine the next sampling time point t. i,c+1 ;
[0031] S25: Sample at each sampling time to obtain low-dimensional features of the image and transmit them to the cloud.
[0032] Furthermore, S3 specifically includes:
[0033] S31: Utilize the sampling nodes s in the local sensor network S of each S22 i The images sampled at each sampling time are used as input to the pre-trained X-VAE model, and the error of the output compressed image is sorted to quantitatively evaluate the contribution of each data sample to the network update.
[0034] S32: Sort the reconstruction errors in descending order so that the system can quickly and accurately identify the data that is most valuable for model training and performance improvement in the context of the current task;
[0035] S33: Transmit the reconstruction error data of each node to the cloud to obtain image data that is greater than the user-specified threshold range;
[0036] S34: This error is transmitted back to the local sensor. Images with local image reconstruction errors greater than this error are filtered out and transmitted to the cloud, thus completing the data filtering process.
[0037] S35: After obtaining the selected images from the cloud, the pre-trained X-VAE model is iteratively trained, thereby achieving incremental updates to the network.
[0038] Furthermore, S4 specifically includes:
[0039] S41: Use the low-dimensional features at the sampling time and the incrementally updated network as necessary parts of the feature input;
[0040] S42: Based on the feature vector curves of each dimension constructed at the sampling time points, the low-dimensional features at non-sampling time points are obtained. From the low-dimensional features at each time point, the complete temporal feature vector is obtained.
[0041] S43: Using the incrementally updated Decoder, the complete temporal feature vector is upgraded to recover image data, thus completing cloud image storage.
[0042] The present invention also provides a storage medium comprising a stored program, wherein, when the program is executed, the autonomous deep learning image sensing data acquisition method described in any of the preceding claims is performed.
[0043] The present invention also provides an electronic device, including a memory, a processor, and a computer program stored in the memory and executable on the processor, wherein the processor executes any of the above-described autonomous deep learning image sensing data acquisition methods through the computer program.
[0044] Compared with the prior art, the present invention has the following advantages:
[0045] This invention first improves the Xception network by combining it with the generative model VAE to design a novel feature extraction model, X-VAE, which can more effectively extract features from high-dimensional images. Furthermore, we divide data processing into two parts: a cloud-based system and a local sensor network S. For the sensor network, we propose a sampling algorithm that allows sampling only at the next sampling moment calculated by the algorithm, thus reducing the sampling frequency of the sensor network, saving energy, and extending its lifespan. Simultaneously, we design an effective data filtering mechanism to effectively reduce network bandwidth consumption during data transmission without significantly affecting data recovery accuracy.
[0046] In summary, the technical solution of this invention can balance energy and bandwidth limitations with changes in the real physical world. Experimental results show that this method performs excellently in extending network lifetime and improving image restoration quality. Regarding sensor data acquisition, we adopted an adaptive sampling algorithm that allows the system to adjust sensor parameters and algorithms according to changes in the current environment, thereby adapting to different physical environments and providing more accurate image restoration results. Simultaneously, an incremental learning image compression algorithm is used in the cloud, which continuously learns new image information to approximate real-world data as closely as possible. Through flexible adjustments to the working modes of the sensors and the cloud, we successfully achieved efficient use of energy and bandwidth, thereby improving the overall system efficiency. Experimental verification shows that the autonomous deep learning method provided in this study can significantly extend network lifetime, reduce energy consumption, and improve image restoration quality. Compared with traditional methods, our method excels in detail preservation and noise suppression in image restoration, providing clearer and more accurate image restoration results, and providing strong support for research and application in the field of image sensor acquisition methods. Attached Figure Description
[0047] To more clearly illustrate the technical solutions in the embodiments of the present invention or the prior art, the drawings used in the description of the embodiments or the prior art will be briefly introduced below. Obviously, the drawings described below are some embodiments of the present invention. For those skilled in the art, other drawings can be obtained based on these drawings without creative effort.
[0048] Figure 1 This is a schematic diagram of the method flow of the present invention.
[0049] Figure 2 This is a diagram of the overall model of the missing value filling method of the present invention. Detailed Implementation
[0050] To enable those skilled in the art to better understand the present invention, the technical solutions of the present invention will be clearly and completely described below with reference to the accompanying drawings of the embodiments of the present invention. Obviously, the described embodiments are only some embodiments of the present invention, and not all embodiments. Based on the embodiments of the present invention, all other embodiments obtained by those skilled in the art without creative effort should fall within the scope of protection of the present invention.
[0051] It should be noted that the terms "first," "second," etc., in the specification, claims, and accompanying drawings of this invention are used to distinguish similar objects and are not necessarily used to describe a specific order or sequence. It should be understood that such data can be interchanged where appropriate so that the embodiments of the invention described herein can be implemented in orders other than those illustrated or described herein. Furthermore, the terms "comprising" and "having," and any variations thereof, are intended to cover a non-exclusive inclusion; for example, a process, method, system, product, or apparatus that comprises a series of steps or units is not necessarily limited to those steps or units explicitly listed, but may include other steps or units not explicitly listed or inherent to such processes, methods, products, or apparatus.
[0052] like Figure 1 As shown, this invention provides an autonomous deep learning method for acquiring image sensing data, comprising the following steps:
[0053] S1: Cloud-based compressed network pre-training: Given a sensor network detection region with cloud and local image sensor nodes, the local image sensor network consists of several sensor nodes {s1, s2, ... s... n The system consists of a cloud-based server cluster with powerful computing capabilities. Given historical data collected by the sensor network, the historical data I(t0) is standardized to obtain preprocessed data, which is then input into the X-VAE network for training. A converged pre-trained X-VAE model is obtained and distributed to the local sensor node network. Finally, the trained compressed X-VAE network is transmitted to the local nodes.
[0054] Regarding X-VAE: The Xception network is used as the encoder in X-VAE for feature extraction. However, due to the depth of the Xception network, regularization techniques such as LayerScaler and DropPath are added to accelerate training and improve reconstruction results in order to achieve the desired reconstruction outcome. In the decoder part, to make the image reconstruction more closely resemble real-world data, a deep supervision mechanism is used to process the outputs of each intermediate layer, and auxiliary reconstruction branches are added after the output. These branches directly generate image descriptions and run in parallel with the main decoding path.
[0055] S11: Obtain the historical data I(t0) that was originally stored in the cloud;
[0056] S12: Perform preprocessing operations such as standardization and normalization on the dataset obtained in S11;
[0057] S13: In the X-VAE model, the feature extraction capability is improved by repeating the block structure, residual connection, and separable convolution in Xception, thereby capturing rich features. Regularization techniques such as LayerScaler and DropPath are used to accelerate training and improve the flexibility and robustness of the model.
[0058] S14: The trained X-VAE model is transmitted to the local image sensor node network S.
[0059] S2: Local image sensor network S performs adaptive sampling: local image sensor node s i Using the compressed network model X-VAE received at S14, the image I obtained at the current sampling time is... i Image dimensionality reduction to obtain low-dimensional features F(t) i,c According to the low-dimensional feature F(t) i,c ) Calculate the next sampling time point t c+1 Then at the next sampling time point t c+1 Sampling is performed, and the process is repeated iteratively.
[0060] S21: Local image sensor node i Received the compressed network model X-VAE sent from the cloud;
[0061] S22: Use the Encoder part of X-VAE to reduce the dimensionality of the image at the current sampling time, and obtain the image I at the current sampling time. i Low-dimensional features F(t) i,c );
[0062] S23: Based on the current sampling time t i,c The previous sampling time t i,c-1 , and its corresponding eigenvector F(t) i,c ), F(t) i,c-1 ), calculate the previous sampling interval [t i,c-1 ,t i,c The maximum fourth derivative Since time-series data are often similar, we will consider the previous sampling interval [t] i,c-1 ,t i,c The maximum fourth derivative The default value is equal to the next sampling interval [t]. i,c ,t i,c+1 The maximum fourth derivative
[0063] S24: Calculate the value in the previous sampling interval [t] using the maximum fourth derivative. i,c-1 ,t i,cThe curve of the i-th eigenvector changing with time is obtained by using Hermit interpolation with the maximum fourth derivative to calculate the next sample t. i,c+1 ;
[0064] S25: Image data I obtained from sampling at the sampling time point i As input, in local node s i Enter the compressed network X-VAE;
[0065] S3: The local image sensor network S filters out a portion of the data and uploads it to the cloud: Data I is obtained at each sampling time. i Then, the sampled data is filtered, and data with errors exceeding the user-specified threshold is transmitted to the cloud.
[0066] S31: Image data I entering the X-VAE i Compression is performed, and the compression error of each image data is obtained.
[0067] S32: Image data with compression error greater than the user-specified threshold ∈ is transmitted to the cloud as filtered data FI(t).
[0068] S4: Using the filtered data transmitted to the cloud as input, incrementally update the compressed network, then use the updated network to restore the data at non-sampling time points, then transmit the updated network to the local node, and repeat the above S1-S3.
[0069] S41: The filtered image FI(t) transmitted to the cloud is used as input to iteratively train the compressed network;
[0070] S42: Using the low-dimensional features at the sampling time points as input, and employing Hermit interpolation, the feature vector values at the non-sampling time points are obtained from the feature vector curve, thus recovering the low-dimensional features at the non-sampling time points. Ultimately, complete low-dimensional features at each time step are obtained;
[0071] S43: Input the interpolated and restored complete low-dimensional features into the updated X-VAE decoder to recover the images at all times and store them in the cloud;
[0072] S44: Distribute the updated network to the local node.
[0073] S45: Iterate through S1-S4 above to ensure continuous image acquisition and storage in the cloud.
[0074] Figure 2The overall framework diagram of the proposed model is shown, which shows the detailed structure of the encoder and decoder networks, including the compression network, and finally the image is recovered. Figure 2 Medium-light dashed arrows represent the transformation from the output of the current intermediate layer to the final output, while heavy dashed arrows represent the transformation from the output of every two blocks to the final output.
[0075] Table 1 shows a comparison of the image compression performance of this invention, and Table 2 shows a comparison of the image restoration performance. Two datasets were used in the experiments: the pedestrian dataset Prid and the pedestrian tracking dataset Mars. Experiments were conducted on both compression error and image restoration. The criterion for judgment was the RMSE value between the original image and the inpainted image; the smaller the value, the smaller the difference between the inpainted image and the original image, indicating better performance. This demonstrates the superiority and effectiveness of the current model in image restoration.
[0076] Table 1 Comparison of Compression Error Effects of Single-Channel and Three-Channel Datasets in this Invention
[0077] CAEs X-VAE CAEs X-VAE 64 0.0435135 0.000548 0.122752792 0.000367 128 0.037462969 0.000549 0.108059 0.000347 256 0.031050925 0.000546 0.098027 0.000355 512 0.032280153 0.000546 0.1004345 0.000346 1024 0.036291422 0.000545 0.080825 0.000351
[0078] CAEs X-VAE CAEs X-VAE 64 0.087562758 0.002514 0.1999 0.001919 128 0.077311994 0.002565 0.20973387 0.001943 256 0.065903321 0.002549 0.188218 0.001929 512 0.066838603 0.02523 0.214714216 0.001958 1024 0.072513942 0.002491 0.272204507 0.001991
[0079] Table 2 Interpolation Recovery Error Table
[0080] 0.017261 0.050809 0.017363 0.0505 0.017 0.050047 0.026767 0.050418 0.023 0.051647
[0081] The present invention also provides a storage medium comprising a stored program, wherein, when the program is executed, an image sensing data acquisition method based on autonomous deep learning is performed.
[0082] The present invention also provides an electronic device, including a memory, a processor, and a computer program stored in the memory and executable on the processor, wherein the processor executes an image sensing data acquisition method that performs autonomous deep learning through the computer program.
[0083] Finally, it should be noted that the above embodiments are only used to illustrate the technical solutions of the present invention, and not to limit them; although the present invention has been described in detail with reference to the foregoing embodiments, those skilled in the art should understand that modifications can still be made to the technical solutions described in the foregoing embodiments, or equivalent substitutions can be made to some or all of the technical features; and these modifications or substitutions do not cause the essence of the corresponding technical solutions to deviate from the scope of the technical solutions of the embodiments of the present invention.
Claims
1. A method for acquiring image sensing data through autonomous deep learning, characterized in that, Includes the following steps: S1. Cloud-based compressed network pre-training: Given historical data collected by the sensor network, the historical data is input into the X-VAE network for training, resulting in a converged pre-trained X-VAE model, which is then distributed to the local sensor node network. The Xception network, incorporating LayerScaler and Droppath regularization, is used as the encoder in X-VAE to extract features from the input image and output the corresponding low-dimensional features. These low-dimensional features are then input into the decoder, which incorporates a deep supervision mechanism. The decoder performs upscaling and recovery on the low-dimensional features to obtain the pre-trained X-VAE network. S2. The local image sensor network S performs adaptive sampling: The image is sampled using an image compression model to obtain the image data at the current sampling time. Based on the image data at the current sampling time, the image data at the previous sampling time, and the X-VAE model transmitted from the cloud, the current and previous sampling time data are used as inputs into the compression network to obtain two low-dimensional features. The two low-dimensional features are then input into the sampling algorithm to calculate the next sampling time, and sampling is only performed at the sampling time to obtain the next sampled data. S3, Local Image Sensor Network S filters out some data and uploads it to the cloud: Based on the X-VAE model obtained by S1 and the next sampling data collected by S2, the first sampling data is used as input into the X-VAE model to obtain the reconstruction error. The reconstruction error of each image is recorded and judged according to the error threshold ε specified by the user. If the reconstruction error is greater than the error threshold ε specified by the user, it is marked as filterable data and the filterable data is transmitted to the cloud. The filterable data is used as input to replace the historical data of S1 and incrementally update the X-VAE model. Thus, the X-VAE model is continuously iterated in the order of S1, S2, and S3, and the iterated X-VAE model is continuously output to ensure continuous operation and acquisition of image sensing data. S4. Perform incremental updates of X-VAE based on the partially uploaded data, and restore all data using X-VAE: Data restoration is performed periodically. In the cloud, after each period when the X-VAE model completes the incremental update, the non-sampling time features are calculated based on the uploaded sampling time features. The decoder in the updated X-VAE model is then used to restore the data, and the restored data is stored in the cloud to obtain complete image sensing data.
2. The image sensing data acquisition method based on autonomous deep learning according to claim 1, characterized in that, S1 specifically includes the following steps: S11: Use the historical data stored at time zero as the initial training dataset, and pre-train the network based on the initial training dataset. S12: Perform preprocessing operations such as standardization and normalization on the initial training dataset; S13: Input the dataset output from S12 into the X-VAE model. Improve feature extraction capability by repeating the block structure, residual connections, and separable convolution techniques in Xception to capture rich features. Accelerate training using LayerScaler and DropPath regularization techniques to improve the model's flexibility and robustness. At the same time, add a deep supervision mechanism to the decoder to process the output of each intermediate layer. Add an auxiliary reconstruction branch to the output, which runs in parallel with the main decoding path to directly generate image descriptions. The combination of complex encoder and decoder structures improves the model's accuracy. Finally, the pre-trained network X-VAE is obtained. S14: Transfer the pre-trained X-VAE model to the local sensor node network.
3. The image sensing data acquisition method based on autonomous deep learning according to claim 2, characterized in that, In S13: The separable convolution consists of Depthwise Convolution and Pointwise Convolution. Depthwise Convolution is used to learn the spatial relationships in the input feature map, using a 3×3 convolution kernel to handle one channel of the input feature map. Pointwise Convolution is used to learn the channel relationships in the input feature map, using a 1×1 convolution kernel, meaning that each point is traversed, and the depth of the convolution kernel is equal to the number of channels in the input image. The block structure consists of residual connections and separable convolutions, which gradually reduce the spatial resolution of the feature map to obtain a more abstract feature representation. Deep supervision mechanisms achieve feature representation at multiple scales by directly transforming the output of intermediate layers into output scales. LayerScaler scales the input feature map using learnable parameters, dynamically adjusting the scale of the feature map according to the input data and task requirements; DropPath randomly masks the output of convolutional blocks with a certain probability to perform model regularization and avoid overfitting. In X-VAE, LayerScaler is combined with DropPath to scale the output of each layer, thereby avoiding overfitting. X-VAE consists of an Encoder, latent variables, and a Decoder. The X-VAE Encoder is composed of separable convolutions and blocks, and the Decoder is composed of separable inverse convolutions and inverse convolutions. It also incorporates a deep supervision mechanism.
4. The image sensing data acquisition method based on autonomous deep learning according to claim 1, characterized in that, S2 specifically includes: S21: Train an image compression model using the Prid dataset to obtain a pre-trained image compression model. This makes the extracted features more consistent with the main features of the image. The pre-trained image compression model is then distributed to the local image sensor network S, which consists of several sensor nodes {s1, s2, ... s...}. n }composition; S22: The local image sensor S uses the pre-trained image compression model X-VAE to extract features from the historical images acquired by S at time zero, and then uses the current sampling node s i Sampling time t i,c The obtained low-dimensional image features F(t) i,c ) as input to the sampling algorithm; S23: Execute the sampling algorithm, based on the current sampling node s. i Sampling time t i,c The previous sampling time t i,c-1 , and its corresponding eigenvector F(t) i,c ), F(t) i,c-1 ), calculate the previous sampling interval [t] of the current sample. i,c-1 ,t i,c The maximum fourth derivative F (4) (ξ i,k ); S24: Calculate the value in the previous sampling interval [t] using the maximum fourth derivative. i,c-1 ,t i,c The curve of the i-th eigenvector changing with time is obtained to determine the next sampling time point t. i,c+1 ; S25: Sample at each sampling time to obtain low-dimensional features of the image and transmit them to the cloud.
5. The image sensing data acquisition method based on autonomous deep learning according to claim 1, characterized in that, S3 specifically includes: S31: Utilize the sampling nodes s in the local sensor network S of each S22 i The images sampled at each sampling time are used as input to the pre-trained X-VAE model, and the error of the output compressed image is sorted to quantitatively evaluate the contribution of each data sample to the network update. S32: Sort the reconstruction errors in descending order so that the system can quickly and accurately identify the data that is most valuable for model training and performance improvement in the context of the current task; S33: Transmit the reconstruction error data of each node to the cloud to obtain image data that is greater than the user-specified threshold range; S34: This error is transmitted back to the local sensor. Images with local image reconstruction errors greater than this error are filtered out and transmitted to the cloud, thus completing the data filtering process. S35: After obtaining the selected images from the cloud, the pre-trained X-VAE model is iteratively trained, thereby achieving incremental updates to the network.
6. The image sensing data acquisition method based on autonomous deep learning according to claim 1, characterized in that, S4 specifically includes: S41: Use the low-dimensional features at the sampling time and the incrementally updated network as necessary parts of the feature input; S42: Based on the feature vector curves of each dimension constructed at the sampling time points, the low-dimensional features at non-sampling time points are obtained. From the low-dimensional features at each time point, the complete temporal feature vector is obtained. S43: Using the incrementally updated Decoder, the complete temporal feature vector is upgraded to recover image data, thus completing cloud image storage.
7. A storage medium, characterized in that, The storage medium includes a stored program, wherein when the program is executed, it performs the autonomous deep learning image sensing data acquisition method according to any one of claims 1 to 6.
8. An electronic device comprising a memory, a processor, and a computer program stored in the memory and executable on the processor, characterized in that, The processor executes the autonomous deep learning image sensing data acquisition method according to any one of claims 1 to 6 through the computer program.
Citation Information
Patent Citations
Compressed sensing reconstruction method and system based on deep learning
CN113052925A
Distributed machine learning model
EP3944146A1