Ascend-based High-Reliability Light Field Super-Resolution Edge Computing Method and System

By adopting Astron-based high-reliability light field super-resolution edge computing method in the edge computing environment, a number of innovative mechanisms are integrated to solve multiple challenges of light field super-resolution technology in the edge computing environment, achieving high-quality and consistent super-resolution imaging, and improving the reliability and adaptability of the system.

CN119904358BActive Publication Date: 2025-05-30深圳市斯贝达电子有限公司
View PDF 2 Cites 0 Cited by

Patent Information

Application Number
CN202510388991.3
Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Filing Date
2025-03-31
Publication Date
2025-05-30
Estimated Expiration
2045-03-31

AI Technical Summary

Technical Problem

The existing light field super-resolution technology faces problems such as network latency and bandwidth pressure caused by huge data volume, limited computing resources, large data quality differences, lack of reliability guarantee mechanisms, and neglect of real-time, stability and resource efficiency in edge computing environments.

Method used

Using Astron-based high-reliable light field super-resolution edge computing method, innovative mechanisms such as adaptive quality evaluation, dual-path super-resolution processing and intelligent fusion, and heterogeneous resource dynamic scheduling are integrated. Through DnCNN denoising processing, Transformer-CNN calibration, FPGA accelerated EfficientNet model and FFT processing, Astron NPU executes ESRGAN model and U-Net fusion network for quality score weighting, forming super-resolution imaging of 8K light field, and dynamically allocates resources through PSNR, SSIM evaluation and task scheduling modules to build an edge computing redundancy mechanism.

Benefits of technology

It significantly improves the reliability and adaptability of light field super-resolution processing in edge environments, improves the quality consistency of super-resolution results, and realizes a high reliability, high adaptability and high efficiency light field super-resolution system.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN119904358B_ABST
    Figure CN119904358B_ABST
Patent Text Reader

Abstract

This application relates to the field of data processing technology, and discloses a high-reliability light field super-resolution edge computing method and system based on Ascend. The method includes: collecting RAW data through a light field camera and denoising it with DnCNN; calibrating with Transformer-CNN to generate error coefficients and correction matrices; accelerating EfficientNet and FFT processing with FPGA to output scores and frequency domain results; the Ascend NPU executes ESRGAN to form a deep learning image; U-Net fusion obtains 8K imaging; quality assessment is carried out and resources are dynamically allocated to construct a redundancy mechanism. By constructing a high-reliability light field super-resolution edge computing method based on Ascend, this application integrates innovative mechanisms such as adaptive quality assessment, dual-path super-resolution processing and intelligent fusion, and heterogeneous resource dynamic scheduling, effectively improving the reliability and adaptability of light field super-resolution processing in the edge environment.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] This application relates to the field of data processing technologies, and particularly to a high-reliability light field super-resolution edge computing method and system based on Ascend. Background Art

[0002] In recent years, with the rapid development of fields such as computer vision, virtual reality, and autonomous driving, the demand for high-quality images has been increasing day by day. As a new imaging technology, the light field camera can simultaneously record the direction and intensity information of light in a scene, providing a rich data basis for subsequent image processing. Compared with traditional two-dimensional images, the four-dimensional light field data (spatial coordinates and angular coordinates) captured by the light field imaging technology provides a richer scene representation, making three-dimensional scene reconstruction and depth perception possible. The light field super-resolution reconstruction technology aims to utilize this rich light field data to break through the resolution limitation of traditional imaging systems and obtain higher-definition images, which has broad application prospects in fields such as medical imaging, security monitoring, and industrial inspection. The development routes of related technologies include traditional methods (based on frequency domain processing, interpolation, and regularization) and deep learning methods (convolutional neural networks, generative adversarial networks, etc.), but most research focuses on the cloud processing environment.

[0003] However, the existing light field super-resolution technologies face many challenges in the edge computing environment. First, the light field data is huge in volume, and the traditional cloud processing mode brings high network latency and bandwidth pressure; second, the computing resources of edge devices are limited, making it difficult to directly deploy complex deep learning models; third, the quality of light field data varies greatly in different scenarios, and a single processing algorithm is difficult to adapt to diverse application scenarios; fourth, the existing methods generally lack a reliability guarantee mechanism and cannot handle common edge scenario problems such as data anomalies and device overload; finally, the existing technologies mostly take processing quality as the only goal and ignore the comprehensive performance requirements such as real-time performance, stability, and resource efficiency in practical applications. Especially in the context of the increasing popularity of new AI chips such as Ascend, how to give full play to the advantages of heterogeneous computing resources and build a high-reliability edge light field super-resolution computing framework has become a key technical problem to be solved urgently. Summary of the Invention

[0004] This application provides a high-reliability light field super-resolution edge computing method and system based on Ascend, which is used to build a high-reliability light field super-resolution edge computing method based on Ascend, integrating innovative mechanisms such as adaptive quality assessment, dual-path super-resolution processing and intelligent fusion, and heterogeneous resource dynamic scheduling, effectively improving the reliability and adaptability of light field super-resolution processing in the edge environment.

[0005] In a first aspect, the present application provides a high-reliability light field super-resolution edge computing method based on Ascend. The high-reliability light field super-resolution edge computing method based on Ascend includes: collecting RAW format light field data through a light field camera, performing DnCNN denoising processing on the light field data to obtain preprocessed data in HDF5 format; calibrating through a Transformer-CNN model according to the preprocessed data to generate a depth error compensation coefficient and a distortion correction matrix; using the calibrated light field data, accelerating the EfficientNet model and performing FFT processing through an FPGA to output a reliability score and a frequency-domain super-resolution result; based on the reliability score and the frequency-domain super-resolution result, performing ESRGAN model calculation by an Ascend NPU to form a deep learning super-resolution image; weighting the quality scores of the frequency-domain super-resolution result and the deep learning super-resolution image through a U-Net fusion network to obtain a super-resolution image of an 8K light field; evaluating the super-resolution image of the 8K light field for PSNR and SSIM, and dynamically allocating Ascend NPU and FPGA resources by a task scheduling module to construct an edge computing redundancy mechanism.

[0006] In a second aspect, the present application provides a high-reliability light field super-resolution edge computing system based on Ascend. The high-reliability light field super-resolution edge computing system based on Ascend includes:

[0007] A denoising module, configured to collect RAW format light field data through a light field camera, perform DnCNN denoising processing on the light field data to obtain preprocessed data in HDF5 format;

[0008] A calibration module, configured to calibrate through a Transformer-CNN model according to the preprocessed data to generate a depth error compensation coefficient and a distortion correction matrix;

[0009] A processing module, configured to use the calibrated light field data, accelerate the EfficientNet model and perform FFT processing through an FPGA to output a reliability score and a frequency-domain super-resolution result;

[0010] A calculation module, configured to perform ESRGAN model calculation by an Ascend NPU based on the reliability score and the frequency-domain super-resolution result to form a deep learning super-resolution image;

[0011] A weighting module, configured to weight the quality scores of the frequency-domain super-resolution result and the deep learning super-resolution image through a U-Net fusion network to obtain a super-resolution image of an 8K light field;

[0012] An allocation module, configured to evaluate the super-resolution image of the 8K light field for PSNR and SSIM, and dynamically allocate Ascend NPU and FPGA resources by a task scheduling module to construct an edge computing redundancy mechanism.

[0013] In a third aspect of the present invention, a computer device is provided, including: a memory and at least one processor, wherein instructions are stored in the memory; the at least one processor invokes the instructions in the memory to cause the computer device to execute the above-mentioned Ascend-based high-reliability light field super-resolution edge computing method.

[0014] In a fourth aspect of the present invention, a computer-readable storage medium is provided, wherein instructions are stored in the computer-readable storage medium, and when it runs on a computer, it causes the computer to execute the above-mentioned Ascend-based high-reliability light field super-resolution edge computing method.

[0015] In the technical solution provided by this application, the present invention collects RAW-format light field data through a light field camera and performs DnCNN denoising processing to obtain structured preprocessed data in HDF5 format, ensuring the data quality and consistency for subsequent processing; calibrates the preprocessed data based on the Transformer-CNN hybrid architecture model to generate accurate depth error compensation coefficients and distortion correction matrices, effectively eliminating systematic errors in the light field imaging process; utilizes the FPGA-accelerated EfficientNet model and FFT processing technology to achieve hardware-level parallel processing capabilities, greatly improving the real-time performance of quality assessment and frequency-domain processing. At the same time, the output reliability score and frequency-domain super-resolution results provide double guarantees for subsequent processing; performs ESRGAN model calculations based on the Ascend NPU, making full use of the advantages of the NPU in deep learning inference to form a deep learning super-resolution image with rich texture details; innovatively weights the quality scores of the frequency-domain results and deep learning images through a U-Net fusion network, realizing the complementary enhancement of the advantages of different processing paths and solving the problem that a single algorithm is difficult to adapt to complex scenarios; adopts PSNR, SSIM evaluation and task scheduling module dynamic resource allocation strategies to construct a complete edge computing redundancy mechanism, significantly improving the fault tolerance and stability of the system. In view of the characteristics of light field data, the present invention has made a number of innovations at the algorithm design level: the DnCNN denoising model is optimized for the noise distribution characteristics of light field images, effectively retaining angular information; the Transformer-CNN hybrid architecture gives full play to the complementary advantages of Transformer in global relationship modeling and CNN in local feature extraction; the dynamic loss function adjustment mechanism based on reliability scores enables the ESRGAN model to adaptively balance fidelity and perceptual quality according to data quality; the channel attention mechanism of the U-Net fusion network realizes the optimal fusion of frequency-domain processing and deep learning results. The combined action of these algorithm characteristics enables the system to automatically select the best processing strategy for light field data of different qualities in resource-constrained edge environments, significantly improving the quality consistency of super-resolution results. Overall, the present invention not only overcomes the problems of resource limitations and insufficient reliability in the existing edge environment technology, but also realizes the high reliability, high adaptability and high efficiency of the light field super-resolution system through the collaborative optimization of heterogeneous computing resources and the dynamic adjustment of intelligent processing strategies. BRIEF DESCRIPTION OF THE DRAWINGS

[0016] To more clearly illustrate the technical solutions of the embodiments of the present invention, the following will briefly introduce the drawings required for the description of the embodiments. Obviously, the following drawings are some embodiments of the present invention. For those of ordinary skill in the art, without creative efforts, other drawings can also be obtained based on these drawings.

[0017] Figure 1 Schematic diagram of an embodiment of the Ascend-based high-reliability light field super-resolution edge computing method in an embodiment of the present application;

[0018] Figure 2 Schematic diagram of an embodiment of the Ascend-based high-reliability light field super-resolution edge computing system in an embodiment of the present application;

[0019] Figure 3 Block diagram showing the structure of a computer device in an embodiment of the present invention. Detailed implementation manners

[0020] The embodiments of the present application provide an Ascend-based high-reliability light field super-resolution edge computing method and system. Terms such as "first", "second", "third", "fourth", etc. (if any) in the specification, claims and the above drawings of the present application are used to distinguish similar objects, and do not necessarily describe a specific order or sequence. It should be understood that such data can be interchanged under appropriate circumstances so that the embodiments described herein can be implemented in an order other than that illustrated or described herein. In addition, the terms "comprising" or "having" and any variations thereof are intended to cover non-exclusive inclusion. For example, a process, method, system, product or device that includes a series of steps or units does not necessarily limit to those clearly listed steps or units, but may include other steps or units not clearly listed or inherent to these processes, methods, products or devices.

[0021] For ease of understanding, the specific process of the embodiments of the present application is described below. Please refer to Figure 1 An embodiment of the Ascend-based high-reliability light field super-resolution edge computing method in an embodiment of the present application includes:

[0022] Step S101: Collect RAW format light field data through a light field camera, perform DnCNN denoising processing on the light field data to obtain preprocessed data in HDF5 format;

[0023] Step S102: Calibrate through a Transformer-CNN model according to the preprocessed data to generate a depth error compensation coefficient and a distortion correction matrix;

[0024] Step S103: Use the calibrated light field data, accelerate the EfficientNet model and perform FFT processing through FPGA to output a reliability score and a frequency domain super-resolution result;

[0025] Step S104: Based on the reliability score and the frequency domain super-resolution result, perform ESRGAN model calculation by an Ascend NPU to form a deep learning super-resolution image;

[0026] Step S105: Weight the quality scores of the frequency-domain super-resolution result and the deep learning super-resolution image through a U-Net fusion network to obtain the super-resolution imaging of the 8K light field;

[0027] Step S106: Evaluate the super-resolution imaging of the 8K light field in terms of PSNR and SSIM. The task scheduling module dynamically allocates Ascend NPU and FPGA resources to build an edge computing redundancy mechanism.

[0028] It can be understood that the execution entity of this application can be a high-reliability light field super-resolution edge computing system based on Ascend, or a terminal or a server. Specifically, it is not limited here. In this embodiment of the application, the server is used as the execution entity for illustration.

[0029] Specifically, RAW format light field data is collected by a high-precision light field camera such as the Lytro or Raytrix series. These data contain multi-view and depth information at 4K resolution. During the collection process, the operating environment parameters of the light field camera are monitored. When the temperature exceeds 60°C or the power supply fluctuation exceeds ±5%, resulting in a signal-to-noise ratio lower than 20 dB, the system will automatically switch to a backup light field camera to ensure the continuity of data input. The collected RAW format light field data is then transmitted to an edge server based on Huawei Atlas 500, and the RK3588 edge computing chip performs DnCNN denoising processing. DnCNN is a deep convolutional neural network optimized specifically for the noise characteristics of light field images, and the single-frame processing delay is controlled within 10 ms. During the processing, abnormal sub-images with a perspective shift exceeding 2 pixels are screened out based on the epipolar constraint rule, and low-quality regions are filtered through a deep confidence threshold of 0.8. The processed light field data is appended with metadata tags such as dynamic range and contrast statistics values, and finally encapsulated into preprocessed data in HDF5 format.

[0030] Calibration processing is performed based on the preprocessed data in HDF5 format. Here, the preprocessed data is parsed into a sequence of sub-aperture images in an N×N array to form a multi-view representation of the light field. Subsequently, a Transformer-CNN model with a hybrid architecture is used for light field parameter calibration. The Vision Transformer branch is responsible for analyzing the global perspective relationship and extracting spatial consistency features; while the ResNet-18 variant CNN branch focuses on local distortion localization and identifying high-frequency distortion regions. After fusing the two-way features, a perspective correlation mapping diagram is generated, thereby calculating the depth error compensation coefficient and the distortion correction matrix. These two key parameters are used to correct the optical distortion and depth error in the light field data.

[0031] Using the calibrated light field data, two-way parallel processing is performed on the Xilinx UltraScale+ FPGA. First, the quality of the light field data is evaluated through the FPGA-accelerated EfficientNet-B3 feature extraction layer, and the feature calculation time is compressed from 15 ms on the CPU to 5 ms. The quality evaluation constructs a weight-adjustable scoring function based on PSNR, SSIM, and edge sharpness metrics to calculate the reliability score. At the same time, the light field data is segmented into image blocks of n×n pixel size, and the fast Fourier transform (FFT) is performed to convert the spatial domain data into a frequency domain representation. Interpolation upsampling processing is carried out in the frequency domain, and then reconstruction is performed through the inverse FFT transform to generate the frequency domain super-resolution result. The internal dynamic resource control module of the FPGA flexibly allocates computing resources according to the processing load. By default, 80% is used for quality evaluation acceleration, and 20% is used for frequency domain processing. The reliability score and the frequency domain super-resolution result are transmitted to the Ascend NPU processing unit to construct a light field enhancement processing stream. An improved version of the ESRGAN network is deployed on the Ascend NPU, and the residual dense block structure is used to enhance the feature extraction ability. During the processing, the reliability score is incorporated into the loss function calculation to form a weighted loss function combining pixel loss, perceptual loss, and adversarial loss. Multilayer feature extraction is performed on the frequency domain result through the residual dense block to retain the light field angle information, and the phase consistency loss constraint algorithm is used to reduce texture artifacts in the multi-view data. Finally, a deep learning super-resolution image containing high-frequency details and accurate depth information is generated.

[0032] The frequency domain super-resolution result and the deep learning super-resolution image are fused. First, the high-frequency feature map is extracted from the frequency domain result, and the texture feature map is extracted from the deep learning image. Both are input into the fusion network with a lightweight U-Net structure. In the encoder stage of the fusion network, key information is identified through the channel attention mechanism; at the same time, the weight coefficient λ is dynamically adjusted based on the reliability score. When the score is greater than 0.9, the weight of the deep learning branch is set to 0.8, and when the score is less than or equal to 0.6, the weight is reduced to 0.3, relying more on the fidelity of the frequency domain result. The decoder of the fusion network reconstructs the processed features into an 8K resolution light field super-resolution image.

[0033] Perform multi-dimensional quality assessment on super-resolution imaging, including PSNR assessment (threshold 32 dB), SSIM assessment (threshold 0.92), and edge sharpness assessment (threshold 90). If the standard is not met, trigger the recalculation process, call the calibrated data from the storage module to re-execute the super-resolution process, and retry up to 3 times. At the same time, the task scheduling module dynamically allocates Ascend NPU and FPGA resources through the priority queue mechanism. In the normal mode, 70% of the NPU computing power is allocated for real-time super-resolution inference, and 30% is used for model incremental training. The system also continuously monitors the hardware operation parameters. When the temperature exceeds 85 °C or the utilization rate exceeds 90%, start the hardware redundancy switching mechanism, migrate the computing task to the standby computing node, and build a complete edge computing redundancy mechanism.

[0034] For example, after collecting the RAW format light field data of an indoor scene, the DnCNN denoising process improves the noise signal-to-noise ratio from the original 18 dB to 28 dB, and the epipolar constraint filtering removes approximately 8% of the abnormal sub-images. After calibrating the Transformer-CNN model, the distortion correction matrix corrects the radial distortion of 1.2 pixels in the edge region. In the FPGA acceleration processing stage, the quality assessment calculates a reliability score of 0.85, and the frequency-domain super-resolution processing retains the high-frequency details but has a slight ringing effect. After the Ascend NPU executes the ESRGAN model, the generated deep learning super-resolution image performs excellently in texture details but is slightly blurred in the depth transition region. The U-Net fusion network sets the weight coefficient to 0.7 according to the reliability score of 0.85, effectively combines the edge sharpness of the frequency-domain result with the texture richness of the deep learning result. Finally, the PSNR of the 8K super-resolution imaging reaches 33.5 dB, the SSIM reaches 0.94, and the edge sharpness reaches 95, meeting the quality requirements without recalculation. The entire processing flow is completed on the edge device, and the total computing time is controlled within 200 ms, fully leveraging the high efficiency of the Ascend NPU in deep learning inference and the advantages of the FPGA in parallel computing.

[0035] In the embodiments of the present application, the present invention collects RAW format light field data through a light field camera and performs DnCNN denoising processing to obtain structured HDF5 format preprocessed data, ensuring the data quality and consistency of subsequent processing; calibrates the preprocessed data based on a Transformer-CNN hybrid architecture model to generate accurate depth error compensation coefficients and distortion correction matrices, effectively eliminating systematic errors in the light field imaging process; utilizes the FPGA-accelerated EfficientNet model and FFT processing technology to achieve hardware-level parallel processing capabilities, greatly improving the real-time performance of quality assessment and frequency domain processing. At the same time, the output reliability score and frequency domain super-resolution results provide double guarantees for subsequent processing; performs ESRGAN model calculations based on the Ascend NPU, making full use of the advantages of the NPU in deep learning inference to form a deep learning super-resolution image with rich texture details; innovatively weights the quality scores of the frequency domain results and deep learning images through a U-Net fusion network, realizing the complementary enhancement of the advantages of different processing paths and solving the problem that a single algorithm is difficult to adapt to complex scenarios; adopts a PSNR, SSIM evaluation and task scheduling module dynamic resource allocation strategy to construct a complete edge computing redundancy mechanism, significantly improving the fault tolerance and stability of the system. The present invention makes a number of innovations at the algorithm design level for the characteristics of light field data: the DnCNN denoising model is optimized for the noise distribution characteristics of light field images, effectively retaining angular information; the Transformer-CNN hybrid architecture gives full play to the complementary advantages of Transformer in global relationship modeling and CNN in local feature extraction; the dynamic loss function adjustment mechanism based on reliability scores enables the ESRGAN model to adaptively balance fidelity and perceptual quality according to data quality; the channel attention mechanism of the U-Net fusion network realizes the optimal fusion of frequency domain processing and deep learning results. The combined action of these algorithm characteristics enables the system to automatically select the best processing strategy for light field data of different qualities in resource-constrained edge environments, significantly improving the quality consistency of super-resolution results. Overall, the present invention not only overcomes the problems of resource limitations and insufficient reliability in the existing edge environment technology, but also realizes the high reliability, high adaptability and high efficiency of the light field super-resolution system through the collaborative optimization of heterogeneous computing resources and the dynamic adjustment of intelligent processing strategies.

[0036] In a specific embodiment, the process of executing step S101 may specifically include the following steps:

[0037] (1) Collect 4K resolution RAW format light field data, and the sensor records the light direction and intensity information to form multi-view light field raw data;

[0038] (2)Monitor the operating environment parameters of the light field camera. When the temperature exceeds 60°C or the power supply fluctuates by more than ±5%, resulting in a signal-to-noise ratio lower than 20 dB, switch the acquisition task to the standby light field camera;

[0039] (3)Transmit the RAW format light field data to an edge server based on Huawei Atlas 500, and the RK3588 edge computing chip executes the computing task;

[0040] (4)Apply the DnCNN deep convolutional neural network model to the RAW format light field data to eliminate sensor noise and control the single-frame processing delay within 10 ms;

[0041] (5)Analyze the sub-images in the light field data according to the epipolar constraint rule, and filter out abnormal data regions with a perspective shift exceeding 2 pixels;

[0042] (6)Perform quality screening on the light field data according to the depth confidence threshold of 0.8, remove low-quality regions, and ensure data reliability;

[0043] (7)Attach metadata tags of dynamic range and contrast statistical values to the light field data and encapsulate it into preprocessed data in HDF5 format.

[0044] Specifically, collect RAW format light field data with a resolution of 4K. This process utilizes professional light field cameras such as the Lytro or Raytrix series, whose built-in microlens arrays can record both the spatial position and angle information of light rays. The camera sensor captures not only the light intensity on a two-dimensional plane but also the light ray incident angle data, thus forming multi-view light field raw data containing depth information. Each original RAW data point contains four-dimensional information composed of spatial coordinates (x, y) and angle coordinates (u, v), recording the complete path characteristics of light rays in the scene. Light field cameras are prone to being affected by environmental interference during the data acquisition process, so it is crucial to monitor the camera's operating environment parameters in real-time. The system continuously monitors key indicators such as temperature and power supply stability through built-in sensors. When the temperature exceeds 60°C or the power supply fluctuates by more than ±5%, the performance of the internal circuit of the camera deteriorates, resulting in a reduction of the signal-to-noise ratio of the acquired data to below 20 dB. At this time, the image noise significantly increases and a large amount of details are lost. After the monitoring program detects these abnormal parameters, it immediately triggers the standby camera switching mechanism, seamlessly transferring the data acquisition task to the standby light field camera with normal environmental parameters to ensure the continuity and reliability of data acquisition. During the switching process, the spatial positions and parameter configurations of the two cameras are kept synchronized to avoid perspective jumps.

[0045] The collected RAW format light field data is sent to the edge server based on Huawei Atlas 500 through a high-speed data transmission channel. Atlas 500 is an edge computing platform launched by Huawei, with powerful AI computing capabilities. Inside the edge server, there is an RK3588 edge computing chip, which is a system-on-chip integrating an octa-core CPU, a high-performance GPU, and an NPU, optimized for edge computing scenarios. After receiving the light field RAW data, RK3588 loads the data into memory and assigns computing tasks, improving the processing efficiency through multi-core parallel processing.

[0046] The primary task executed by the RK3588 chip is to denoise the RAW format light field data using the DnCNN deep convolutional neural network model. DnCNN is a residual learning framework specifically designed for image denoising, consisting of multiple convolutional layers, batch normalization layers, and ReLU activation functions. The model first decomposes the RAW data into different sub-aperture views, and then performs independent noise feature learning on each view. DnCNN learns the noise distribution characteristics through residual connections instead of directly learning the original image, enabling the model to more accurately identify and remove sensor noise, especially the photon shot noise and readout noise under low-light conditions in light field data. The NPU unit of RK3588 hardware-accelerates the DnCNN model, controlling the single-frame processing delay within 10 ms to meet the real-time processing requirements. After denoising, it enters the analysis and processing stage based on the epipolar constraint rule. The sub-images in the light field data should satisfy the epipolar geometry constraint, that is, the corresponding points should be distributed along the epipolar line under different perspectives. The processor calculates the corresponding positions of the feature points in each sub-image in the adjacent perspectives to form a disparity map. When the disparity value in a certain area deviates from the theoretical expectation by more than 2 pixels, it is marked as abnormal data. Such abnormalities are usually caused by factors such as sensor defects, light scattering, or object reflection, which will seriously affect the accuracy of subsequent super-resolution processing. The system automatically detects and filters these abnormal areas, retaining the valid data that meets geometric consistency.

[0047] Next, a deep confidence evaluation mechanism is used for quality screening. The light field data can calculate the scene depth through disparity analysis, but the depth estimation accuracy varies in different regions. The depth confidence algorithm is used to quantitatively evaluate the reliability of the depth estimation for each region, and the calculation factors include disparity consistency, texture richness, and gradient change, etc. When the depth confidence of a certain region is lower than the set threshold of 0.8, it indicates that the depth information in this region is unreliable, usually appearing in reflective surfaces, transparent objects, or textureless areas. The system marks and removes these low-confidence regions to ensure the reliability of the subsequent processed data.

[0048] Finally, metadata tags are attached to the processed light field data, including dynamic range information (recording the difference between the maximum and minimum pixel values) and contrast statistics (describing the degree of light and dark changes in the image). These information provide important references for subsequent super-resolution processing. The processed data is encapsulated in the HDF5 format. HDF5 is a file format used to store and organize large amounts of data, supporting multi-dimensional array storage and complex metadata attachment, and is particularly suitable for the structured preservation of light field data. During the encapsulation process, the light field sub-images, depth maps, quality markers, and metadata are uniformly organized to form structured HDF5-format preprocessed data. When collecting light fields for an indoor scene containing complex textures and objects at different depths, the Lytro Illum camera generates RAW-format light field data of 5MB in size, containing a 9×9 array of sub-images. During the processing, the edge server detects that the signal-to-noise ratio in some areas is only 16dB due to window reflection, far lower than the acceptable threshold, and immediately switches to the backup camera to re-collect the data. The DnCNN denoising process identifies and removes the sensor noise pattern, boosting the overall signal-to-noise ratio to 32dB. Epipolar constraint analysis finds that there is a perspective shift of about 3 pixels in the window edge area, and the system automatically marks and filters these areas. Depth confidence evaluation shows that the confidence levels of the smooth glass surface and the low-texture wall in the distance are 0.65 and 0.72 respectively, both lower than the 0.8 threshold, so these areas are marked as low-reliability areas. The finally generated HDF5-format preprocessed data contains clear multi-perspective information, accurate depth data, and complete quality assessment markers, providing high-quality input data for subsequent super-resolution processing based on the Ascend NPU. The entire preprocessing process makes full use of the parallel processing capabilities of the RK3588 edge computing chip, with a total processing time of only 85ms, effectively supporting real-time light field super-resolution computing.

[0049] In a specific embodiment, the process of executing step S102 may specifically include the following steps:

[0050] (1) Parse the preprocessed data in the HDF5 format into a sequence of sub-aperture images in an N×N array, and construct a multi-perspective representation of the light field;

[0051] (2) Perform global perspective relationship analysis on the sequence of sub-aperture images through the Vision Transformer branch to extract the spatial consistency features of the light field;

[0052] (3) Use the ResNet-18 variant CNN branch to perform local distortion localization on the sequence of sub-aperture images to identify high-frequency distortion regions;

[0053] (4) Combine the spatial consistency features with the information of the high-frequency distortion regions to generate a perspective correlation mapping diagram;

[0054] (5) Calculate the depth error compensation coefficient based on the perspective correlation mapping graph and adjust the light field depth data;

[0055] (6) Construct a distortion correction matrix according to the information of the spatial consistency feature and the high-frequency distortion region to eliminate the optical distortion in the light field data.

[0056] Specifically, the calibration process of the HDF5 format preprocessed data is a key link in the high-reliability light field super-resolution edge computing method. First, the HDF5 format preprocessed data needs to be parsed into a structured N×N array of sub-aperture image sequences. This parsing process extracts the light field information stored in different datasets through an HDF5 file parser, including the image data, disparity map, and quality markers of each perspective. Typical light field data is usually organized as a 5×5, 7×7, or 9×9 sub-aperture array, and each sub-aperture corresponds to an image of a specific perspective. The parsed data forms a complete four-dimensional light field representation, that is, two-dimensional spatial position plus two-dimensional angle information, constructing a complete multi-perspective representation of the light field and providing a basic data structure for subsequent processing.

[0057] Next, the global perspective relationship analysis of the sub-aperture image sequence is carried out through the Vision Transformer branch. Vision Transformer is a deep learning architecture based on the self-attention mechanism. It divides the input image into fixed-size blocks (patches) and captures the long-range dependencies between these blocks through self-attention layers. In light field processing, Vision Transformer regards each sub-aperture image as a sequence input, and the self-attention mechanism enables the model to consider the information of all perspectives simultaneously and effectively capture the global relationships between different perspectives. Through multi-head self-attention calculation, the model generates feature maps representing the spatial consistency of the light field. These feature maps contain the geometric consistency information that the light rays should maintain under different perspectives and are crucial for identifying systematic deviations caused by inaccurate calibration. The local distortion localization of the sub-aperture image sequence is carried out using the ResNet-18 variant CNN branch. ResNet-18 is a convolutional neural network based on residual connections, which solves the problem of gradient disappearance in the training of deep networks through skip connections. In this method, ResNet-18 is modified to adapt to the light field data structure and focuses on extracting local spatial features. This branch processes each sub-aperture image separately, and through stacked convolutional operations and feature map comparison, it identifies local regions with severe distortion, especially paying attention to positions prone to high-frequency distortion such as the lens edge region and the uneven illumination region. Through multi-scale feature extraction, the CNN branch can accurately locate problems such as optical distortion, chromatic aberration, and pixel offset, and output the accurate position map and distortion degree quantization value of the high-frequency distortion region.

[0058] Subsequently, by combining the spatial consistency features with the information of the high-frequency distortion regions, a view correlation mapping graph is generated. This process performs weighted merging of the global features output by the Vision Transformer and the local features output by the ResNet-18 variant through a feature fusion network. The fusion adopts an attention mechanism, assigning different weights to different regions according to the feature confidence, ensuring the balanced integration of global geometric consistency and local distortion information. The generated view correlation mapping graph is a multi-channel tensor, where each channel corresponds to the mapping relationship between a pair of views, precisely describing the spatial correspondence and distortion degree between different views in the light field, providing an accurate reference basis for subsequent calibration steps.

[0059] The depth information of the light field data is usually obtained by calculating the disparity between sub-aperture images. However, due to factors such as calibration errors and lens distortions, the initial depth information often has systematic biases. By analyzing the difference between the relative displacement in the view correlation mapping graph and the theoretically expected displacement, the depth correction values at each spatial position are calculated. This process first constructs the disparity-depth mapping relationship under the ideal light field model, and then generates spatially varying depth error compensation coefficients based on the deviation between the actual mapping and the ideal mapping. These coefficients are applied to the original depth map to adjust the depth value of each pixel, correcting the depth estimation errors caused by lens parameter deviations.

[0060] Finally, a distortion correction matrix is constructed based on the spatial consistency features and the information of the high-frequency distortion regions. Optical distortions mainly include radial distortion and tangential distortion, which can cause straight lines in the image to appear curved or deformed. By analyzing the distortion patterns in the view correlation mapping graph, the system constructs specific correction matrices for each sub-aperture image. These matrices are transformation matrices of 3×3 or higher dimensions, describing the mapping relationship from the distorted image to the ideal image. Applying these correction matrices to the original sub-aperture images can eliminate various optical distortions in the light field data, including barrel distortion, pincushion distortion, and complex mixed distortions, thereby obtaining a geometrically accurate representation of the light field.

[0061] Taking the security scenario in the park as an example, when the system processes a set of light field data containing near and far buildings and complex environments, it first parses the HDF5 format data into a 7×7 sub-aperture image array. Vision Transformer analysis finds that the perspective relationship in the central area is stable, but there are systematic offsets in the edge area; the ResNet-18 variant identifies a locally distorted area in the lower right corner due to light reflection. The perspective correlation mapping graph generated by combining the information of both shows that there is a systematic deviation of about 1.5 pixels in the parallax calculation at the building edge. Based on this discovery, the system calculates the depth error compensation coefficient that varies with space and successfully corrects the problem of underestimated depth of distant buildings. At the same time, the constructed distortion correction matrix effectively eliminates about 2% of the radial distortion caused by the lens, making the straight edges of the building maintain geometric straightness characteristics at all perspectives. The calibrated light field data enables accurate depth estimation and spatial reconstruction on edge computing devices.

[0062] In a specific embodiment, the process of executing step S103 may specifically include the following steps:

[0063] (1) Apply the depth error compensation coefficient and the distortion correction matrix to the calibrated light field data to generate a corrected light field image;

[0064] (2) Deploy a lightweight EfficientNet-B3 feature extraction layer on a Xilinx UltraScale+ FPGA, and parallelize the processing of the corrected light field image through logic units;

[0065] (3) Construct a weight-adjustable scoring function based on PSNR, SSIM, and edge sharpness metrics, and calculate the reliability score of the corrected light field image;

[0066] (4) Divide the corrected light field image into image blocks of n×n pixel size for frequency domain processing preparation;

[0067] (5) Perform a fast Fourier transform on the image blocks to convert the spatial domain light field data into a frequency domain representation;

[0068] (6) Perform interpolation upsampling processing on the frequency domain representation in the frequency domain, and then reconstruct through inverse Fourier transform to obtain a frequency domain super-resolution result.

[0069] Specifically, quality assessment and frequency-domain super-resolution processing of calibration data are performed with FPGA acceleration. First, the depth error compensation coefficient and distortion correction matrix calculated in the previous step are applied to the calibrated light field data. This process is achieved through matrix transformation operations. The corresponding distortion correction matrix is applied to each sub-aperture image for spatial remapping, and at the same time, the light field depth information is adjusted according to the depth error compensation coefficient. In specific operations, the distortion correction matrix acts on each pixel coordinate to calculate its new position in the corrected image, and the corrected pixel value is obtained through bilinear or bicubic interpolation; the depth error compensation coefficient directly acts on the values of the depth map to correct the depth estimation deviation. These operations generate a geometrically consistent and depth-accurate corrected light field image, laying a foundation for subsequent evaluation and frequency-domain processing. A lightweight EfficientNet-B3 feature extraction layer is deployed on the Xilinx UltraScale+ FPGA. EfficientNet is an efficient convolutional neural network architecture that balances network depth, width, and resolution through a compound scaling method. The B3 variant provides a good balance between accuracy and efficiency. Deploying it to the FPGA requires network pruning and quantization optimization to reduce the computational complexity while retaining the key feature extraction ability. The Xilinx UltraScale+ FPGA provides highly parallel computing capabilities through programmable logic units, which are particularly suitable for accelerating convolutional operations. In the FPGA implementation, the convolutional layer is mapped to the programmable logic unit, and the DSP slices are used to accelerate the multiply-accumulate operations, and the image data stream is processed through a pipelined architecture. This parallel processing method can reduce the feature extraction calculation time of EfficientNet-B3 from 15 ms on a general CPU to 5 ms, greatly improving the real-time processing ability.

[0070] Based on the extracted features, the system constructs a weight-adjustable scoring function to calculate the reliability score of the corrected light field image. This scoring function comprehensively considers three key quality indicators: Peak Signal-to-Noise Ratio (PSNR) measures the difference between the image and the noise-free reference; Structural Similarity Index (SSIM) evaluates the structural, luminance, and contrast similarities of the image; the edge sharpness index quantifies the sharpness of the image edges. The scoring function adopts a weighted summation form, i.e., Score = α×PSNR_norm + β×SSIM + γ×Edge_Sharpness, where α, β, and γ are adjustable weight coefficients that are adaptively adjusted according to different scenarios, and PSNR_norm is the normalized PSNR value. The system calculates the score for each sub-aperture image and comprehensively obtains the reliability score of the entire light field data, which usually ranges from 0 to 1, and a high value indicates reliable data quality.

[0071] In parallel with the quality assessment, the system prepares the corrected light field image for frequency domain super-resolution processing. First, the corrected light field image is segmented into image blocks of size n×n pixels, with a typical block size of 64×64 or 128×128 pixels. Block processing helps reduce memory requirements, improve computational efficiency, and allow adaptive processing for local characteristics. When segmenting, appropriate overlapping areas are considered (usually 10-20% of the block size) to avoid discontinuities when processing block boundaries. For each light field subview, the system generates a set of regularly arranged image blocks to form a two-dimensional block matrix structure. Each block retains its position information in the original image to facilitate subsequent reconstruction.

[0072] A fast Fourier transform (FFT) is performed on each image block to convert the spatial domain data into a frequency domain representation. FFT is an algorithm that efficiently calculates discrete Fourier transforms and can decompose spatial domain signals into sinusoidal components of different frequencies. FPGA hardware accelerates the FFT operation and efficiently calculates two-dimensional FFT through butterfly operation units. The transformed frequency domain data contains an amplitude spectrum and a phase spectrum. The amplitude spectrum represents the intensity of each frequency component, and the phase spectrum represents the corresponding phase information. The advantage of frequency domain analysis is that it is easier to distinguish the low-frequency structural information and high-frequency detail information of the image, which facilitates targeted processing. The frequency domain representation is interpolated and upsampled in the frequency domain. Frequency domain upsampling is achieved by inserting zero padding in the center area of ​​the spectrum, and then applying frequency domain constraints to ensure that the newly generated high-frequency components transition smoothly with the edges of the original spectrum. This method retains all the information of the original spectrum while reserving space for new high-frequency details. After processing, the frequency domain data is converted back to the spatial domain through an inverse fast Fourier transform (IFFT) to obtain a higher resolution image block. Finally, each image block is reassembled according to its original position, and the overlapping area is weighted averaged when necessary to obtain the frequency domain super-resolution result.

[0073] Taking the autonomous driving scenario as an example, when processing complex light field data containing vehicles and pedestrians near and far, the system first applies the depth error compensation coefficient to correct the problem of underestimated depth of distant objects, and applies the distortion correction matrix to eliminate about 3% of edge distortion to generate a geometrically accurate corrected light field image. EfficientNet-B3 deployed on the FPGA processes 128 sub-image blocks in parallel and extracts key features for quality assessment. The system calculates that the PSNR of the vehicle area is 38dB, the SSIM is 0.96, and the edge sharpness is 92. The scoring function with weights of 0.3, 0.4, and 0.3 is applied, and the final reliability score is 0.92, indicating that the data quality is good. For frequency domain processing, the corrected light field image is divided into 64×64 pixel image blocks. After performing FFT, the sampling points are doubled in the frequency domain. Clear license plate details and pedestrian facial features are reconstructed through IFFT, which provides high-quality frequency domain processing results and accurate reliability assessment for subsequent deep learning super-resolution based on Ascend NPU.

[0074] In a specific embodiment, the process of executing step S104 may specifically include the following steps:

[0075] (1) Transmit the reliability score and the frequency-domain super-resolution result to the Ascend NPU processing unit to construct a light field enhancement processing stream;

[0076] (2) Deploy the ESRGAN network with a residual dense block structure on the Ascend NPU and receive the light field enhancement processing stream as input data;

[0077] (3) Construct a weighted loss function according to the reliability score and incorporate the reliability score into the calculation processes of pixel loss, perceptual loss, and adversarial loss;

[0078] (4) Perform multi-layer feature extraction on the frequency-domain super-resolution result through residual dense blocks to retain the angular information of the light field image;

[0079] (5) Use the phase consistency loss constraint algorithm to process the multi-layer features and reduce texture artifacts in the light field multi-view data;

[0080] (6) Upsample and reconstruct the processed feature map to generate a deep learning super-resolution image containing high-frequency details and accurate depth information.

[0081] Specifically, transmit the reliability score and the frequency-domain super-resolution result obtained from the previous-stage FPGA processing to the Ascend NPU processing unit. The Ascend NPU is a neural network processor designed by Huawei, adopting the Da Vinci architecture, optimized for deep learning inference, and having high-efficient matrix calculation capabilities and low-power consumption characteristics. The data transmission process is realized through a high-speed PCIe bus or on-chip interconnect to ensure data integrity and transmission efficiency. The transmitted data includes the frequency-domain super-resolution result (including the preliminarily enhanced light field image) and the reliability score (a value in the 0-1 interval representing the data credibility), and these data together constitute the light field enhancement processing stream as the input for deep learning processing. After receiving the data, the Ascend NPU loads it into the on-chip memory and organizes the processing flow according to the predefined data flow graph.

[0082] Deploying the ESRGAN network with a residual dense block structure on the Ascend NPU is the second step. ESRGAN (Enhanced Super-Resolution Generative Adversarial Network) is an advanced super-resolution generative adversarial network that introduces residual dense blocks (RDBs) and related improvements on the basis of the original SRGAN. The deployment process requires first converting the pre-trained model into a format supported by the Ascend NPU. Through steps such as operator mapping, graph optimization, and quantization compression, the network can run efficiently on the NPU. The residual dense block is the core structure of the network, consisting of multiple convolutional layers. The output of each layer is connected to the outputs of all previous layers, forming a dense connection pattern, which enhances the feature extraction ability and gradient flow. The ESRGAN network architecture includes a feature extraction front end, a main body with multiple residual dense blocks, and a reconstruction back end. It receives the light field enhancement processing stream as input data. Each residual dense block is mapped to different computing units of the Ascend NPU to achieve parallel processing. A weighted loss function is constructed according to the reliability score. Traditional ESRGAN uses fixed weights to combine multiple loss functions, but the quality of light field data varies greatly and needs to be adaptively adjusted. This method incorporates the reliability score into the calculation of the loss function, combining pixel loss, perceptual loss, and adversarial loss. Pixel loss measures the direct difference between the reconstructed image and the target image; perceptual loss measures the perceptual similarity through the feature differences extracted by the pre-trained VGG network; adversarial loss comes from the discriminator's evaluation of the authenticity of the generated image. The reliability score serves as a dynamic weight adjustment factor. When the score value is high, it enhances the influence of perceptual loss and adversarial loss to promote the generation of texture details; when the score value is low, it reduces these two weights and relies more on the stability of pixel loss. This dynamic loss function enables the network to intelligently adjust the reconstruction strategy according to the reliability of the input data.

[0083] Performing multi-layer feature extraction on the frequency-domain super-resolution results through residual dense blocks is the fourth step. Each residual dense block contains multiple convolutional layers, and the output feature maps of each layer are connected to all subsequent layers, forming a dense connection network. The Ascend NPU efficiently executes convolutional operations through computing units. The feature extraction process is divided into multiple scales: first, shallow features are extracted to capture the basic structure and edges; then, middle-layer features are extracted to capture textures and local patterns; finally, deep-layer features are extracted to capture abstract semantic information. For light field data, feature extraction needs to particularly preserve the angular information. Therefore, the network maintains perspective consistency when processing multi-view sub-images. By designing special feature channels to store perspective relationship information, it ensures that the angular information is not lost during the reconstruction process.

[0084] Next, the phase consistency loss constraint algorithm is used to process the multi-layer features, reducing the texture artifacts in the light field multi-view data. Light field super-resolution is prone to generating inconsistent textures in different views, especially on complex surfaces. The phase consistency loss constraint algorithm ensures consistent texture generation across views by analyzing the phase spectrum relationship of features in different views. The algorithm first transforms the features of each view into the frequency domain, extracts the phase information, and then enforces the minimization of the phase difference in similar regions between different views. This constraint mechanism keeps the texture patterns coherent across different views, reducing the artifacts generated by independently processing each view, especially in complex scenarios such as surface reflections and transparent objects.

[0085] The processed feature maps are upsampled and reconstructed to generate a deep learning super-resolution image. The upsampling process uses a sub-pixel convolutional layer, providing finer detail reconstruction ability compared to traditional bilinear interpolation. The reconstruction module fuses multi-layer features, retains the original information through residual connections, and pays special attention to the retention of the light field depth information during the reconstruction process. The super-resolution result is ensured to be consistent with the original depth through the depth perception loss. Figure 1 The finally output deep learning super-resolution image retains the light field angle information while having enhanced spatial resolution, clear high-frequency details, and accurate depth information.

[0086] Taking the intelligent monitoring application as an example, when processing a light field scene containing multiple moving objects, the area with a reliability score of 0.88 (the main moving object) is given a higher perception loss weight by the Ascend NPU to generate more texture details; while the area with a score of 0.62 (shadow and reflective areas) relies more on pixel loss to maintain the stability of detail generation. The multi-layer features extracted by the residual dense block retain the detail changes of the monitored objects from different views. The phase consistency constraint ensures that the vehicle surface texture remains consistent across different views. The finally reconstructed 8K super-resolution image can clearly present the license plate number and the characteristics of the people 10 meters away, while accurately retaining the scene depth information to support subsequent 3D reconstruction and object recognition tasks.

[0087] In a specific embodiment, the process of executing step S105 may specifically include the following steps:

[0088] (1) Extract the high-frequency feature map from the frequency domain super-resolution result to retain the spatial detail information;

[0089] (2) Extract the texture feature map from the deep learning super-resolution image to obtain the details of the deep learning reconstruction;

[0090] (3) Input the high-frequency feature map and the texture feature map into the fusion network with a lightweight U-Net structure;

[0091] (4)In the encoder stage of the fusion network, key information in the high-frequency feature map and the texture feature map is identified through a channel attention mechanism;

[0092] (5)Based on the reliability score, the weight coefficient λ is dynamically adjusted. When the score is greater than 0.9, the weight of the deep learning branch is set to 0.8, and when the score is less than or equal to 0.6, the weight of the deep learning branch is reduced to 0.3;

[0093] (6)The fused features are reconstructed into an 8K resolution light field super-resolution image through the decoder of the fusion network.

[0094] Specifically, the high-frequency feature map is extracted from the frequency-domain super-resolution result. This process involves processing the frequency-domain super-resolution result through a feature extraction network, which consists of multiple convolutional layers and high-pass filters, specifically designed to capture the high-frequency components in the image. The high-frequency feature map focuses on retaining edge, texture, and detail information, which are crucial for image sharpness. The frequency-domain path processing can effectively preserve the spatial structure and geometric characteristics of the original image, especially showing advantages in clear boundary and regular texture regions. During the extraction process, while retaining high-frequency information, artifacts and noise that may be introduced during frequency-domain processing are filtered out, ensuring that the feature map has high information content and low noise.

[0095] The texture feature map is extracted from the deep learning super-resolution image. The super-resolution image generated by the ESRGAN network through the deep learning path often performs excellently in terms of texture details and perceptual quality. The texture feature extraction network focuses on capturing the detailed textures, material properties, and local structural features reconstructed by deep learning. These texture features often go beyond the details visible in the original low-resolution data and are generated through the prior knowledge learned by the deep network. The extraction process adopts a multi-scale feature extraction strategy to ensure the capture of texture information at different scales, from macroscopic structures to microscopic details, providing a comprehensive texture description. Compared with frequency-domain features, deep learning texture features perform better in complex and irregular texture regions, but their geometric accuracy may be inferior to that of frequency-domain features. The extracted high-frequency feature map and texture feature map are input into the fusion network with a lightweight U-Net structure. U-Net is a neural network with an encoder-decoder structure, originally designed for medical image segmentation, and is widely used in various image processing tasks due to its efficient feature extraction and reconstruction capabilities. The lightweight U-Net adapts to the resource limitations in the edge computing environment by reducing the network depth and the number of channels. The fusion network receives the two feature maps as inputs and unifies them into the network processing stream through a designed multi-channel input layer. To adapt to the processing characteristics of the Ascend NPU, the network structure is optimized to reduce memory requirements and improve computational efficiency while maintaining the effectiveness of feature fusion.

[0096] In the encoder stage of the fusion network, key information in the high-frequency feature map and texture feature map is identified through the channel attention mechanism. Channel attention is an adaptive feature weighting method that assigns weights according to the importance of different channel features. Specifically, the encoder first extracts hierarchical representations of the input features through downsampling convolutions, then calculates global statistical information (such as mean and standard deviation) for each feature channel, and learns channel weights based on this. Important channels (channels containing key information) obtain higher weights, while the weights of unimportant channels are reduced. This mechanism enables the network to automatically focus on the clear edges in high-frequency features and the rich details in texture features, selectively retain the advantageous parts of both features, and form a complementary and enhanced feature representation. Dynamically adjusting the weight coefficient λ based on the reliability score is a key adaptive mechanism. This mechanism uses the previously calculated reliability score to dynamically determine the contribution ratio of the frequency domain and deep learning paths. The specific rule is set as follows: when the score is greater than 0.9, the weight of the deep learning branch is set to 0.8, indicating that the data is highly reliable and more dependent on the rich textures generated by deep learning; when the score is less than or equal to 0.6, the weight of the deep learning branch is reduced to 0.3, indicating that the data reliability is low and more dependent on the fidelity of the frequency domain path. When the reliability score is between 0.6 and 0.9, the weight is calculated by linear interpolation. This dynamic adjustment strategy ensures that the system can intelligently switch processing strategies according to the quality of the input data and achieve the best balance between fidelity and perceptual quality.

[0097] Finally, the fused features are reconstructed into an 8K resolution light field super-resolution image through the decoder of the fusion network. The decoder contains a series of upsampling convolutional layers that gradually restore the spatial resolution of the features. Each upsampling layer is followed by a residual block to enhance the network's expressive power and alleviate the problem of gradient vanishing. During the decoding process, skip connections from the encoder provide detailed spatial information to supplement the details that may be lost during the upsampling process. The final output layer uses sub-pixel convolution to achieve precise detail reconstruction and generate an 8K resolution (7680×4320 pixels) light field super-resolution image. The imaging result combines the geometric accuracy of frequency domain processing and the rich textures of deep learning, while retaining the angular and depth information of the light field data.

[0098] Taking the product inspection application as an example, when processing the light field data containing a fine circuit board, the system first extracts the clear circuit line edges and geometries from the frequency domain results, and extracts the complex component textures and marking details from the deep learning results. The fusion network, through the channel attention mechanism, strengthens the frequency domain features in the circuit line area and the deep learning features in the component identification area. Since the reliability score of this batch of data is 0.85, the system sets the weight of the deep learning branch to 0.7 to appropriately balance the two-way features. The finally generated super-resolution imaging can clearly present the 0.1-mm-wide circuit lines and the tiny text markings on the components, while accurately retaining the three-dimensional depth information of the circuit board, enabling the inspection system to accurately identify defects and locate their spatial positions, greatly improving the efficiency and accuracy of product quality inspection.

[0099] In a specific embodiment, the process of executing step S106 may specifically include the following steps:

[0100] (1) Calculate the peak signal-to-noise ratio (PSNR) value of the super-resolution imaging of the 8K light field, and set 32 dB as the qualified quality threshold;

[0101] (2) Measure the structural similarity (SSIM) index of the super-resolution imaging of the 8K light field, and set 0.92 as the qualified quality threshold;

[0102] (3) Evaluate the edge sharpness index of the super-resolution imaging of the 8K light field, and set 90 as the qualified quality threshold;

[0103] (4) Comprehensively analyze the PSNR value, SSIM index, and edge sharpness index. If the set threshold is not reached, trigger the recalculation process, and call the calibrated data from the storage module to perform super-resolution processing again;

[0104] (5) Allocate Ascend NPU computing resources through the priority queue mechanism. In the normal mode, allocate 70% of the computing power for real-time super-resolution inference and 30% for model incremental training;

[0105] (6) Monitor the operating parameters of the FPGA and NPU. When the temperature exceeds 85 °C or the utilization rate exceeds 90%, start the hardware redundancy switching mechanism to migrate the computing task to the standby computing node to ensure the continuous and stable operation of the edge computing system.

[0106] Specifically, the sixth stage of the Ascend-based high-reliability light field super-resolution edge computing method is quality assessment and resource management to ensure the reliability of system output and the stability of operation. First, the peak signal-to-noise ratio (PSNR) of the generated super-resolution imaging is calculated. PSNR is a basic index for evaluating image quality. By calculating the mean square error (MSE) between the reconstructed image and the reference image and then converting it to the signal-to-noise ratio in logarithmic scale. In the light field super-resolution scenario, the reference image usually adopts a preset standard sample or the high-quality result processed in the previous stage. The PSNR calculation covers all sub-view images, and its average value is taken as the PSNR evaluation of the overall light field data. The system sets 32dB as the quality qualification threshold. This threshold is based on a large number of experimental verifications and can ensure that details are distinguishable and there is no obvious noise interference. A PSNR lower than the threshold usually means that key details are lost or too many artifacts are introduced during the reconstruction process.

[0107] Then, the structural similarity (SSIM) index of the super-resolution imaging is measured. SSIM is a quality assessment method that is more in line with human visual perception than PSNR. By comparing the brightness, contrast, and structural similarity of images, a score value between 0 and 1 is generated, where 1 means exactly the same. The SSIM calculation takes into account the local characteristics of the image. The SSIM value is calculated for each sub-view image, and then weighted and averaged to obtain the overall score. The system sets 0.92 as the SSIM quality qualification threshold. This value ensures that the reconstructed image is highly similar to the ideal result in structure and maintains the consistency of key visual features. An SSIM lower than the threshold usually indicates that the image structure is deformed or the texture features are unnatural, even if the PSNR may meet the standard.

[0108] In the third step, the edge sharpness index of the super-resolution imaging is evaluated. Edge sharpness is a special index for measuring image clarity and is calculated by measuring the gradient amplitude of the image edge transition region. The calculation process first uses an edge detection operator (such as Sobel or Canny) to locate the image edge, then measures the average intensity of the pixel gradient at the edge, and finally normalizes it to a score value between 0 and 100. The system sets 90 as the edge sharpness quality qualification threshold. This threshold ensures that the image edge is sharp and clear, and the detail contours are distinct. An edge sharpness lower than the threshold usually means that the image is blurred or lacks details, especially in high-frequency detail areas such as text and lines.

[0109] A quality decision-making mechanism is formed by comprehensively analyzing the PSNR value, SSIM metric, and edge sharpness metric. The system determines that the output is qualified only when all three metrics reach their respective thresholds simultaneously. If any one of the metrics fails to meet the standard, a recalculation process is triggered: the previously saved calibrated data (the processing results of the second stage) is called from the storage module, and the super-resolution processing is executed again. Different parameter configurations or processing algorithms may be used in the recalculation process. The system attempts recalculation up to 3 times, adjusting the relevant parameters each time to improve the success rate. If the standard is still not met, it is marked as a low-quality output and a detailed analysis report is recorded. This loop optimization mechanism ensures the high-quality standard of the system output and prevents unqualified results from flowing downstream to applications.

[0110] The task scheduling module allocates Ascend NPU computing resources through a priority queue mechanism. This mechanism classifies and queues all processing tasks according to their priorities, and real-time tasks usually obtain the highest priority. In the normal working mode, the system allocates 70% of the computing power of the Ascend NPU for real-time super-resolution inference tasks to ensure the timely response of the main functions; the other 30% of the computing power is used for model incremental training to continuously optimize the network parameters to adapt to new scenarios. Incremental training adopts an online learning strategy, using processed high-quality samples as training data to continuously improve the model performance without disturbing real-time tasks. When the pressure of real-time tasks increases, the system will dynamically adjust this allocation ratio to ensure that the core functions are guaranteed resources first.

[0111] Finally, the system continuously monitors the operating parameters of the FPGA and NPU and implements a hardware redundancy guarantee mechanism. The monitoring metrics include key parameters such as processor temperature, resource utilization rate, power supply status, and processing delay. When it is detected that the temperature exceeds 85°C or the resource utilization rate continuously exceeds 90%, the system determines that the hardware is in an overloaded state and activates the hardware redundancy switching mechanism. This mechanism migrates the current processing tasks and status data to a standby computing node, which can be a redundant FPGA / NPU module or other nodes in the edge server cluster. The task migration process is achieved through a status synchronization protocol to ensure that the processing continuity is not affected. The hardware redundancy mechanism is linked with dynamic resource scheduling to form a complete reliability guarantee framework for the edge computing system, enabling the system to maintain continuous and stable operation even in harsh environments or high-load conditions.

[0112] Taking medical imaging applications as an example, when the system processes the patient's light field fundus image, the generated 8K super-resolution imaging is evaluated for quality, obtaining a PSNR value of 33.8 dB (exceeding the 32 dB threshold), an SSIM value of 0.94 (exceeding the 0.92 threshold), and an edge sharpness value of 95 (exceeding the 90 threshold). All three indicators meet the standards, indicating that the reconstruction quality meets the requirements of medical diagnosis. At the same time, the monitoring shows that the NPU temperature reaches 82°C, approaching but not exceeding the 85°C threshold, and the system controls heat dissipation by dynamically adjusting the fan speed. During the peak diagnosis period, the system adjusts the NPU computing power allocation to 85% for real-time inference and 15% for model training, giving priority to ensuring clinical diagnosis needs. This example demonstrates the complete processing flow of the Ascend-based high-reliability light field super-resolution edge computing method in terms of quality assurance and resource management, ensuring high-quality and high-reliability light field super-resolution imaging services in resource-constrained edge environments.

[0113] The above describes the Ascend-based high-reliability light field super-resolution edge computing method in the embodiments of the present application. Next, the Ascend-based high-reliability light field super-resolution edge computing system in the embodiments of the present application will be described. Please refer to Figure 2 One embodiment of the Ascend-based high-reliability light field super-resolution edge computing system in the embodiments of the present application includes:

[0114] A denoising module, configured to collect RAW format light field data through a light field camera, perform DnCNN denoising processing on the light field data, and obtain preprocessed data in HDF5 format;

[0115] A calibration module, configured to perform calibration through a Transformer-CNN model according to the preprocessed data, and generate a depth error compensation coefficient and a distortion correction matrix;

[0116] A processing module, configured to use the calibrated light field data, perform FPGA-accelerated EfficientNet model and FFT processing, and output a reliability score and a frequency domain super-resolution result;

[0117] A computing module, configured to perform ESRGAN model calculation by an Ascend NPU based on the reliability score and the frequency domain super-resolution result to form a deep learning super-resolution image;

[0118] A weighting module, configured to weight the quality scores of the frequency domain super-resolution result and the deep learning super-resolution image through a U-Net fusion network to obtain a super-resolution imaging of an 8K light field;

[0119] An allocation module, configured to evaluate the super-resolution imaging of the 8K light field for PSNR and SSIM, and dynamically allocate Ascend NPU and FPGA resources by a task scheduling module to construct an edge computing redundancy mechanism.

[0120] Through the collaborative cooperation of the above-mentioned various components, the present invention acquires RAW-format light field data through a light field camera and performs DnCNN denoising processing to obtain structured preprocessed data in HDF5 format, ensuring the data quality and consistency for subsequent processing; calibrates the preprocessed data based on the Transformer-CNN hybrid architecture model to generate accurate depth error compensation coefficients and distortion correction matrices, effectively eliminating systematic errors in the light field imaging process; utilizes the FPGA-accelerated EfficientNet model and FFT processing technology to achieve hardware-level parallel processing capabilities, greatly improving the real-time performance of quality assessment and frequency domain processing. At the same time, the output reliability score and frequency domain super-resolution results provide double guarantees for subsequent processing; performs ESRGAN model calculations based on the Ascend NPU, making full use of the advantages of the NPU in deep learning inference to form a deep learning super-resolution image with rich texture details; innovatively fuses the frequency domain results and the deep learning image through a U-Net fusion network for quality score weighting, realizing the complementary enhancement of the advantages of different processing paths and solving the problem that a single algorithm is difficult to adapt to complex scenarios; adopts PSNR, SSIM evaluation and task scheduling module dynamic resource allocation strategies to construct a complete edge computing redundancy mechanism, significantly improving the fault tolerance and stability of the system. The present invention makes a number of innovations at the algorithm design level for the characteristics of light field data: the DnCNN denoising model is optimized for the noise distribution characteristics of light field images, effectively retaining angular information; the Transformer-CNN hybrid architecture gives full play to the complementary advantages of Transformer in global relationship modeling and CNN in local feature extraction; the dynamic loss function adjustment mechanism based on reliability score enables the ESRGAN model to adaptively balance fidelity and perceptual quality according to data quality; the channel attention mechanism of the U-Net fusion network realizes the optimal fusion of frequency domain processing and deep learning results. The combined action of these algorithm characteristics enables the system to automatically select the best processing strategy for light field data of different qualities in resource-constrained edge environments, significantly improving the quality consistency of super-resolution results. Overall, the present invention not only overcomes the problems of resource limitations and insufficient reliability in the existing edge environment technology, but also realizes the high reliability, high adaptability and high efficiency of the light field super-resolution system through the collaborative optimization of heterogeneous computing resources and the dynamic adjustment of intelligent processing strategies.

[0121] Referring to Figure 3 , in the embodiment of the present invention, a computer device is further provided. This computer device can be a server, and its internal structure can be as Figure 3As shown in the figure. The computer device includes a processor, a memory, a display screen, an input device, a network interface, and a database connected via a system bus. Among them, the processor of the computer design is used to provide computing and control capabilities. The memory of the computer device includes a non-volatile storage medium and an internal memory. The non-volatile storage medium stores an operating system, a computer program, and a database. The internal memory provides an environment for the operation of the operating system and the computer program in the non-volatile storage medium. The database of the computer device is used to store the corresponding data in this embodiment. The network interface of the computer device is used to communicate with an external terminal via a network connection. When the computer program is executed by the processor, the above method is implemented.

[0122] Those skilled in the art can understand that Figure 3 the structure shown in the figure is only a block diagram of a part of the structure related to the solution of the present invention, and does not constitute a limitation on the computer device to which the solution of the present invention is applied.

[0123] An embodiment of the present invention further provides a computer-readable storage medium, on which a computer program is stored. When the computer program is executed by a processor, the above method is implemented. It can be understood that the computer-readable storage medium in this embodiment can be a volatile readable storage medium or a non-volatile readable storage medium.

[0124] Those of ordinary skill in the art can understand that all or part of the processes in the methods of the above embodiments can be completed by instructing relevant hardware through a computer program. The computer program can be stored in a non-volatile computer-readable storage medium. When the computer program is executed, it can include the processes of the embodiments of the above methods. Among them, any reference to memory, storage, database, or other media provided by the present invention and used in the embodiments can include non-volatile and / or volatile memory. Non-volatile memory can include read-only memory (ROM), programmable ROM (PROM), electrically programmable ROM (EPROM), electrically erasable programmable ROM (EEPROM), or flash memory. Volatile memory can include random access memory (RAM) or an external cache memory. By way of illustration and not limitation, RAM is available in many forms, such as static RAM (SRAM), dynamic RAM (DRAM), synchronous DRAM (SDRAM), double data rate SDRAM (SSRSDRAM), enhanced SDRAM (ESDRAM), synchronous link (Synchlink) DRAM (SLDRAM), Rambus direct RAM (RDRAM), direct memory bus dynamic RAM (DRDRAM), and memory bus dynamic RAM, etc.

[0125] Those skilled in the art can clearly understand that for the convenience and brevity of description, the specific working processes of the above-described systems, systems, and units can refer to the corresponding processes in the foregoing method embodiments, and will not be described herein again.

[0126] If the integrated unit is implemented in the form of a software functional unit and sold or used as an independent product, it can be stored in a computer-readable storage medium. Based on such an understanding, the technical solution of the present invention, in essence, or the part that contributes to the prior art, or all or part of the technical solution, can be embodied in the form of a software product. The computer software product is stored in a storage medium and includes several instructions for causing a computer device (which may be a personal computer, a server, or a network device, etc.) to execute all or part of the steps of the methods described in various embodiments of the present invention. The foregoing storage medium includes: various media that can store program codes, such as USB flash drives, mobile hard disks, read-only memory (ROM), random access memory (RAM), magnetic disks, or optical discs.

[0127] As described above, the above embodiments are only used to illustrate the technical solutions of the present application, rather than to limit them; although the present application has been described in detail with reference to the foregoing embodiments, those of ordinary skill in the art should understand that they can still modify the technical solutions recorded in the foregoing embodiments, or perform equivalent replacements for some of the technical features; and these modifications or replacements do not cause the essence of the corresponding technical solutions to deviate from the spirit and scope of the technical solutions of the various embodiments of the present application.

Claims

1. A highly reliable light field super-resolution edge computing method based on Ascend, characterized in that: The Ascend-based high-reliability light field super-resolution edge computing method includes: Collecting RAW format light field data through a light field camera, performing DnCNN denoising on the light field data, and obtaining pre-processed data in HDF5 format; According to the preprocessed data, calibration is performed through the Transformer-CNN model to generate a depth error compensation coefficient and a distortion correction matrix; Using the calibrated light field data, the FPGA-accelerated EfficientNet model and FFT processing are used to output reliability scores and frequency domain super-resolution results. Based on the reliability score and frequency domain super-resolution results, the Ascend NPU performs ESRGAN model calculations to form a deep learning super-resolution image, including: Transmitting the reliability score and the frequency domain super-resolution result to the Ascend NPU processing unit to construct a light field enhancement processing flow; Deploy an ESRGAN network with a residual dense block structure on the Ascend NPU and receive the light field enhancement processing stream as input data; Constructing a weighted loss function according to the reliability score, and incorporating the reliability score into the calculation process of pixel loss, perceptual loss, and adversarial loss; Performing multi-layer feature extraction on the frequency domain super-resolution result through a residual dense block to retain the angle information of the light field image; Processing the multi-layer features using a phase consistency loss constraint algorithm to reduce texture artifacts in light field multi-view data; Upsample and reconstruct the processed feature map to generate a deep learning super-resolution image containing high-frequency details and accurate depth information; The frequency domain super-resolution result and the deep learning super-resolution image are weighted by quality scoring through the U-Net fusion network to obtain 8K light field super-resolution imaging, including: Extracting high-frequency feature maps from the frequency domain super-resolution results to retain spatial detail information; Extracting a texture feature map from the deep learning super-resolution image to obtain details of deep learning reconstruction; Inputting the high-frequency feature map and the texture feature map into a fusion network with a lightweight U-Net structure; In the encoder stage of the fusion network, key information in the high-frequency feature map and the texture feature map is identified through a channel attention mechanism; Dynamically adjust the weight coefficient λ based on the reliability score, set the deep learning branch weight to 0.8 when the score is greater than 0.9, and reduce the deep learning branch weight to 0.3 when the score is less than or equal to 0.6; Reconstructing the fused features into 8K resolution light field super-resolution imaging through a decoder of the fusion network; The PSNR and SSIM are evaluated for the super-resolution imaging of the 8K light field, and the task scheduling module dynamically allocates Ascend NPU and FPGA resources to build an edge computing redundancy mechanism.

2. The Ascend-based high-reliability light field super-resolution edge computing method according to claim 1, characterized in that: The light field camera is used to collect RAW format light field data, and the light field data is subjected to DnCNN denoising processing to obtain pre-processed data in HDF5 format, including: Collect 4K resolution RAW format light field data, and the sensor records the light direction and intensity information to form multi-view light field raw data; Monitor the operating environment parameters of the light field camera. When the temperature exceeds 60°C or the power supply fluctuation exceeds ±5%, resulting in a signal-to-noise ratio lower than 20dB, switch the acquisition task to the backup light field camera. The RAW format light field data is transmitted to the edge server based on Huawei Atlas 500, and the RK3588 edge computing chip performs the computing task; Applying the DnCNN deep convolutional neural network model to the RAW format light field data to eliminate sensor noise and control the single frame processing delay to less than 10ms; Analyze the sub-images in the light field data according to the epipolar constraint rule, and filter out abnormal data areas with a viewing angle deviation exceeding 2 pixels; The light field data is quality screened according to a depth confidence threshold of 0.8, and low-quality areas are removed to ensure data reliability; Dynamic range and contrast statistics metadata tags are added to the light field data, and the data is packaged into HDF5 format pre-processed data.

3. The Ascend-based high-reliability light field super-resolution edge computing method according to claim 1, characterized in that: The method of calibrating the preprocessed data through the Transformer-CNN model to generate a depth error compensation coefficient and a distortion correction matrix includes: Parsing the HDF5 format preprocessed data into a sub-aperture image sequence of an N×N array, and constructing a light field multi-view representation; Performing global perspective relationship analysis on the sub-aperture image sequence through the Vision Transformer branch to extract light field spatial consistency features; Using a ResNet-18 variant CNN branch to locate local distortion of the sub-aperture image sequence and identify high-frequency distortion areas; Combining the spatial consistency feature with information of the high-frequency distortion area to generate a viewing angle correlation map; Calculating a depth error compensation coefficient based on the perspective association map, and adjusting light field depth data; A distortion correction matrix is ​​constructed according to the spatial consistency feature and information of the high-frequency distortion area to eliminate optical distortion in the light field data.

4. The Ascend-based high-reliability light field super-resolution edge computing method according to claim 1, characterized in that: The calibrated light field data is used to accelerate the EfficientNet model and FFT processing through FPGA to output reliability scores and frequency domain super-resolution results, including: Applying the depth error compensation coefficient and the distortion correction matrix to the calibrated light field data to generate a corrected light field image; Deploy a lightweight EfficientNet-B3 feature extraction layer on a Xilinx UltraScale+ FPGA to parallelize the processing of the corrected light field image through logic units; Constructing a weight-adjustable scoring function based on PSNR, SSIM and edge sharpness indicators to calculate the reliability score of the corrected light field image; Segmenting the corrected light field image into image blocks of size n×n pixels in preparation for frequency domain processing; Performing a fast Fourier transform on the image block to convert the spatial domain light field data into a frequency domain representation; The frequency domain representation is subjected to interpolation up-sampling processing in the frequency domain, and then reconstructed through inverse Fourier transform to obtain a frequency domain super-resolution result.

5. The Ascend-based high-reliability light field super-resolution edge computing method according to claim 1, characterized in that: The PSNR and SSIM evaluations are performed on the super-resolution imaging of the 8K light field, and the Ascend NPU and FPGA resources are dynamically allocated by the task scheduling module to build an edge computing redundancy mechanism, including: Calculate the peak signal-to-noise ratio (PSNR) value of the super-resolution imaging of the 8K light field, and set 32 ​​dB as the quality threshold; The structural similarity SSIM index of the super-resolution imaging of the 8K light field is measured, and 0.92 is set as the quality acceptance threshold; Evaluate the edge sharpness index of the super-resolution imaging of the 8K light field, and set 90 as the quality acceptance threshold; The PSNR value, SSIM index and edge sharpness index are comprehensively analyzed. If the set threshold is not reached, a recalculation process is triggered, and the calibrated data is called from the storage module to perform super-resolution processing again; Ascend NPU computing resources are allocated through a priority queue mechanism. In normal mode, 70% of computing power is allocated for real-time super-resolution inference and 30% for incremental model training. Monitor the operating parameters of FPGA and NPU. When the temperature exceeds 85°C or the utilization rate exceeds 90%, start the hardware redundancy switching mechanism to migrate computing tasks to the backup computing nodes to ensure the continuous and stable operation of the edge computing system.

6. A high-reliability light field super-resolution edge computing system based on Ascend, used to implement the high-reliability light field super-resolution edge computing method based on Ascend as described in any one of claims 1-5, characterized in that: The Ascend-based high-reliability light field super-resolution edge computing system includes: A denoising module is used to collect RAW format light field data through a light field camera, perform DnCNN denoising on the light field data, and obtain pre-processed data in HDF5 format; A calibration module, used to calibrate the preprocessed data through a Transformer-CNN model to generate a depth error compensation coefficient and a distortion correction matrix; The processing module is used to use the calibrated light field data, accelerate the EfficientNet model and FFT processing through FPGA, and output the reliability score and frequency domain super-resolution results; A computing module, configured to execute ESRGAN model calculation by the Ascend NPU based on the reliability score and the frequency domain super-resolution result to form a deep learning super-resolution image; A weighting module is used to weight the frequency domain super-resolution result and the deep learning super-resolution image through a U-Net fusion network to obtain 8K light field super-resolution imaging; The allocation module is used to perform PSNR and SSIM evaluation on the super-resolution imaging of the 8K light field. The task scheduling module dynamically allocates Ascend NPU and FPGA resources to build an edge computing redundancy mechanism.

7. A computer device, characterized in that: The method comprises a memory and a processor, wherein the memory stores a computer program that can be run on the processor, and when the processor executes the computer program, the high-reliability light field super-resolution edge computing method based on Ascend described in any one of claims 1 to 5 is implemented.

8. A computer-readable storage medium, characterized in that: A computer program is stored thereon, and when the computer program is executed by a processor, the processor executes the Ascend-based high-reliability light field super-resolution edge computing method as described in any one of claims 1 to 5.

Citation Information

Patent Citations

  • Light field multi-view image super-resolution reconstruction method based on deep learning

    CN112750076A

  • Image super-resolution reconstruction method based on improved ESRGAN

    CN114463176A