Image collector, intelligent light field array camera system and display equipment
By adopting a regular hexagonal array structure of primary and secondary collectors in the light field camera, combined with the dynamic processing module of the FPGA and Ascend 310 NPU processor, high-resolution multi-viewpoint light field acquisition and efficient naked-eye 3D display are achieved. This solves the problems of resolution loss, uneven sampling and high power consumption of existing light field cameras, and realizes multi-scene adaptive and low-latency light field imaging and display.
Patent Information
- Application Number
- CN202511074552.1
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-08-01
- Publication Date
- 2025-09-12
AI Technical Summary
Existing light field cameras have problems such as resolution loss, uneven sampling, rigid algorithms, high system latency, high power consumption and poor display effects. They are unable to meet the needs of high-resolution multi-viewpoint light field acquisition, multi-scene adaptive hardware and algorithm switching, and multi-person naked-eye 3D display.
A regular hexagonal array structure consisting of a primary and secondary collector is adopted, combined with the FPGA system's dynamic processing module and the Ascend 310 NPU processor to achieve multi-viewpoint light field acquisition and naked-eye 3D display. ISP pre-processing, light field sub-image synthesis, and scene classification are performed through the FPGA system's reconfigurable partitioning. A lightweight neural network is combined for real-time scene label recognition and hardware start-stop, driving the display module for efficient display.
It realizes high-resolution multi-viewpoint light field acquisition, real-time adaptive algorithm switching and high-fidelity naked-eye 3D display, improving the system's flexibility, real-time performance and energy efficiency, adapting to multiple scenario requirements, reducing power consumption and ensuring system stability.
Smart Images

Figure CN120640146A_ABST
Abstract
Description
Technical Field
[0001] The present disclosure relates to the field of intelligent imaging and embedded AI collaborative technology, and in particular to an image collector, an intelligent light field array camera system, and a display device. Background Art
[0002] With the application of technologies such as virtual reality, unmanned driving, and intelligent monitoring, the demand for three-dimensional information acquisition and interactive experience is also increasing. Light field imaging technology has attracted much attention because it can simultaneously record the spatial position and direction information of light.
[0003] Existing light field cameras are mainly divided into two categories: single-camera solutions based on microlens arrays and multi-camera solutions based on multi-camera arrays. Typical representatives of the former include products such as Lytro and Raytrix. They integrate a fixed-format microlens array in front of the traditional photosensitive chip. By collecting light field data on a single physical camera, they realize back-end functions such as variable focus, depth of field adjustment, and 3D reconstruction. However, due to the inherent "one-to-many" mapping relationship between microlenses and pixels, the final effective resolution is usually only one-tenth of the bare metal pixel or even lower, which makes it difficult to meet the needs of detail capture in high-resolution imaging scenarios. In addition, the fixed microlens size and focal length make it impossible to adjust imaging parameters online when facing low light, high dynamic range or motion scenes, limiting its adaptability in complex environments.
[0004] Another approach involves arranging standard camera modules at equal intervals to form a multi-camera array, such as the common 3×3, 5×5, or even larger rectangular grid arrays. This architecture can significantly improve the spatial resolution and disparity density of light field data by leveraging the full pixel output of independent lenses and image sensors, thereby improving depth estimation accuracy and rendering quality. However, the rectangular grid itself has limitations in sampling geometry: the disparity baseline in the center is compact and uniform, while the disparity baseline in the edge areas is elongated and non-uniform, resulting in a sharp increase in the complexity of 3D reconstruction and sub-viewpoint interpolation algorithms.
[0005] Furthermore, multi-camera arrays typically use fixed FPGAs or ASICs to implement a predefined ISP pipeline during hardware deployment. These processing algorithms are fixed during the design phase and cannot be dynamically loaded or updated during operation. While some FPGA systems support partial reconfiguration (DPR), these are primarily used for simple function switching, making them incapable of adapting to ever-changing external environments and diverse imaging requirements.
[0006] The back-end processing and display of light field data currently relies heavily on GPUs or high-performance CPUs for offline or near-real-time rendering of large-scale captured data streams. This struggles to meet the demanding requirements of embedded or mobile edge scenarios, both in terms of rendering latency and power consumption. Furthermore, existing solutions for naked-eye 3D display of light field content rely on head tracking systems or glasses-free screens with limited viewpoint output, failing to achieve a truly immersive experience with multiple viewers, high refresh rates, and wide viewing angle coverage.
[0007] Given the problems of resolution loss, uneven sampling, algorithm rigidity, high system latency, high power consumption and poor display effects in the above-mentioned solutions, there is an urgent need for an overall solution that can provide high-resolution multi-viewpoint light field acquisition, realize multi-scene adaptive hardware and algorithm switching, and have low latency, low power consumption and multi-person naked-eye 3D display capabilities.
[0008] Based on the above problems, the present disclosure provides an image collector, an intelligent light field array camera system and a display device. Summary of the Invention
[0009] The present disclosure provides an image collector, an intelligent light field array camera system, and a display device to solve the problems of existing light field cameras such as resolution loss, uneven sampling, algorithm rigidity, high system latency, high power consumption, and poor display quality.
[0010] According to a first aspect of the present disclosure, an image collector is provided, which is applied to an intelligent light field array camera system, comprising: 1 main collector and several secondary collectors; The main collector is used to capture detail images, and the secondary collector is used to capture close-range parallax images and overall field of view images; N1 secondary collectors are distributed at equal intervals around the main collector to form an inner ring of a regular hexagon; N2 secondary collectors are distributed at equal intervals outside the inner ring to form an outer ring of a regular hexagon; Where N1 and N2 are both integers, and N1 ≥ 4, N2 ≤ 32, N2 = 2N 1。
[0011] According to a second aspect of the present disclosure, an intelligent light field array camera system is provided, the system comprising the following modules: An image acquisition module, comprising the image collector as described in the first aspect, for acquiring original images, including detail images, close-range parallax images, and overall field of view images; The dynamic processing module, built on an FPGA system, includes three reconfigurable partitions, which respectively deploy ISP preprocessing, light field sub-image synthesis, and scene classification functions, and are used to output sub-viewpoint feature streams and scene labels of the original image; A management module, configured to generate a control signal and an acquisition module start / stop mask according to the scene tag; the control signal is used to drive the corresponding DPR logic function in the FPGA system, and the start / stop mask is used to control the start / stop of the primary and secondary acquisition modules of the acquisition module; The reconstruction module is used to input the sub-viewpoint feature stream into the dual-branch network for multi-viewpoint depth estimation and multi-focus image synthesis, and output a 40×40 viewpoint parallel image array.
[0012] As described above and any possible implementation method, a further implementation method is provided, in which the ISP preprocessing function simultaneously completes denoising, white balance, geometric correction and HDR fusion through a row-level deep pipeline.
[0013] As described above and any possible implementation method, an implementation method is further provided, in which the light field sub-image synthesis function crops the original images collected by the main collector and the secondary collector after ISP preprocessing into sub-viewpoint images, and aligns them according to the calibrated disparity baseline to obtain the sub-viewpoint feature stream.
[0014] According to the above aspects and any possible implementation, there is further provided an implementation, further comprising: The sensor module includes: a temperature sensor for collecting the brightness of the environment; a brightness sensor for collecting the brightness of the environment; an IMU sensor for collecting acceleration and angular velocity data of the intelligent light field array camera system; The original image collected by the main collector after ISP preprocessing, the temperature and brightness of the environment, and the acceleration and angular velocity data are input into the MobileNetV2 lightweight neural network to realize the scene classification function, and the real-time output includes scene labels including "daytime", "low light", "high speed", "long distance" or "panoramic".
[0015] According to the above aspects and any possible implementation, a further implementation is provided, which generates a control signal and an acquisition module start / stop mask according to the scene tag, including: In the "daylight" mode, enable the main collector and the secondary collector in the acquisition module and load daylight_standard.bit; In "low light" mode, only enable the inner loop sub-collector in the acquisition module and load low_light_fusion.bit; Enable the main collector and inner loop secondary collector in the acquisition module in "high speed" mode and load motion_stab.bit; Enable the main collector and outer ring secondary collector to collect in "long distance" mode and load super_res_recon.bit; In the "Panorama" mode, enable the main and secondary collectors in the acquisition module again and load panorama_render.bit.
[0016] According to the above aspects and any possible implementation, an implementation is further provided, wherein the start / stop mask controls the switch matrix of the acquisition module through the GPIO interface, and the switch matrix enables or disconnects the corresponding sub-array channels in the acquisition module according to the start / stop mask; the sub-array channels control the start and stop of the primary collector and the secondary collector.
[0017] According to the above aspects and any possible implementation, an implementation is further provided, in which, in the reconstruction module, the dual-branch network runs on the Ascend 310 NPU processor, and the Ascend 310 NPU processor is connected to the FPGA system via a PCIe interface.
[0018] According to a third aspect of the present disclosure, there is provided a display device, comprising: The intelligent light field array camera system according to the second aspect is used to output a 40×40 viewpoint parallel image array; A display module is used to visualize the 40×40 viewpoint parallel array to achieve naked-eye 3D display; the display module includes an LCD panel and a 40×40 microlens array, each microlens of the microlens array is aligned with a pixel area on the LCD panel.
[0019] According to the above aspects and any possible implementation, an implementation is further provided for visualizing the 40×40 viewpoint parallel image array output by the reconstruction module in the intelligent light field array camera system, including: The 40×40 viewpoint parallel array is mapped through a parallel interpolation and parallax allocation algorithm, and is visually displayed through a display; wherein the display automatically corrects the parameters of the microlens array parameters.
[0020] The beneficial effects of the present disclosure are: The present disclosure provides an intelligent array camera and an image collector, which is divided into a primary collector and a secondary collector, and realizes multi-view light field acquisition in the center and two ring areas. The FPGA system in the front-end dynamic processing module is divided into three reconfigurable partitions based on the Xilinx Zynq UltraScale+ platform: ISP pre-processing, light field sub-image synthesis, and scene division. Denoising, white balancing, geometric correction, HDR fusion, and sub-view image extraction are completed in parallel through a row-level deep pipeline, with a processing throughput of up to 5G pixels / second. The mid-end reconstruction module is connected to the FPGA system via a PCIe 3.0×4 interface, runs a lightweight dual-branch network after quantization and channel pruning, performs multi-view depth estimation and multi-focus image synthesis, respectively, expands the sub-view feature stream into a 40×40 viewpoint parallel image array, generates a depth map and a focus level map, and outputs them in less than or equal to 20 ms. The present disclosure also includes a back-end display module that uses an 8K LCD panel and a microlens array, and employs a parallel interpolation algorithm to drive a 30-60 Hz naked-eye multi-viewer 3D display. The FPGA system scene classification function is implemented based on a quantized MobileNetV2 lightweight neural network, which recognizes five scene labels in real time: "daytime, low light, high speed, long distance, and panoramic view." This drives the DPR logic and the start and stop of the sub-array array, achieving closed-loop adaptation of the hardware and algorithm. Through modular design and dynamic collaboration, this paper achieves the organic integration of high-resolution multi-viewpoint light field acquisition, real-time adaptive algorithm switching and high-fidelity naked-eye 3D display, significantly improving the system's flexibility, real-time performance and energy efficiency. This paper organically combines high-density intelligent camera array acquisition, reconfigurable hardware acceleration with AI chip reasoning and multi-viewpoint high-refresh naked-eye display to achieve a multi-scene adaptive, low-latency, high-energy-efficiency and high-reliability light field imaging and display system. This paper adopts copper heat pipes, honeycomb aluminum alloy shell, PID algorithm intelligent air cooling, and cooperates with the clock gating of the FPGA system and the power consumption mode switching of the AI chip to achieve low power consumption and high stability operation in a wide temperature range of -20℃ to 70℃.
[0021] It should be understood that the contents described in the Summary of the Invention section are not intended to limit the key or important features of the embodiments of the present disclosure, nor are they intended to limit the scope of the present disclosure. Other features of the present disclosure will become readily understood through the following description. BRIEF DESCRIPTION OF THE DRAWINGS
[0022] The above and other features, advantages and aspects of the embodiments of the present disclosure will become more apparent with reference to the following detailed description in conjunction with the accompanying drawings. The accompanying drawings are provided for a better understanding of the present disclosure and do not constitute a limitation of the present disclosure. In the accompanying drawings, the same or similar reference numerals represent the same or similar elements, among which: Figure 1 The figure shows a schematic structural diagram of the image collector provided by the present disclosure; Figure 2 The figure shows a schematic structural diagram of the intelligent light field array camera system provided by the present disclosure; Figure 3 The DPR logic control flow chart provided by the present disclosure is shown. DETAILED DESCRIPTION
[0023] To make the purpose, technical solutions, and advantages of the embodiments of the present disclosure more clear, the technical solutions in the embodiments of the present disclosure will be clearly and completely described below in conjunction with the accompanying drawings. Obviously, the described embodiments are only part of the embodiments of the present disclosure, not all of the embodiments. Based on the embodiments of the present disclosure, all other embodiments obtained by ordinary technicians in this field without making any creative efforts are within the scope of protection of the present disclosure.
[0024] In this document, the term "and / or" simply describes a relationship between related objects, indicating that three possible relationships exist. For example, "A and / or B" can represent: A exists alone, A and B exist simultaneously, or B exists alone. Furthermore, the character " / " in this document generally indicates that the related objects are in an "or" relationship.
[0025] The present disclosure provides an image collector, see Figure 1 ,include: 1 main collector and several secondary collectors, the main collector is used to capture detail images, and the secondary collectors are used to capture close-range parallax images and overall field of view images; Figure 1 The large circle in the center is the primary collector, and the small circles around it are the secondary collectors.
[0026] N1 secondary collectors are distributed around the main collector at equal intervals to form an inner ring of a regular hexagon; N2 secondary collectors are distributed outside the inner ring at equal intervals to form an outer ring of a regular hexagon; wherein N1 and N2 are both integers, and N1≥4, N2≤32, N2=2N 1。
[0027] In the present disclosure, the placement angles of the inner and outer regular hexagons can be staggered, and the sides do not need to be set parallel. Figure 1 This is just an example. The present disclosure does not limit the total number of sub-collectors. In practical applications, they can be flexibly configured according to the required spatial resolution, parallax baseline, system cost, bandwidth, depth accuracy, etc. to meet different parallax density and field of view coverage requirements.
[0028] By limiting the relationship between the number of sub-collectors in the inner and outer rings, the parallax density of the two-level baselines formed by the main collector and the two levels of surrounding sub-collectors can be increased. A consistent light field effect can be obtained by simply ensuring that the sub-collectors in the inner and outer rings are equally spaced and proportional in number.
[0029] The following descriptions are based on the example of N1=6 and N2=12. Figure 1 The entire array camera adopts a hybrid inner and outer ring structure, with the primary collector located in the center. The primary collector is a high-resolution camera responsible for capturing details of key viewpoints. The secondary collectors are also cameras. The six sub-collectors in the inner ring are evenly spaced in a regular hexagon to provide high-density close-range parallax samples. The 12 sub-collectors in the outer ring are also evenly spaced in a regular hexagon to expand the overall field of view and enhance long-range depth measurement. The disclosed structure ensures the accuracy of depth estimation while avoiding the problem of uneven parallax at the edge of traditional rectangular grids.
[0030] Compared with the square design, there are problems: the parallax baseline in the central area is short and uniform, but the edge baseline is elongated and inconsistent in direction, which leads to a decrease in edge accuracy during depth reconstruction; the redundancy in the diagonal direction is high, that is, the utilization rate of the corner sub-collectors is low; the regular hexagonal design in the present disclosure makes the parallax baseline length of the edge area closer to the central area, avoiding the edge distortion of the square array.
[0031] Compared with the problem of circular design: the sub-collectors cannot be arranged closely, and there is wasted gaps; the regular hexagonal design in the present disclosure can accommodate the maximum number of sub-collectors in the smallest area, and can reduce cable crossing and reduce interference.
[0032] It should be noted that the radius of the inner and outer rings is set based on actual conditions. In a preferred embodiment, while ensuring the available space and physical dimensions of the camera module, the ratio of the radius of the outer and inner hexagons is set to 2. This enables centimeter-level depth resolution and sufficient field of view coverage in both close-range (e.g., tens of centimeters) and long-range (e.g., several meters) scenarios.
[0033] The design of the disclosed array camera provides high-density near-range parallax in the inner ring and expands the long-range depth in the outer ring. While ensuring field of view coverage, the parallax sampling is evenly distributed in space, significantly improving the accuracy and robustness of 3D reconstruction.
[0034] The present disclosure also provides an intelligent light field array camera system 200, see Figure 2 , including the following modules: The image acquisition module 201 includes the aforementioned image collector, and is used to acquire original images, including detail images, close-range parallax images, and overall field of view images.
[0035] The dynamic processing module 202 is built based on the FPGA system and includes three reconfigurable partitions, which respectively deploy ISP preprocessing, light field sub-image synthesis and scene classification functions, and are used to output the sub-viewpoint feature stream and scene label of the original image.
[0036] In one embodiment, the dynamic processing module 202 is constructed based on the Xilinx Zynq UltraScale+ platform, which includes an FPGA system and a processor system. The FPGA system is programmable logic. Apparently, the present disclosure includes 19 neutron viewpoint feature streams, which come from the original images collected by the primary and secondary collectors, respectively.
[0037] (1) The ISP pre-processing function simultaneously completes denoising, white balancing, geometric correction and HDR fusion through a row-level deep pipeline; pixel-level processing refers to the processing operation for each pixel in the image, and the color, brightness, depth and other attributes of each pixel are calculated individually or based on its neighborhood to achieve a specific effect. For example, when denoising, the noise can be judged and removed based on the characteristics of each pixel and its surrounding pixels. Row-level processing refers to processing the image by row, applying a specific algorithm to remove the noise in the row based on the distribution characteristics and adjacent relationships of the pixels in the row, and then processing each row in turn to finally complete the denoising of the entire image.
[0038] (2) The light field sub-image synthesis function crops the pre-processed original images collected by the main collector and the secondary collector into sub-viewpoint images, and aligns them according to the calibrated parallax baseline to obtain the sub-viewpoint feature stream; the pre-processed original image may include invalid edge pixels, background beyond the scene range, etc. Cropping is to intercept the corresponding effective field of view area in the images from different collectors, generate the sub-viewpoint feature stream, and calibrate the baseline according to the position difference caused by different viewing angles.
[0039] (3) The original image collected by the main collector after ISP preprocessing, the temperature and brightness of the environment, and the acceleration and angular velocity data are input into the MobileNetV2 lightweight neural network to realize the scene classification function, and the scene labels including "daytime", "low light", "high speed", "long distance" or "panoramic" are output in real time.
[0040] In a specific embodiment, a sensor module is further included, including: a temperature sensor for collecting the brightness of the environment; a brightness sensor for collecting the brightness of the environment; an IMU sensor (i.e., an inertial measurement unit) for collecting acceleration and angular velocity data of the intelligent light field array camera system (to reflect the motion state of the intelligent array camera or the object being photographed); the sensor module can be arranged on a bracket provided outside the intelligent light field array camera system, adjacent to the image acquisition module.
[0041] Among them, the sensor component passes the standard I2 C or SPI bus interface is connected to the scenario classification reconfigurable partition of the FPGA system.
[0042] During each classification cycle (e.g., 100 milliseconds), the dynamic processing module 202 simultaneously pulls the raw image from the main image collector, along with the ambient temperature and brightness data, acceleration, and angular velocity data. These data are used as inputs for the MobileNetV2 lightweight neural network to achieve reliable scene classification. In the subsequent management module, the classification results drive the DPR logic switching algorithm bitstream and the activation and deactivation of the corresponding camera subarray channels.
[0043] During the switching process, other partitions that have not been reconfigured continue to maintain their original functions to ensure continuous operation of the system. That is, the light field sub-image synthesis function and the scene classification function are completely parallel, and the two do not interfere with each other in different partitions. In addition, the DPR logic switching interval triggered by the scene classification function is less than 10 ms and only acts on the corresponding partition, without interrupting the sub-image synthesis or ISP preprocessing process, ensuring high throughput, low latency and smooth switching of the end-to-end pipeline.
[0044] In one specific embodiment, three reconfigurable partitions work together on the AXI bus, and the DPR logic completes bitstream switching within 10ms through the ICAP interface without affecting the continuous and stable operation of the other partitions. The AXI bus and the ICAP interface have a connection and control relationship. The AXI bus can serve as a bridge, allowing the Xilinx Zynq UltraScale+ processor system to access the FPGA system's configuration memory, namely the BRAM memory, through the ICAP interface, thereby realizing dynamic modification of the circuit structure and function of the FPGA system.
[0045] After generating the sub-viewpoint feature streams, the FPGA system packages 19 parallel feature streams and sends them to the Ascend 310 NPU processor through the PCIe interface for depth estimation and multi-focus reconstruction.
[0046] Reconstruction module 203, including a dual-branch network, for inputting the sub-viewpoint feature stream into the dual-branch network to perform multi-viewpoint depth estimation and multi-focus image synthesis, and outputting a 40×40 viewpoint parallel image array; The dual-branch network runs on the Ascend 310 NPU processor, which is connected to the FPGA system via a PCIe interface. For example, a PCIe 3.0×4 interface is used.
[0047] In the present disclosure, the dual-branch network needs to undergo channel pruning and INT8 quantization to reduce its size. In a specific embodiment, the dual-branch network after channel pruning and INT8 quantization is approximately 4 MB in size. It uses 19×3×3 sub-viewpoint input parallel reconstruction to expand to a 40×40 viewpoint parallel array to generate high-precision depth maps and multi-level focused images. The single-frame inference delay is less than or equal to 15 ms, providing complete light field information for back-end display.
[0048] Management module 204 is used to generate a control signal and an acquisition module start / stop mask based on the scene tag; the control signal is used to drive the corresponding DPR logic function in the FPGA system, and the start / stop mask is used to control the start / stop of the acquisition module sub-array channel to realize the start / stop control of the main collector and the secondary collector.
[0049] The subarray channel is the hardware carrier of the DPR logic, specifically the signal lines from each camera received by the FPGA system. The DPR logic is used to load different hardware functions according to different scenario requirements. The DPR logic loads or unloads the bitstream of the corresponding reconfigurable partition through the ICAP interface. The bitstream name and address of the DPR logic to be loaded are sent to the ICAP interface in the form of a control signal to initiate the reconfiguration of the reconfigurable partition corresponding to the DPR logic.
[0050] The switch matrix is a hardware switch placed between the smart camera array and the FPGA system interface. Based on the start / stop mask, it connects the LVDS / CSI signal lines of the corresponding subarray channels to the input pins or disconnect pins of the FPGA system, effectively turning the physical channels on or off. The start / stop mask is a bit pattern generated by the decision unit after a table lookup based on scene classification results. Each bit corresponds to a subarray channel. For example, a "1" indicates that the corresponding subarray channel is connected to the FPGA system, while a "0" indicates that the corresponding subarray channel is disconnected. The start / stop mask is output to the switch matrix via the GPIO interface to enable or disable the specified subarray channel.
[0051] The FPGA system includes a decision-making unit (DMU), consisting of a policy mapping table, state machine / lookup logic, and control signal generation. The policy mapping table stores the bitstream names corresponding to each scenario and a mask mapping table for each subarray channel, enabling functions such as starting or stopping ISP preprocessing, light field sub-image synthesis, and scene classification. It also synchronously triggers the initialization and switching of algorithm modules within the corresponding partition to switch image preprocessing functions for different scenarios. It also selectively writes specified bitstream files from the onboard QSPI flash or BRAM to the corresponding reconfigurable partition.
[0052] The scene classification results are refreshed every 100 ms. After the scene label is updated, the decision unit looks up the policy mapping table in the BRAM memory to generate the control signal of the DPR logic and the start and stop mask of the acquisition module: In the "Daytime" mode, enable the main collector and the secondary collector in the acquisition module and load daylight_standard.bit; In "low light" mode, only enable the inner loop sub-collector in the acquisition module and load low_light_fusion.bit; Enable the main collector and inner loop secondary collector in the acquisition module in "high speed" mode and load motion_stab.bit; Enable the main collector and outer ring secondary collector to collect in "long distance" mode and load super_res_recon.bit; In the "Panorama" mode, enable the main and secondary collectors in the acquisition module again and load panorama_render.bit.
[0053] Different modes correspond to the start and stop of different main collectors and secondary collectors, and load different control signals.
[0054] During operation, the start / stop mask controls the switch matrix of the acquisition module through the GPIO interface. The switch matrix enables or disconnects the corresponding sub-array channels in the acquisition module according to the start / stop mask. The sub-array channels control the start / stop of the main collector and the secondary collector. The refined start / stop control of the main collector and the secondary collector is achieved through the camera physical channel, switch hardware and management logic.
[0055] This disclosure also includes: using the FPGA system's clock gating technology to dynamically shut down unused sub-array channels and drive the Ascend 310 NPU processor to switch to low-power mode, reducing overall power consumption by approximately 30%.
[0056] In this disclosure, copper heat pipes are used to directly contact the FPGA system and the Ascend 310 NPU processor (AI chip) for heat dissipation. The exterior is equipped with a honeycomb aluminum alloy shell and dual 9 cm PWM fans. The speed is dynamically adjusted by the built-in PID algorithm of the FPGA system to ensure stable operation in a wide temperature range of -20℃ to 70℃, and the chip temperature is always less than or equal to 65℃.
[0057] In summary, scene classification is triggered every 100 ms, and the decision unit looks up the table in the BRAM memory to generate the control signal of the ICAP interface and the start and stop mask of the GPIO interface, which correspond to the sub-array combination and bit stream switching of modes such as "daytime", "low light", "high speed", "long distance", and "panoramic", realizing true seamless switching of multiple scenes. When the scene label is updated, the above process is repeated. The scene classification function and management module work together to achieve closed-loop adaptation of hardware and algorithms.
[0058] The present disclosure also provides a display device, comprising: As previously described, the intelligent light field array camera system is used to output a 40×40 viewpoint parallel image array; A display module is used to visualize the 40×40 viewpoint parallel array to achieve naked-eye 3D display; the display module includes an LCD panel and a 40×40 microlens array, each microlens of the microlens array is aligned with a pixel area on the LCD panel.
[0059] In a specific embodiment, the light field display consists of an 8K (7680×4320) high refresh LCD panel and a 40×40 microlens array, where each microlens is approximately 0.5 mm in size and corresponds to a viewpoint.
[0060] The 40×40 viewpoint parallel array output by the Ascend 310 NPU is mapped to the corresponding viewing angles through parallel interpolation and parallax distribution algorithms, enabling multi-person, glasses-free 3D display at 30-60 Hz. Simultaneously, the display automatically adjusts microlens array parameters based on actual viewing distance and parallax data, ensuring an optimal stereoscopic viewing experience with no crosstalk and continuous 3D within a range of 10-200 cm.
[0061] In a specific embodiment, the image collector provided by the present disclosure is arranged on a planar bracket. The main collector, i.e., a high-resolution camera, is installed at the center of the planar bracket. Its photosensitive chip and optical lens are precisely calibrated, with a pixel pitch of 4.8 μm and a field of view of 45°. Six small cameras are arranged in an inner ring at equal intervals along the radius R1 of a regular hexagon, R1≈50 mm, and the angle between adjacent small cameras in the regular hexagon is 60°, which is used to capture close-range parallax information. Twelve small cameras are arranged in an outer ring at equal intervals along the radius R2, R2≈100 mm, and the angle between adjacent small cameras in the regular hexagon is 30°, which expands the overall field of view and improves the long-range depth resolution. All cameras are aligned with the unified optical axis reference line through M12 bayonet mounts, and cables are reserved on the back of the bracket to connect to the LVDS signal harness channel to ensure neat wiring and minimal interference during synchronous acquisition.
[0062] The 19 LVDS camera signals captured by the 19 cameras, or raw image data, first enter the FPGA system's ISP pre-processing reconfigurable partition, which uses a row-level deep pipeline structure. Each channel of raw image data undergoes denoising, white balancing, geometric distortion correction, and HDR fusion in sequence on the pipeline, with processing latency less than 5 ms and a total throughput of 5 Gpixels per second.
[0063] The raw image data after ISP preprocessing enters the reconfigurable partition for light field sub-image synthesis. According to the pre-loaded calibrated disparity baseline, each channel image is cropped into a sub-viewpoint image with a resolution of 480×270 and aligned in real time through the AXI4-Stream bus.
[0064] In the parallel scene classification reconfigurable partition, the FPGA system stores the weights of the quantized MobileNetV2 lightweight neural network in BRAM memory and reads the 320×240 preview frame from the central camera and data from the sensor module via the AXI4-Lite bus. It completes an inference cycle within 100 milliseconds, outputs the scene label, and drives the DPR control. The DPR logic loads the corresponding bitstream via the ICAP interface. The entire switching process takes less than 10 milliseconds and does not affect the continued operation of other reconfigurable partitions. The FPGA system transmits 19 sub-viewpoint feature streams to the Ascend 310NPU processor via the PCIe 3.0×4 interface.
[0065] In the reconstruction module, a dual-branch lightweight network on the Ascend 310 NPU performs depth estimation and multi-focus reconstruction, respectively. The depth estimation branch uses a five-layer EPI-Net architecture to output a 640×360 layered depth map. The focus branch uses a multi-scale U-Net to generate a three-level focus map. The weights of both branches undergo channel pruning and INT8 quantization before being stored in 16GB of HBM memory, achieving inference latency of less than or equal to 15ms.
[0066] The Ascend 310 NPU processor sends the 40×40 viewpoint parallel image array to the DDR memory of the FPGA system through the PCIe interface or AXI DMA controller for cache. The FPGA system or external GPU processor then completes the final parallel interpolation and disparity allocation algorithm.
[0067] See also Figure 3The DPR logic executes in coordination with the FPGA system's internal logic and the QSPI flash memory. Specifically, the DPR logic performs the following steps: the scene classification reconfigurable partition performs real-time inference on the scene and writes the classification results to the status register; the decision unit reads the policy mapping table in the BRAM memory to generate control signals for the ICAP interface and start / stop masks for the GPIO interface. Based on the control signals, the ICAP interface pulls the corresponding bitstream (such as "low_light_fusion.bit" or "motion_stab.bit") from the QSPI flash memory and performs reconfiguration in the corresponding reconfigurable partition. Simultaneously, the GPIO interface outputs signals that drive the switch matrix to enable or disable the corresponding camera channel, ensuring that only the corresponding subarray channels in the array are active and reducing bus load. After reconfiguration, the control flow returns to the scene classification reconfigurable partition for the next round of steps, forming a closed-loop adaptive switching process that occurs every 100ms, dynamically adjusting the subarray channels and algorithms to adapt to environmental changes.
[0068] The display module receives a 40×40 viewpoint parallel image array from the Ascend 310 NPU processor via a DisplayPort 1.4 interface, driving a 7680×4320 high-refresh rate LCD panel and a 40×40 microlens array for synchronous refresh. The microlens array is precisely aligned with FZ precision (micrometer-level positioning accuracy is required in the focal plane direction), with each microlens aligned to a single viewpoint on the LCD panel. Each microlens measures 0.5 mm×0.5 mm, corresponding to a 0.5° viewing angle, ensuring a cross-talk-free, auto-glasses 3D effect for multiple viewers at distances of 10 to 200 cm. The display module's interpolation and parallax allocation algorithms are implemented in the FPGA system based on pre-compiled IP cores, enabling full frame output within 5 ms, with an overall latency of less than or equal to 25 ms.
[0069] In addition, by combining the FPGA system's clock gating with the Ascend 310 NPU processor's power mode switching, the clocks of unused sub-array channels are shut down, saving energy on the chip and reducing overall power consumption by approximately 30%. A built-in PID algorithm dynamically adjusts the casing's heat dissipation design, ensuring that the system's peak temperature remains below 65°C over a wide temperature range of -20°C to 70°C, ensuring both high performance and high reliability.
[0070] Based on the above technical solution, the present disclosure provides an image collector, which is divided into a primary collector and a secondary collector, and realizes multi-view light field acquisition in the center and two ring areas. The FPGA system in the front-end dynamic processing module is divided into three reconfigurable partitions based on the Xilinx Zynq UltraScale+ platform: ISP pre-processing, light field sub-image synthesis, and scene division. It completes denoising, white balance, geometric correction, HDR fusion and sub-view image extraction in parallel through a row-level deep pipeline, with a processing throughput of up to 5G pixels / second. The mid-end reconstruction module is connected to the FPGA system through a PCIe 3.0×4 interface, runs a lightweight dual-branch network after quantization and channel pruning, performs multi-view depth estimation and multi-focus image synthesis respectively, expands the sub-view feature stream into a 40×40 viewpoint parallel image array, generates a depth map and a focus hierarchy map, and outputs them in less than or equal to 20 ms. The present disclosure also includes a back-end display module that uses an 8K LCD panel and a microlens array, and employs a parallel interpolation algorithm to drive a 30-60 Hz naked-eye multi-viewer 3D display. The FPGA system scene classification function is implemented based on a quantized MobileNetV2 lightweight neural network, which recognizes five scene labels in real time: "daytime, low light, high speed, long distance, and panoramic view." This drives the DPR logic and the start and stop of the sub-array array, achieving closed-loop adaptation of the hardware and algorithm. Through modular design and dynamic collaboration, this paper achieves the organic integration of high-resolution multi-viewpoint light field acquisition, real-time adaptive algorithm switching and high-fidelity naked-eye 3D display, significantly improving the system's flexibility, real-time performance and energy efficiency. This paper organically combines high-density intelligent camera array acquisition, reconfigurable hardware acceleration with AI chip reasoning and multi-viewpoint high-refresh naked-eye display to achieve a multi-scene adaptive, low-latency, high-energy-efficiency and high-reliability light field imaging and display system. This paper adopts copper heat pipes, honeycomb aluminum alloy shell, PID algorithm intelligent air cooling, and cooperates with the clock gating of the FPGA system and the power consumption mode switching of the AI chip to achieve low power consumption and high stability operation in a wide temperature range of -20℃ to 70℃.
[0071] Various embodiments of the systems and techniques described herein can be implemented in digital electronic circuit systems, integrated circuit systems, application specific integrated circuits (ASICs), application specific standard products (ASSPs), systems on chips (SOCs), programmable logic devices (CPLDs), computer hardware, firmware, software, and / or combinations thereof. These various embodiments can include being implemented in one or more computer programs that are executable and / or interpreted on a programmable system comprising at least one programmable processor, which can be a special purpose or general purpose programmable processor that can receive data and instructions from a storage system, at least one input device, and at least one output device, and transmit data and instructions to the storage system, the at least one input device, and the at least one output device.
[0072] The program code for implementing the method of the present disclosure can be written in any combination of one or more programming languages. These program codes can be provided to a processor or controller of a general-purpose computer, a special-purpose computer, or other programmable data processing device so that when the program code is executed by the processor or controller, the functions / operations specified in the flow chart and / or block diagram are implemented. The program code can be executed entirely on the machine, partially on the machine, as a stand-alone software package, partially on the machine and partially on a remote machine, or entirely on a remote machine or server.
[0073] In the context of the present disclosure, a machine-readable medium may be a tangible medium that may contain or store a program for use by or in conjunction with an instruction execution system, apparatus, or device. A machine-readable medium may be a machine-readable signal medium or a machine-readable storage medium. A machine-readable medium may include, but is not limited to, an electronic, magnetic, optical, electromagnetic, infrared, or semiconductor system, apparatus, or device, or any suitable combination of the foregoing. More specific examples of machine-readable storage media may include an electrical connection based on one or more wires, a portable computer disk, a hard disk, a random access memory (RAM), a read-only memory (ROM), an erasable programmable read-only memory (EPROM or flash memory), optical fibers, a portable compact disk read-only memory (CD-ROM), an optical storage device, a magnetic storage device, or any suitable combination of the foregoing.
[0074] To provide interaction with a user, the systems and techniques described herein can be implemented on a computer having: a display device for displaying information to the user; and a keyboard and pointing device (e.g., a mouse or trackball) through which the user can provide input to the computer. Other types of devices can also be used to provide interaction with the user; for example, feedback provided to the user can be any form of sensory feedback (e.g., visual feedback, auditory feedback, or tactile feedback); and input from the user can be received in any form (including acoustic input, voice input, or tactile input).
[0075] It should be understood that the various forms of the processes shown above can be used to reorder, add, or delete steps. For example, the steps described in this disclosure can be performed in parallel, sequentially, or in a different order, as long as the desired results of the technical solutions disclosed in this disclosure can be achieved. This is not limited herein.
[0076] The above specific embodiments do not constitute a limitation on the scope of protection of this disclosure. Those skilled in the art will appreciate that various modifications, combinations, sub-combinations, and substitutions may be made based on design requirements and other factors. Any modifications, equivalent substitutions, and improvements made within the spirit and principles of this disclosure shall be included within the scope of protection of this disclosure.
Claims
1. An image collector, applied to an intelligent light field array camera system, comprising: 1 main collector and several secondary collectors; The main collector is used to capture detail images, and the secondary collector is used to capture close-range parallax images and overall field of view images; N1 secondary collectors are distributed at equal intervals around the main collector to form an inner ring of a regular hexagon; N2 secondary collectors are distributed at equal intervals outside the inner ring to form an outer ring of a regular hexagon; Where N1 and N2 are both integers, and N1 ≥ 4, N2 ≤ 32, N2 = 2N 1。 2. An intelligent light field array camera system, comprising the following modules: An image acquisition module, comprising the image acquisition device according to claim 1, for acquiring original images, including detail images, close-range parallax images, and overall field of view images; The dynamic processing module, built on an FPGA system, includes three reconfigurable partitions, which respectively deploy ISP preprocessing, light field sub-image synthesis, and scene classification functions, and are used to output sub-viewpoint feature streams and scene labels of the original image; A management module, configured to generate a control signal and an acquisition module start / stop mask according to the scene tag; the control signal is used to drive the corresponding DPR logic function in the FPGA system, and the start / stop mask is used to control the start / stop of the primary and secondary acquisition modules of the acquisition module; The reconstruction module is used to input the sub-viewpoint feature stream into the dual-branch network for multi-viewpoint depth estimation and multi-focus image synthesis, and output a 40×40 viewpoint parallel image array.
3. The intelligent light field array camera system according to claim 2, characterized in that: The ISP pre-processing function simultaneously completes denoising, white balance, geometric correction and HDR fusion through a row-level deep pipeline.
4. The intelligent light field array camera system according to claim 2, characterized in that: The light field sub-image synthesis function crops the original images collected by the primary and secondary collectors after ISP preprocessing into sub-viewpoint images, and aligns them according to the calibrated disparity baseline to obtain the sub-viewpoint feature stream.
5. The intelligent light field array camera system according to claim 2, characterized in that: Also includes: The sensor module includes: a temperature sensor for collecting the brightness of the environment; a brightness sensor for collecting the brightness of the environment; an IMU sensor for collecting acceleration and angular velocity data of the intelligent light field array camera system; The original image collected by the main collector after ISP preprocessing, the temperature and brightness of the environment, and the acceleration and angular velocity data are input into the MobileNetV2 lightweight neural network to realize the scene classification function, and the real-time output includes scene labels including "daytime", "low light", "high speed", "long distance" or "panoramic".
6. The intelligent light field array camera system according to claim 5, characterized in that: Generating a control signal and an acquisition module start / stop mask according to the scene tag includes: In "daylight" mode, enable the main and secondary collectors in the acquisition module and load daylight_standard.bit; In "low light" mode, only enable the inner loop sub-collector in the acquisition module and load low_light_fusion.bit; Enable the main collector and inner loop secondary collector in the acquisition module in "high speed" mode and load motion_stab.bit; Enable the main collector and outer ring secondary collector to collect in "long distance" mode and load super_res_recon.bit; In "Panorama" mode, enable the primary and secondary collectors in the acquisition module again and load panorama_render.bit.
7. The intelligent light field array camera system according to claim 2, characterized in that: The start / stop mask controls the switch matrix of the acquisition module through the GPIO interface. The switch matrix enables or disconnects the corresponding sub-array channels in the acquisition module according to the start / stop mask. The sub-array channels control the start / stop of the main collector and the secondary collector.
8. The intelligent light field array camera system according to claim 2, wherein: In the reconstruction module, the dual-branch network runs on the Ascend 310 NPU processor, which is connected to the FPGA system via a PCIe interface.
9. A display device comprising: The intelligent light field array camera system according to any one of claims 2 to 8, configured to output a 40×40 viewpoint parallel array; A display module is used to visualize the 40×40 viewpoint parallel array to achieve naked-eye 3D display; the display module includes an LCD panel and a 40×40 microlens array, each microlens of the microlens array is aligned with a pixel area on the LCD panel.
10. The display device according to claim 9, wherein: Visualizing the 40×40 viewpoint parallel image array output by the reconstruction module in the intelligent light field array camera system, including: The 40×40 viewpoint parallel array is mapped through a parallel interpolation and parallax allocation algorithm, and is visually displayed through a display; wherein the display automatically corrects the parameters of the microlens array parameters.