Image processing method based on light field, computer device and computer readable storage medium
Through the image processing methods of multi-camera video surveillance system, including resolution enhancement and viewing angle integration, the problem that a single-camera system cannot effectively capture viewing angle changes and depth information is solved, and high-precision scene depth detection and high-definition video output are achieved.
Patent Information
- Application Number
- CN202210600260.7
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Filing Date
- 2022-05-30
- Publication Date
- 2025-05-09
- Estimated Expiration
- 2042-05-30
AI Technical Summary
The single-camera video surveillance system cannot effectively capture parallax changes caused by changes in view angles, which limits the accurate detection of object depth information. Due to the sensor resolution, it is difficult to accurately detect information from objects far away, increasing the difficulty of monitoring.
Multiple cameras are used to capture image data, and through resolution enhancement processing and viewing angle integration, shallow features are extracted and feature extraction and fusion are performed, super-resolution image data is generated, and scene depth detection is performed through fast depth estimation algorithm.
It realizes low-cost and accurate scene depth calculation, improves the accuracy and reliability of the video surveillance system for object distance detection, and provides high-definition and ultra-definition video output.
Smart Images

Figure CN114897958B_ABST
Abstract
Description
Technical Field
[0001] The present invention relates to the technical field of image processing, and in particular to a light field-based image processing method applied to video surveillance, and a computer device and a computer-readable storage medium for implementing the method. Background Art
[0002] As people's safety awareness increases, intelligent video surveillance systems are widely used in security systems. At present, in the field of video surveillance, video image resolution reconstruction and distance monitoring are important monitoring directions of intelligent video surveillance systems. Since current video surveillance equipment is often a fixed scene captured by only a single camera, and the resolution of components also limits the clarity of the image obtained by the video surveillance equipment, the video surveillance system cannot capture more and more detailed video information. Specifically, the currently commonly used single-camera video surveillance system has the following disadvantages:
[0003] First, compared with the image data captured by a multi-camera camera, the image data obtained by a single camera lacks the change of parallax in the image data caused by the change of perspective. The depth information of the specific target obtained by the single camera is equivalent to inferring the depth information of the three-dimensional space from the two-dimensional image. This feature hinders the single-camera video surveillance system from detecting the distance of the object.
[0004] Second, single-camera cameras are limited by the resolution of the sensor. For objects that are far away in the data, more accurate information cannot be obtained, which seriously affects accurate distance detection and increases the difficulty of monitoring. Although replacing high-definition cameras can solve this problem to a certain extent, it will bring higher costs.
[0005] Third, although there are many options for super-resolution algorithms for single-camera cameras, performing super-resolution calculations on images from each perspective separately will destroy the structural characteristics of the image and reduce the accuracy of depth estimation. Summary of the invention
[0006] The first object of the present invention is to provide a light field-based image processing method that is low-cost and can accurately calculate scene depth.
[0007] A second objective of the present invention is to provide a computer device for implementing the above-mentioned light field-based image processing method.
[0008] A third object of the present invention is to provide a computer-readable storage medium for implementing the above-mentioned light field-based image processing method.
[0009] To achieve the first purpose of the present invention, the light field-based image processing method provided by the present invention includes obtaining initial image data output by multiple cameras, performing resolution enhancement processing on the initial image data of the multiple cameras: integrating the perspectives of the initial image data of the multiple cameras, and obtaining shallow features of the image after intensive feature extraction, and obtaining super-resolution image data after feature extraction and fusion of the shallow features; using a video containing super-resolution image data to perform scene depth estimation and prediction to obtain scene depth information of the center position; storing the super-resolution image data and the scene depth information in a preset data storage module; obtaining a video output instruction, and obtaining and displaying the corresponding super-resolution image data and scene depth information according to the video output instruction.
[0010] As can be seen from the above scheme, the present invention draws on the characteristics of light field acquisition equipment and adds a certain amount of video capture equipment, such as setting up multiple cameras, without increasing the cost as much as possible. Since the present invention adopts a lightweight multi-eye super-resolution algorithm, that is, an algorithm for performing resolution enhancement processing on the initial image data, an algorithm for extracting shallow features, etc., it can simultaneously realize the switching of 2K high-definition or 4K ultra-clear scene resolution of each camera, and can obtain scene information with higher reliability and higher definition than a single camera.
[0011] A preferred solution is that integrating the perspective of the initial image data of multiple cameras includes: adjusting the initial image data of the multiple cameras into image data of tensors having multiple dimensions.
[0012] It can be seen that integrating the perspective of images, especially integrating data of multiple dimensions into a tensor data, is conducive to improving the speed of data processing.
[0013] A further solution is that dense feature extraction is performed on the image data after perspective integration, including: using a dense feature extraction module to perform dense feature extraction, the dense feature extraction module includes a first dilute spatial pyramid pooling module, a second dilute spatial pyramid pooling module, a first convolution module and a residual convolution module that are cascaded in sequence.
[0014] It can be seen that the first atrous spatial pyramid pooling module is used to improve the receptive field of the initial feature, and the second atrous spatial pyramid pooling module can further extract deep feature representation, which is a combination of the initial feature and the output feature, and can further extract multi-level feature information.
[0015] A further solution is that the first atrous spatial pyramid pooling module has the same structure as the second atrous spatial pyramid pooling module; the first atrous spatial pyramid pooling module includes three parallel dilated convolutions with different dilation rates.
[0016] It can be seen that using dilated convolutions with different dilation rates makes the three dilated convolutions have different receptive fields and can obtain more image information.
[0017] A further solution is to extract and fuse the shallow features, including: extracting features between viewpoints and extracting features within viewpoints respectively, extracting and integrating the high-frequency information of each viewpoint, and fusing the feature information obtained by extracting features between viewpoints and extracting features within viewpoints.
[0018] By performing feature extraction between viewpoints and feature extraction within viewpoints respectively, the information between different viewpoints and the information within each viewpoint can complement each other to obtain more accurate image information.
[0019] A further solution is that the inter-view feature extraction of shallow features is implemented by applying the inter-view feature extraction module, and the inter-view feature extraction module includes a third space hole pyramid pooling module, a first deconvolution module, a fourth space hole pyramid pooling module, a second deconvolution module and a second convolution module which are cascaded in sequence; wherein the third space hole pyramid pooling module is used to extract differential features with rich text information, the first deconvolution module is used to map the feature map containing the differential features to a high-dimensional space, the fourth space hole pyramid pooling module is used to extract the differential features again, the second deconvolution module is used to map the high-dimensional features to the low-dimensional features, and the second convolution module is used to increase the depth of the model.
[0020] A further solution is to use an intra-view feature extraction module to extract intra-view features based on shallow features. The structure of the intra-view feature extraction module is the same as that of the inter-view feature extraction module, and the input dimension of the intra-view feature extraction module is different from the input dimension of the inter-view feature extraction module.
[0021] It can be seen that different input dimensions are set according to the difference between feature extraction within a view and feature extraction between views to meet the needs of image extraction and make the extraction of image features more accurate.
[0022] A further solution is to perform the operations of extracting features between views, extracting features within views, and fusing feature information obtained by extracting features between views and extracting features within views twice or more.
[0023] In this way, a residual map with rich information between and within each perspective can be formed. The residual map is added to the result of upsampling between the previous perspectives, and finally the resolution of multiple perspectives is improved through the upsampling module, providing video output in both HD and UHD modes.
[0024] To achieve the second objective, the computer device provided by the present invention includes a processor and a memory, wherein the memory stores a computer program, and when the computer program is executed by the processor, each step of the above-mentioned light field-based image processing method is implemented.
[0025] To achieve the third objective mentioned above, the present invention provides a computer readable storage medium storing a computer program, and when the computer program is executed by a processor, each step of the above light field-based image processing method is implemented. BRIEF DESCRIPTION OF THE DRAWINGS
[0026] Figure 1 It is a system structure block diagram of an embodiment of the image processing method based on light field of the present invention.
[0027] Figure 2 It is a schematic diagram of the system structure of an embodiment of the image processing method based on light field of the present invention.
[0028] Figure 3 It is a flow chart of an embodiment of the image processing method based on light field of the present invention.
[0029] Figure 4 It is a structural schematic diagram of a video acquisition module according to an embodiment of the light field-based image processing method of the present invention.
[0030] Figure 5 It is a structural block diagram of a resolution enhancement module according to an embodiment of the light field-based image processing method of the present invention.
[0031] Figure 6 It is a structural block diagram of a dense feature extraction module according to an embodiment of the light field-based image processing method of the present invention.
[0032] Figure 7 It is a structural block diagram of a first atrous spatial pyramid pooling module according to an embodiment of the light field-based image processing method of the present invention.
[0033] Figure 8 It is a structural block diagram of a residual convolution module of an embodiment of the light field-based image processing method of the present invention.
[0034] Fig. 9 It is a structural block diagram of an inter-view feature extraction module according to an embodiment of the light field-based image processing method of the present invention.
[0035] The present invention is further described below in conjunction with the accompanying drawings and embodiments. DETAILED DESCRIPTION
[0036] The image processing method based on light field of the present invention is used to process the image of the surveillance video. Preferably, the method can be implemented by a computer device, for example, the computer device is provided with a processor and a memory, the memory stores a computer program, and the processor implements the above-mentioned image processing method based on light field by executing the computer program.
[0037] Embodiment of the image processing method based on light field:
[0038] The main idea of this embodiment is to improve the existing video surveillance equipment with the help of the characteristics of the existing light field acquisition equipment, and to use the structural characteristics of the light field data to increase multiple perspectives to capture information about the scene without increasing the cost as much as possible. At the same time, this embodiment is a new lightweight multi-perspective super-resolution algorithm, which simultaneously performs high-definition 2K and ultra-clear 4K super-resolution processing on the entire 720P monitoring equipment. In addition, based on each perspective after super-resolution, a fast depth estimation algorithm is introduced to detect the distance of multiple targets in the scene.
[0039] See also Figure 1 and Figure 2 The video surveillance system using this embodiment includes a video acquisition module 21, a resolution enhancement module 22, a scene depth estimation and prediction module 23, a data storage module 24, a system management module 25, a display module 26, a video database 27 and a user database 28, and is provided with a data transmission channel 91 and a data management server 92.
[0040] The video acquisition module 21 is used to acquire initial image data, and send the acquired initial image data to the resolution enhancement module 22 through the data transmission channel 91. The data transmission channel 91 is preferably a high-speed Ethernet device to ensure fast information transmission between various modules.
[0041] The resolution enhancement module 22 is used to perform resolution enhancement calculation on the initial image data, for example, to form a high-resolution image. The scene depth estimation prediction module 23 uses the high-resolution image output by the resolution enhancement module 22 to estimate and predict the scene depth, and obtain scene depth estimation prediction information. The high-resolution image output by the resolution enhancement module 22 and the scene depth estimation prediction information obtained by the scene depth estimation prediction module 23 are stored in the data storage module 24.
[0042] The system management server 25 is used to manage the entire system, for example, to receive user instructions, obtain corresponding data according to the user's instructions, and send the obtained data to the display module 26 for display by the display module 26. Preferably, the display module 26 can be a display screen or a terminal device with a display function. The video database 27 is used to store video data. For example, the data obtained by the data storage module 24 can be stored in the video database 27 for subsequent retrieval and reference. The user database 28 can store user data, such as storing user ID information. The data management server 92 can manage the data of the video database 27 and the user database 28.
[0043] Combine the following Figure 3 The workflow of this embodiment is described. First, step S11 is executed to obtain initial image data of multiple cameras in the video acquisition module 21. Figure 4 The video acquisition module 21 includes multiple cameras, for example, 9 cameras 31, and the 9 cameras 31 are arranged in a 3×3 grid architecture, and the resolution of each camera 31 is 720P. In the grid architecture of the video acquisition module 21, the 9 cameras 31 are evenly distributed in 3 rows and 3 columns, all cameras 31 remain on the same plane, and the maximum parallax of each frame image taken by adjacent cameras 31 at the same time is less than 15. In the grid architecture, the spacing of each grid is a, that is, in each row, the distance between two adjacent cameras 31 is a, and in each column, the distance between two adjacent cameras 31 is also a. In addition, the optical axes of each camera 31 remain parallel to each other and perpendicular to the grid architecture.
[0044] Each camera 31 can obtain independent initial image data, and send the obtained initial image data to the resolution enhancement module 22. At this time, the resolution enhancement module 22 executes step S12 to perform resolution enhancement processing on the obtained initial image data. Specifically, the resolution enhancement module 22 performs super-resolution processing on the video image data of multiple perspectives collected by multiple cameras 31, and adopts a super-resolution network with a novel fusion and distribution strategy. Through dense feature extraction, inter-perspective feature extraction operator (AIO) and intra-perspective feature extraction operator (SIO), fusion and distribution strategy (IIFDB), the features of the initial image data obtained by each camera 31 are fully extracted, and the fusion and supplement of high-frequency information between multiple perspectives are realized, and finally a solution for super-resolution of all cameras at the same time is realized.
[0045] The network structure of the resolution enhancement module 22 is as follows: Figure 5As shown, each camera 31 outputs the initial image data of a viewing angle, and the initial image data of all viewing angles are first integrated by the viewing angle, and the input of the 9 independent initial image data is adjusted to a tensor with 9 dimensions as input. For example, the dimension of the initial image data output by each camera 31 is 3*H*W, where H and W are the size of the initial image data, that is, the height and width of the initial image data. The function of the video integration module is to adjust the dimension of the 9 independent initial image data to 9*32*H*W.
[0046] Specifically, the video integration module performs two operations, namely stacking and convolution. The stacking operation can be implemented using the torch.stack() function of pytorch. This operation can add a new data dimension for stacking. For example, the video input module outputs a total of 9 perspectives of initial image data. The dimension of the initial image data of each perspective is 3*H*W. After the stacking operation, it becomes 9*3*H*W, which is conducive to parallel computing of data. The second operation is the convolution operation. The purpose of the convolution operation is to change the number of channels from 3 to 32, so that the dimension of the final output data becomes 9*32*H*W. Such an operation is conducive to the network learning deeper feature representation.
[0047] The video integration module can convert the channel dimensions of the input data on the one hand, and integrate the 9-dimensional data into a tensor data on the other hand, which is conducive to improving the speed of data processing.
[0048] After video integration, the image data is subjected to two dense feature extraction modules for feature extraction. The structure of each dense feature extraction module is as follows: Figure 6 The dense feature extraction module includes a first atrous spatial pyramid pooling module 51, a second atrous spatial pyramid pooling module 52, a first convolution module 53 and a residual convolution module 54 which are cascaded in sequence.
[0049] The first atrous spatial pyramid pooling module 51 has the same structure as the second atrous spatial pyramid pooling module 52. Figure 7 As shown, the first dilated spatial pyramid pooling module 51 includes three parallel dilated convolutions 61, 62, and 63 with different dilation rates, wherein the convolution kernel size of each dilated convolution 61, 62, and 63 is 3×3, and the dilation rates of the three dilated convolutions 61, 62, and 63 are 1, 2, and 4, respectively. In this way, the three dilated convolutions 61, 62, and 63 have different receptive fields, respectively, so that more image information can be obtained. The output ends of the three dilated convolutions 61, 62, and 63 are respectively connected to activation functions 64, 65, and 66, and the three activation functions 64, 65, and 66 are all output to connection 67 to form output data.
[0050] In this embodiment, two cascaded first atrous space pyramid pooling modules 51 and second atrous space pyramid pooling modules 52 are arranged because the feature extraction adopts density connection. The first atrous space pyramid pooling module 51 is used to improve the receptive field of the initial feature, and the second atrous space pyramid pooling module 52 is used to further extract deep feature representation, which is a combination of the initial feature and the output feature, and can further extract multi-level feature information.
[0051] The data passed through the second atrous spatial pyramid pooling module 52 passes through the first convolution module 53 and is input into the residual convolution module 54. The structure of the residual convolution module 54 is as follows: Figure 8 As shown, it includes two third convolution modules, namely, third convolution modules 71 and 72. The convolution kernel size of the two third convolution modules 71 and 72 is 3×3, the step size is 1, and the edge filling of the feature map is 0. After the feature extraction by the dense feature extraction module, the shallow features of the image data are obtained.
[0052] The resolution enhancement module 22 is also provided with an inter-view feature extraction module, an intra-view feature extraction module, and a fusion and distribution module. Figure 5 It can be seen that the operations of inter-view feature extraction, intra-view feature extraction and fusion of feature information obtained by inter-view feature extraction and intra-view feature extraction are executed in a loop more than twice, for example, four times, that is, the number of inter-view feature extraction modules, intra-view feature extraction modules, fusion and distribution modules are all four.
[0053] The shallow features obtained through dense feature extraction are input into the first-level inter-view feature extraction module and the intra-view feature extraction module. The inter-view feature extraction and the intra-view feature extraction are mainly to extract deep-level features in two dimensions: inter-view and intra-view. The angular domain is to realize the integration of feature information between different viewpoints of each channel, while the spatial domain is to realize the information fusion of different feature dimensions under the same viewpoint.
[0054] The structure of the inter-view feature extraction module is as follows: Fig. 9 As shown, it includes a third space hole pyramid pooling module 81, a first deconvolution module 82, a fourth space hole pyramid pooling module 83, a second deconvolution module 84, and second convolution modules 85 and 86 that are cascaded in sequence. Among them, the third space hole pyramid pooling module 81 and the fourth space hole pyramid pooling module 83 have the same structure, and can refer to the structure of the first space hole pyramid pooling module 51, which includes three dilated convolutions with different dilation rates. Each dilated convolution can extract features with different receptive fields, thereby increasing the diversity of convolutions, which is conducive to extracting differentiated features with rich text information.
[0055] In this embodiment, the third space hole pyramid pooling module 81 is used to extract differentiated features with rich text information, and the first deconvolution module 82 is used to map the feature map containing the differentiated features to a high-dimensional space. For example, the dimensions of the input feature map are C, N, H, W, and the dimensions of the feature map mapped to the high-dimensional space are C, N, 2*H, 2*W or C, N, 4*H, 4*W. The application of the first deconvolution module 82 can better explore the mapping relationship between low-resolution images and high-resolution image pairs, thereby improving the quality of the super-resolution image finally output.
[0056] The fourth spatial hole pyramid pooling module 83 has the same function as the third spatial hole pyramid pooling module 81, and is used to extract differential features again. The second deconvolution module 84 is used to map high-dimensional features to low-dimensional features. For example, the dimensions of the input feature map are C, N, 2*H, 2*W or C, N, 4*H, 4*W. After the second deconvolution module 84, the dimensions of the obtained feature map are C, N, H, W. It can be understood that the second deconvolution module 84 draws on the design of the GAN network. The second convolution module 85 is used to increase the depth of the model.
[0057] The structure of the intra-view feature extraction module is the same as that of the inter-view feature extraction module, except that the input dimensions are different. Specifically, the input dimensions of the angular domain operator of the inter-view feature extraction module are C, N, H, W, and the output dimensions are C, N, H, W, while the input dimensions of the spatial domain operator of the intra-view feature extraction module are N, C, H, W, and the output dimensions are N, C, H, W, where N = 9, C = 32, and H and W are the sizes of the video image, i.e., height and width, respectively.
[0058] After extracting features between and within views, it is necessary to perform feature information fusion processing so that the information between different views and within each view can complement each other, and then distribute the features that contain rich information. Figure 5 It can be seen that the above process is cycled 4 times in total, and finally a residual map with rich information between and within each perspective is obtained. The residual map is added to the result of upsampling between the previous perspectives, and finally the resolution of multiple perspectives is improved through the upsampling module, providing video output in two modes of high definition and ultra-high definition, thereby obtaining images of nine super-resolution output perspectives, namely super-resolution output perspective 1 to super-resolution output perspective 9.
[0059] At this point, the resolution enhancement processing of the initial image data is completed. Figure 3After step S12 is completed, step S13 is executed, and the scene depth estimation prediction module 23 predicts the scene depth estimation of the high-resolution image video. Specifically, a 3×3 2K high-definition or 4K ultra-high-definition resolution surveillance video is used as input, and the rotation parallelogram operator is applied to calculate the scene depth information of the central perspective, wherein the operator input is an image of 9 perspectives, and the output is the scene depth information of the central perspective.
[0060] Next, step S14 is executed to store the super-resolution image obtained in step S12 and the scene depth information obtained in step S13 in the data storage module 24. Therefore, the data storage module 24 is mainly used to store high-definition and ultra-high-definition surveillance videos, and to store surveillance videos of the center position containing scene depth information, wherein the high-definition and ultra-high-definition surveillance videos are videos with a resolution of 2048×1152 (2K) and 4090×2160 (4K), respectively. Preferably, the data storage module 24 is directly connected to the video database 27, and can perform operations such as adding, deleting, and querying the video data stored in the video database 27.
[0061] Next, step S15 is executed to determine whether an instruction to display the video is received. If not, continue to wait. If an instruction to display the video is received, step S16 is executed to obtain the required super-resolution video according to the user's instruction, and the super-resolution video and the corresponding scene depth estimation prediction information are displayed in the display module 26. In this embodiment, the display module 26 can realize high-definition monitoring under different cameras, and can also display the distance of different targets in the video shooting scene at the center. The display module 26 consists of a 4K high-resolution display. The entire screen is divided into 9 blocks, representing the video images corresponding to 3×3 different cameras 31, and the entire window can be displayed for real-time monitoring and timed query. Among them, the middle window displays the distance information of different objects in the scene from the camera, and the rest display their respective videos after super-resolution processing. After double-clicking to open, the video data can be converted between 2K high-definition and 4K ultra-clear resolution. In addition, the middle window can also display the depth information in full screen.
[0062] In addition, the system management module 25 of this embodiment can be used in conjunction with the data storage module 24. After the user sets the parameters, the system management module 25 can receive the instructions sent by the user, process the videos stored in the video database 27, such as query and delete, and send the processed video data to the display module 26 to complete the display of the video data required by the user. The system management module 25 can manage the system mode, user level and log, wherein the system mode is divided into real-time monitoring and timed query mode, and the user level is mainly divided into high-authority users and ordinary users, wherein the high-authority user can delete and query the video data, and the ordinary user can only query, and the log will record the operations performed by each user.
[0063] It can be seen that this embodiment draws on the characteristics of light field acquisition equipment, adds a certain number of cameras while controlling costs, and uses a new lightweight multi-eye super-resolution algorithm to simultaneously achieve 2K high-definition or 4K ultra-clear scene resolution switching for each camera. Compared with a monitoring system using a single camera, this embodiment can obtain scene information with higher reliability and higher definition.
[0064] In addition, based on the above-mentioned super-resolution algorithm, the video resolution of 720P with a resolution of 1280×720 can be increased to a 2K high-definition resolution of 2048×1152, or increased to a 4K ultra-clear resolution of 4090×2160. Compared with the video directly captured by 4K high-definition equipment, its Peak Signal-to-Noise Ratio (PSNR) reaches 34.50, and the Structural Similarity Index (SSIM) reaches 0.9748, which has sufficient reliability.
[0065] Finally, for the multi-view video data after the resolution is enhanced, this embodiment uses a multi-view depth estimation algorithm, which can predict the depth estimation of scene targets at 2K high-definition or 4K ultra-high-definition resolution. Compared with the actual distance, its root mean square error (MSE) is less than 0.1, the error of the depth estimation can be ignored, and the result is highly reliable.
[0066] Computer device embodiment:
[0067] The computer device of this embodiment is an electronic device with video acquisition, processing and display functions. The computer device is provided with a processor, a memory and a computer program stored in the memory and executable on the processor. When the processor executes the computer program, each step of the above-mentioned light field-based image processing method is implemented.
[0068] For example, a computer program may be divided into one or more modules, one or more modules are stored in a memory and executed by a processor to complete each module of the present invention. One or more modules may be a series of computer program instruction segments capable of completing a specific function, and the instruction segments are used to describe the execution process of the computer program in a terminal device.
[0069] The processor referred to in the present invention may be a central processing unit (CPU), or other general-purpose processors, digital signal processors (DSP), application-specific integrated circuits (ASIC), field-programmable gate arrays (FPGA) or other programmable logic devices, discrete gate or transistor logic devices, discrete hardware components, etc. A general-purpose processor may be a microprocessor or any conventional processor, etc. The processor is the control center of the terminal device, and uses various interfaces and lines to connect various parts of the entire terminal device.
[0070] The memory can be used to store computer programs and / or modules. The processor realizes various functions of the terminal device by running or executing the computer programs and / or modules stored in the memory, and calling the data stored in the memory. The memory can mainly include a program storage area and a data storage area, wherein the program storage area can store an operating system, at least one application required for a function, etc.; the data storage area can store data created according to the use of the mobile phone, etc. In addition, the memory can include a high-speed random access memory, and can also include a non-volatile memory, such as a hard disk, a memory, a plug-in hard disk, a smart memory card (Smart Media Card, SMC), a secure digital (Secure Digital, SD) card, a flash card (FlashCard), at least one disk storage device, a flash memory device, or other volatile solid-state storage devices.
[0071] Computer readable storage medium embodiment:
[0072] If the computer program stored in the above-mentioned computer device is implemented in the form of a software functional unit and sold or used as an independent product, it can be stored in a computer-readable storage medium. Based on this understanding, the present invention implements all or part of the processes in the above-mentioned embodiment method, and can also be completed by instructing related hardware through a computer program. The computer program can be stored in a computer-readable storage medium. When the computer program is executed by a processor, it can implement each step of the above-mentioned light field-based image processing method.
[0073] Among them, the computer program includes computer program code, which can be in source code form, object code form, executable file or some intermediate form, etc. Computer readable media may include: any entity or device capable of carrying computer program code, recording medium, USB flash drive, mobile hard disk, magnetic disk, optical disk, computer memory, read-only memory (ROM), random access memory (RAM), electric carrier signal, telecommunication signal and software distribution medium, etc. It should be noted that the content contained in the computer readable medium can be appropriately increased or decreased according to the requirements of legislation and patent practice in the jurisdiction. For example, in some jurisdictions, according to legislation and patent practice, computer readable media do not include electric carrier signals and telecommunication signals.
[0074] Finally, it should be emphasized that the present invention is not limited to the above-mentioned embodiments. For example, changes in the internal structure of the inter-view feature extraction module, or changes in the specific structure of the atrous space pyramid pooling module, should also be included in the scope of protection of the claims of the present invention.
Claims
1. A light field-based image processing method, characterized in that: include: Acquire initial image data output by multiple cameras, and perform resolution enhancement processing on the initial image data of the multiple cameras: integrate the initial image data of the multiple cameras in perspective, and obtain shallow features of the image after intensive feature extraction, and obtain super-resolution image data after feature extraction and fusion of the shallow features; Using the video containing the super-resolution image data to perform scene depth estimation and prediction to obtain scene depth information at a center position; Storing the super-resolution image data and the scene depth information in a preset data storage module; Obtaining a video output instruction, and obtaining and displaying the corresponding super-resolution image data and the scene depth information according to the video output instruction; The shallow features are extracted and fused by the following steps: The shallow features are respectively subjected to inter-view feature extraction and intra-view feature extraction to extract and integrate the high-frequency information of each view, and the feature information obtained by the inter-view feature extraction and the intra-view feature extraction is fused; The shallow features are used to extract inter-view features by applying an inter-view feature extraction module, and the inter-view feature extraction module includes a third space hole pyramid pooling module, a first deconvolution module, a fourth space hole pyramid pooling module, a second deconvolution module and a second convolution module which are cascaded in sequence; Among them, the third spatial hole pyramid pooling module is used to extract differentiated features with rich text information, the first deconvolution module is used to map the feature map containing the differentiated features to a high-dimensional space, the fourth spatial hole pyramid pooling module is used to extract differentiated features again, the second deconvolution module is used to map high-dimensional features to low-dimensional features, and the second convolution module is used to increase the depth of the model.
2. The light field-based image processing method according to claim 1, characterized in that: Performing perspective integration on the initial image data of the multiple cameras includes: adjusting the initial image data of the multiple cameras into image data of tensors having multiple dimensions.
3. The light field-based image processing method according to claim 2, characterized in that: The dense feature extraction of the image data after perspective integration includes: using a dense feature extraction module to perform dense feature extraction, wherein the dense feature extraction module includes a first atrous space pyramid pooling module, a second atrous space pyramid pooling module, a first convolution module and a residual convolution module which are cascaded in sequence.
4. The light field-based image processing method according to claim 3, characterized in that: The first atrous space pyramid pooling module has the same structure as the second atrous space pyramid pooling module; The first atrous spatial pyramid pooling module includes three parallel dilated convolutions with different dilation rates.
5. The light field-based image processing method according to claim 1, characterized in that: The shallow feature is used to extract internal view features using an internal view feature extraction module. The structure of the internal view feature extraction module is the same as that of the inter-view feature extraction module, and the input dimension of the internal view feature extraction module is different from the input dimension of the inter-view feature extraction module.
6. The light field-based image processing method according to claim 1, characterized in that: The operations of extracting features between views, extracting features within views, and fusing feature information obtained by extracting features between views and extracting features within views are performed in a cycle more than twice.
7. A computer device, characterized in that The method comprises a processor and a memory, wherein the memory stores a computer program, and when the computer program is executed by the processor, the method implements each step of the light field-based image processing method according to any one of claims 1 to 6.
8. A computer-readable storage medium having a computer program stored thereon, characterized in that: When the computer program is executed by a processor, each step of the light field-based image processing method according to any one of claims 1 to 6 is implemented.