Stereoscopic endoscope device, image processing method, apparatus, and storage medium
The stereoscopic endoscope device, designed with optical lens groups and beam splitting prisms, combined with the super-resolution reconstruction technology of the image processing host, solves the contradiction between the bulky size of traditional 3D electronic endoscopes and high-definition stereoscopic visual imaging, realizing the miniaturization of the endoscope and high-quality stereoscopic visual imaging.
Patent Information
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-11-27
- Publication Date
- 2026-03-17
AI Technical Summary
Existing 3D electronic endoscopes require the integration of two optical links, resulting in an excessively large diameter, making it difficult to simultaneously meet the requirements of high-definition stereoscopic vision imaging and miniaturization.
The design employs an optical lens group and a beam splitter to divide a single light beam into two optical paths. Image signals are acquired by dual image sensors and then reconstructed in super-resolution by an image processing host to generate a stereo image.
It significantly improves image detail and stereoscopic imaging clarity without adding optical components, resolving the contradiction between the complexity and miniaturization of traditional endoscopes.
Smart Images

Figure CN121667595A_ABST
Abstract
Description
TECHNICAL FIELD
[0001] The present application relates to the technical field of endoscopes, in particular to a stereoscopic endoscope device, an image processing method, equipment and a storage medium. BACKGROUND
[0002] In the field of medical endoscopes, achieving high-definition stereoscopic vision imaging is crucial for improving surgical precision. Currently, the 3D electronic endoscopes widely used in clinical practice mainly adopt a scheme of placing two independent imaging systems (double objectives and double sensors) side by side at the front end of the scope to simulate the interpupillary distance of the human eye and obtain stereoscopic images. However, this traditional double-light-path physical imaging scheme requires the integration of two optical links, resulting in an excessively large diameter of the scope, which conflicts with the demand for miniaturized instruments in clinical practice, limiting the applicability of endoscopes in different surgical procedures.
[0003] Therefore, how to meet the demand for miniaturized endoscopes in clinical practice while ensuring high-quality stereoscopic vision imaging is a problem that needs to be solved at present. SUMMARY
[0004] The main purpose of the present application is to provide a stereoscopic endoscope device, an image processing method, equipment and a storage medium, which aims to ensure the clarity of stereoscopic vision imaging while solving the technical problem of complex endoscope structure.
[0005] To achieve the above-mentioned purpose, the present application further provides a stereoscopic endoscope device, which comprises an optical lens group, a light-splitting prism, a first image sensor, a second image sensor and a wire interface.
[0006] The optical lens group is used to collect the reflected light of the target observation object and process it into a single light beam for emission.
[0007] The light-splitting prism is used to split the incident single light beam into a first light path and a second light path.
[0008] The first image sensor is arranged on the propagation path of the first light path and is used to receive the first light path and convert it into a first image signal.
[0009] The second image sensor is arranged on the propagation path of the second light path and is used to receive the second light path and convert it into a second image signal.
[0010] The wire interface is electrically connected to the image processing host at the rear end, and is used to transmit the first image signal and the second image signal to the image processing host, so that the image processing host generates a stereoscopic image according to the first image signal and the second image signal.
[0011] Furthermore, to achieve the above objectives, this application also provides an image processing method applied to an image processing host, the image processing host being electrically connected to the rear end of the stereoscopic endoscope device as described above, the image processing method comprising:
[0012] Acquire the first image signal and the second image signal output by the stereoscopic endoscope device;
[0013] Super-resolution reconstruction is performed on the first image signal and the second image signal to obtain a third image signal and a fourth image signal, respectively, wherein the resolution of the third image signal and the fourth image signal is greater than the resolution of the first image signal and the second image signal, respectively.
[0014] A stereoscopic image is generated based on the third image signal and the fourth image signal.
[0015] In one embodiment, the step of performing super-resolution reconstruction on the first image signal and the second image signal to obtain the third image signal and the fourth image signal respectively includes:
[0016] Image alignment and parallax correction are performed on the first image signal and the second image signal to obtain the registered dual-channel image signal;
[0017] The dual-channel image signals are reconstructed using a preset super-resolution reconstruction algorithm to generate a third image signal and a fourth image signal.
[0018] In one embodiment, the step of performing image alignment and parallax correction processing on the first image signal and the second image signal to obtain the registered dual-channel image signal includes:
[0019] Extract feature points from the first image signal and the second image signal;
[0020] Calculate the disparity map between the first image signal and the second image signal based on the feature points;
[0021] Based on the disparity map, the first image signal and the second image signal are image registered to obtain the registered dual-channel image signal.
[0022] In one embodiment, the step of performing super-resolution reconstruction on the dual-channel image signals based on a preset super-resolution reconstruction algorithm to generate a third image signal and the fourth image signal includes:
[0023] The dual-channel image signals are input into a pre-trained fast super-resolution convolutional neural network;
[0024] The fast super-resolution convolutional neural network is used to perform feature extraction, nonlinear mapping, and upsampling operations on the dual-channel image signals to generate a third image signal and a fourth image signal.
[0025] In one embodiment, the step of performing feature extraction, nonlinear mapping, and upsampling operations on the dual-channel image signals using the fast super-resolution convolutional neural network to generate the third image signal and the fourth image signal includes:
[0026] Feature extraction is performed on the dual-channel image signals to obtain a first number of initial feature maps;
[0027] The initial feature maps are subjected to dimensionality reduction processing, and the first number of initial feature maps are mapped to a second number of feature maps, wherein the second number is less than the first number;
[0028] The feature map after dimensionality reduction is subjected to multi-layer nonlinear mapping, and the feature map after nonlinear mapping is then subjected to dimensionality increase processing to restore the number of feature maps from the second number to the first number, thus obtaining the dimensionality-increased feature map.
[0029] The upsampled feature map is reconstructed to generate the third image signal and the fourth image signal.
[0030] In one embodiment, the specific calculation model for the step of extracting features from the dual-channel image signals to obtain a first number of initial feature maps is as follows:
[0031]
[0032] Where F1 represents the first number of initial feature maps, d represents the first quantity, and X represents the input dual-channel image signal. This indicates a convolutional layer using a 5×5 kernel and having d output channels. Let H represent the element-wise activation function, H represent the row pixels of the dual-channel image signal, W represent the column pixels of the dual-channel image signal, and R represent the set of real numbers.
[0033] In one embodiment, the step of performing dimensionality reduction processing on the initial feature maps, mapping the first number of initial feature maps to a second number of feature maps, wherein the second number is less than the first number, is specifically calculated using the following model:
[0034]
[0035] Where F2 represents the second number of feature maps, s represents the second quantity, and F1 represents the first number of initial feature maps. This indicates a convolutional layer using a 1×1 kernel and having s output channels. Let H represent the element-wise activation function, H represent the row pixels of the dual-channel image signal, W represent the column pixels of the dual-channel image signal, and R represent the set of real numbers.
[0036] In addition, to achieve the above objectives, this application also proposes an image processing apparatus, the apparatus comprising: a memory, a processor, and a computer program stored in the memory and executable on the processor, the computer program being configured to implement the steps of the image processing method as described above.
[0037] In addition, to achieve the above objectives, this application also provides a storage medium, which is a computer-readable storage medium, on which a program implementing an image processing method is stored, and the program implementing the image processing method is executed by a processor to implement the steps of the image processing method as described above.
[0038] This application provides an image processing method. The method involves acquiring reflected light from a target object using an optical lens group and processing it into a single beam. A beam splitter then divides the single beam into a first optical path and a second optical path, which are respectively converted into first and second image signals by a first image sensor and a second image sensor. These signals are transmitted to an image processing host via a ribbon cable interface. Further, stereoscopic image super-resolution reconstruction is performed on the two image signals, generating a stereoscopic image based on the higher-resolution third and fourth image signals. The configuration of the beam splitter and dual image sensors enables stereoscopic image acquisition under a single optical path, effectively simplifying the internal structure of the endoscope and reducing the complexity of the device. Simultaneously, the super-resolution reconstruction of the dual signals by the image processing host significantly improves image detail and stereoscopic imaging clarity without increasing optical components. This solves the technical problem of traditional endoscopes, which are complex in structure and difficult to balance imaging performance and miniaturization, while ensuring stereoscopic vision imaging quality. Attached Figure Description
[0039] The accompanying drawings, which are incorporated in and form part of this specification, illustrate embodiments consistent with this application and, together with the description, serve to explain the principles of this application.
[0040] To more clearly illustrate the technical solutions in the embodiments of this application or the prior art, the drawings used in the description of the embodiments or the prior art will be briefly introduced below. Obviously, for those skilled in the art, other drawings can be obtained based on these drawings without creative effort.
[0041] Figure 1 This is a schematic diagram of the structure of the stereoscopic endoscope device of this application;
[0042] Figure 2This is a schematic diagram of the imaging process of the stereoscopic endoscope device of this application;
[0043] Figure 3 This is a flowchart illustrating an embodiment of the image processing method of this application.
[0044] Figure 4 This is a schematic diagram of the device structure of the hardware operating environment involved in the image processing method in the embodiments of this application.
[0045] The purpose, features, and advantages of this application will be further explained in conjunction with the embodiments and with reference to the accompanying drawings. Detailed Implementation
[0046] It should be understood that the specific embodiments described herein are only used to explain the technical solutions of this application and are not intended to limit this application.
[0047] To better understand the technical solution of this application, a detailed description will be provided below in conjunction with the accompanying drawings and specific implementation methods.
[0048] Optical endoscopes and 3D electronic endoscopes have wide applications in the medical field, especially in endoscopic examinations. Optical endoscopes transmit images to a camera system for display through optical elements such as objective lenses and rod lenses. Their structure is relatively simple, with no front-end electronic components, resulting in low heat generation and good stability. 3D electronic endoscopes, on the other hand, convert light signals into electrical signals through objective lenses and image sensors (CMOS), which are then processed and displayed on a monitor. They provide binocular parallax, producing a strong sense of depth, greatly improving the precision of surgical procedures and reducing the risk of hand-eye coordination problems. Currently, both optical and electronic endoscopes can meet clinical needs, but certain limitations remain. For example, optical endoscopes are limited to 2D planar imaging and cannot provide depth information. Doctors must rely on experience with shadows, occlusions, and motion parallax to judge depth, making operation difficult. Furthermore, dual-optical-path physical imaging schemes require integrating two independent imaging links (objective lens and image sensor) at the front end of the endoscope, achieving binocular parallax through physical separation. However, this scheme requires strict synchronous assembly and adjustment of the two links, resulting in low yield rates and high costs. Traditional solutions for 3D electronic endoscopes require integrating complex optical components (dual image sensors) at the front end of the endoscope, resulting in a bulky structure, extremely high assembly precision requirements, and a larger package size for a 4K resolution image sensor compared to a 1080P image sensor, making integration at the front end impossible and more expensive. Many 3D electronic endoscope products only improve resolution through simple interpolation algorithms, failing to truly restore details. Therefore, how to meet the needs of miniaturized clinical endoscopes while ensuring high-quality stereoscopic imaging is a pressing issue that needs to be addressed.
[0049] This application acquires reflected light from the target object using an optical lens group and processes it into a single beam. A beam splitter then divides the single beam into a first optical path and a second optical path, which are converted into first and second image signals by a first image sensor and a second image sensor, respectively. These signals are transmitted to an image processing host via a ribbon cable interface. Further, stereoscopic image super-resolution reconstruction is performed on the two image signals, generating a stereoscopic image based on the higher-resolution third and fourth image signals. The configuration of the beam splitter and dual image sensors enables stereoscopic image acquisition under a single optical path, effectively simplifying the internal structure of the endoscope and reducing the complexity of the device. Simultaneously, the super-resolution reconstruction of the dual signals by the image processing host significantly improves image detail and stereoscopic imaging clarity without increasing optical components. This solves the technical problem of traditional endoscopes, which, due to their complex structure, struggle to balance imaging performance and miniaturization while ensuring stereoscopic imaging quality.
[0050] This application provides a stereoscopic endoscope device. Please refer to... Figure 1 The stereoscopic endoscope device includes:
[0051] Optical lens assembly, beam splitter, first image sensor, second image sensor, and ribbon cable interface;
[0052] The optical lens group is used to collect the reflected light from the target object and process it into a single beam for emission.
[0053] The beam splitter is used to divide the incident single beam of light into a first optical path and a second optical path.
[0054] The first image sensor is disposed on the propagation path of the first optical path and is used to receive the first optical path and convert it into a first image signal;
[0055] The second image sensor is disposed on the propagation path of the second optical path and is used to receive the second optical path and convert it into a second image signal;
[0056] The ribbon cable interface is electrically connected to the image processing host at the rear end, and is used to transmit the first image signal and the second image signal to the image processing host so that the image processing host can generate a stereoscopic image based on the first image signal and the second image signal.
[0057] It should be noted that, referring to Figure 1The LED light source is located at the front end, emitting cool white light to illuminate the observed object, such as a cavity or tissue. In this embodiment, the optical lens group includes an objective lens, a rod lens, and an eyepiece. The objective lens can collect light reflected from the observed object over a wide range and form an initial optical image; the rod lens can transmit the initial image formed by the objective lens to the back end without loss through internal total internal reflection; the eyepiece performs preliminary magnification and aberration correction on the image transmitted from the rod lens, generating a high-quality intermediate image; the beam splitter uses thin-film interference or reflection / transmission principles to divide the single beam emitted from the eyepiece into a first optical path and a second optical path with different propagation directions according to a certain energy ratio; the first image sensor and the second image sensor are photoelectric conversion devices, receiving the first optical path and the second optical path respectively, and converting the optical signals into electrical signals, namely the first image signal and the second image signal; the ribbon cable interface leads out the electrical signals generated by the two image sensors and transmits them to the image processing host (not shown) at the back end.
[0058] Specifically, refer to Figure 2 The objective lens employs a low magnification design to achieve a large depth of field, ensuring clear imaging of both near and far objects; the rod lens ensures image fidelity after long-distance transmission; the eyepiece provides initial image magnification and optimization; the beam splitter can be used as... Figure 1 The special prism shown enables the acquisition of light signals from images in different directions from two angles, simulating binocular observation of objects by the human eye, ensuring that two images with a preset angle (e.g., simulating an 8° human interpupillary distance) can be received. The first image sensor and the second image sensor are preferably two high-resolution CMOS sensors with the same performance, such as 1080P resolution or higher, fixed on the two light-emitting surfaces of the beam splitter to receive the two light signals and convert the light signals into electrical signals, namely the first image signal and the second image signal. The ribbon cable interface can use a flexible circuit board (FPC) to connect the two first image sensors and the second image sensor, and output the combined signals to the image processing host.
[0059] This embodiment provides a stereoscopic endoscope device. By adopting a design where a single front-end optical link is divided into two optical paths by a beam splitter, this embodiment fundamentally avoids the problem of bulky endoscopes caused by physical space conflicts in traditional dual-objective, dual-sensor solutions, thus achieving a compact and miniaturized endoscope structure. At the same time, by using dual image sensors to receive two images with parallax respectively, a reliable image source is provided for the back-end to generate high-quality stereoscopic vision. Thus, while ensuring stereoscopic vision imaging capabilities, the contradiction between the complexity of the endoscope structure and the need for miniaturization is effectively resolved.
[0060] It should be noted that the execution subject in the embodiments below can be an image processing host, or a computing service device with data processing, network communication, and program execution functions, such as a tablet computer, personal computer, or mobile phone, or an image processing device capable of performing the above functions. This embodiment does not specifically limit it in this way. The following uses an image processing system as the execution subject as an example to describe this embodiment and the following embodiments.
[0061] Based on the first embodiment of this application, in the second embodiment of this application, the content that is the same as or similar to that in the first embodiment described above can be referred to the above description and will not be repeated hereafter. This application proposes an image processing method according to the second embodiment; please refer to... Figure 3 The image processing host is electrically connected to the rear end of the stereoscopic endoscope device as described in the first embodiment, and the image processing method includes steps S10 to S30:
[0062] Step S10: Acquire the first image signal and the second image signal output by the stereoscopic endoscope device;
[0063] Step S20: Perform super-resolution reconstruction on the first image signal and the second image signal to obtain a third image signal and a fourth image signal, respectively, wherein the resolution of the third image signal and the fourth image signal is greater than the resolution of the first image signal and the second image signal, respectively.
[0064] Step S30: Generate a stereoscopic image based on the third image signal and the fourth image signal.
[0065] It should be noted that in the stereoscopic endoscope device of Embodiment 1 described above, the image processing host is a hardware module that performs image algorithm processing. It can be a microprocessor (MCU), field-programmable gate array (FPGA), digital signal processor (DSP), or application-specific integrated circuit (ASIC), etc. Its function is to perform calculations on the input first and second image signals to improve their resolution and synthesize a stereoscopic image. Stereoscopic image super-resolution reconstruction can reconstruct a high-resolution image with richer details based on the complementary information between the low-resolution images (i.e., the first and second image signals).
[0066] Understandably, stereoscopic endoscopes can acquire dual low-resolution image signals for stereo vision in a space-saving manner. Super-resolution reconstruction technology compensates for the insufficient resolution (e.g., 1080P) of the stereoscopic endoscope's image signals. By utilizing the sub-pixel-level complementary information of the dual images, higher-resolution images (e.g., 4K) are generated at the algorithm level. This significantly improves the clarity of stereoscopic imaging without increasing the size of the front-end sensor or the diameter of the endoscope.
[0067] Specifically, steps S10-S30 are executed in the image processing host, where preprocessing (e.g., noise reduction, white balance) is performed on the dual signals (i.e., the first image signal and the second image signal), followed by stereo image super-resolution reconstruction. This reconstruction process can be based on traditional interpolation methods or a deep learning model (such as FSRCNN), which can effectively learn the mapping relationship from low resolution to high resolution. The image processing host outputs two high-resolution image signals (i.e., the third image signal and the fourth image signal), corresponding to the video streams for the left and right eyes respectively, and generates a stereo image based on the third and fourth image signals.
[0068] For example, two 1080P CMOS sensors generate a first image signal (left eye view) and a second image signal (right eye view), which are transmitted to the image processing motherboard via a ribbon cable interface. The image processing host on the motherboard loads a pre-trained super-resolution neural network model, processes the two 1080P signals, and outputs two third and fourth image signals with a resolution of 3840×2160 (4K), which are ultimately displayed on the monitor.
[0069] This embodiment provides an image processing method. By combining a stereo endoscope device with a stereo super-resolution reconstruction algorithm, the method first simplifies and miniaturizes the endoscope structure at the physical level by utilizing the beam splitter prism and dual sensors in the stereo endoscope device. Then, the image processing host performs algorithmic enhancement on the dual low-resolution signals to generate a higher resolution stereo image. This overcomes the shortcomings of traditional dual-optical-path schemes, such as complex structure and bulky endoscopes, in terms of hardware. It also compensates for the deficiencies of low-resolution sensors in terms of imaging effect. Finally, while ensuring the clarity of stereo vision imaging, it systematically solves the contradiction between miniaturization and high-definition of the endoscope.
[0070] In one feasible implementation, step S20 may include steps S21 to S22:
[0071] Step S21: Perform image alignment and parallax correction processing on the first image signal and the second image signal to obtain the registered dual-channel image signal;
[0072] Step S22: Perform super-resolution reconstruction on the dual-channel image signals based on a preset super-resolution reconstruction algorithm to generate a third image signal and the fourth image signal.
[0073] It should be noted that image alignment eliminates global translation and rotation errors caused by minor sensor positional deviations or timing differences between the first and second image signals. Parallax correction addresses pixel position differences related to scene depth caused by dual-viewpoint (simulated interpupillary distance). Its function is to correct the two images to the same imaging plane, ensuring that corresponding objects only have horizontal displacement, thus laying the foundation for subsequent fusion and reconstruction. After correction, the registered dual-channel image signals are obtained.
[0074] Understandably, due to minor physical installation errors in the two sensors behind the beam splitter prism, and the inherent parallax in the dual-channel images caused by simulated interpupillary distance, direct super-resolution reconstruction would lead to information fusion errors, resulting in ghosting or blurring. By employing image alignment and parallax correction, the correspondence between pixels in the dual-channel images is precisely established, eliminating interference from non-ideal factors and ensuring spatial consistency of the dual-channel image signals. Utilizing the complementary information with sub-pixel displacement provided by the registered dual-channel images, a super-resolution reconstruction algorithm is executed, enabling more accurate and effective recovery of high-frequency details and generating higher-resolution third and fourth image signals.
[0075] Specifically, global alignment can be achieved by calculating the homography matrix between two images, followed by calculating a dense disparity map using a stereo matching algorithm, and finally, image reprojection based on the disparity map to achieve disparity correction. Super-resolution reconstruction algorithms can employ interpolation-based methods (such as bicubic interpolation), reconstruction-based methods, or learning-based methods (such as convolutional neural networks). Deep learning-based methods are preferred because they can learn complex image priors, resulting in better reconstruction performance.
[0076] In this embodiment, by performing image alignment and disparity correction first, and then performing super-resolution reconstruction, the dual-path image information can be accurately and effectively fused and utilized. The correction step eliminates the negative impact of physical errors and disparity on the reconstruction process, providing high-quality input for the super-resolution algorithm, thereby significantly improving the overall quality and reliability of stereo vision imaging.
[0077] In one feasible implementation, step S21 may include steps S211 to S213:
[0078] Step S211: Extract feature points from the first image signal and the second image signal;
[0079] Step S212: Calculate the disparity map between the first image signal and the second image signal based on the feature points;
[0080] Step S213: Perform image registration on the first image signal and the second image signal according to the disparity map to obtain the registered dual-channel image signal.
[0081] Understandably, extracting stable and reliable feature points provides a data foundation for subsequent calculations. Utilizing the correspondence between these feature points, the disparity values of all pixels in the image can be further calculated using a stereo matching algorithm, generating a dense disparity map. This allows for precise quantification of the spatial correspondence between the two images. Based on the obtained disparity map, one image (usually the second image signal) is remapped or distorted at the pixel level to place it on the same projection plane as the other image (the first image signal), achieving precise registration and providing the necessary conditions for subsequent super-resolution reconstruction algorithms requiring high-precision pixel correspondence.
[0082] Specifically, feature extraction algorithms can employ Scale Invariant Feature Transform (SIFT), Speed-Up Robust Feature Transform (SURF), or Oriented Fast and Rotated BRIEF (ORB), etc. Disparity maps can be generated using local matching algorithms (such as block matching BM), global matching algorithms (such as graph cut), or semi-global matching algorithms (SGM). The second image signal is then subjected to a disparity map-based inverse mapping (or forward mapping with interpolation) to align its pixel positions with the first image signal.
[0083] In this embodiment, precise geometric correction of dual-channel stereo images is achieved through three steps: feature point extraction, disparity map calculation, and image registration. Feature points provide a stable matching basis, the disparity map fully describes the spatial differences between images, and the final registration operation eliminates these differences, ensuring the spatial consistency of the dual-channel image signals at the pixel level. This provides a crucial guarantee for effectively utilizing the complementary information of the dual-channel images in the subsequent super-resolution reconstruction process, contributing to the improvement of high-definition stereo image quality.
[0084] In one possible implementation, step S22 may include steps S221 to S222:
[0085] Step S221: Input the dual-channel image signals into a pre-trained fast super-resolution convolutional neural network;
[0086] Step S222: The fast super-resolution convolutional neural network is used to perform feature extraction, nonlinear mapping and upsampling operations on the dual-channel image signals to generate a third image signal and a fourth image signal.
[0087] Understandably, a pre-trained Fast Super-Resolution Convolutional Neural Network (FSRCNN model) is used to capture the essential information of the input dual-path images through feature extraction; the complex relationship between low-resolution and high-resolution features is learned through deep nonlinear mapping; and finally, a high-resolution pixel array is reconstructed through upsampling operations. The FSRCNN model can achieve real-time or near-real-time super-resolution processing with limited hardware resources, meeting the smoothness requirements of endoscopic surgery.
[0088] Specifically, the registered dual-channel images can be stitched together along the channel dimension and input as a single multi-channel FSRCNN; alternatively, the registered dual-channel images can be fed into two structurally identical FSRCNN branches for processing. FSRCNN uses small-sized convolutional kernels for fast feature extraction; 1x1 convolutions are used for feature dimensionality reduction to reduce computation; multiple 3x3 convolutions are then used for non-linear mapping; finally, deconvolutional layers (transposed convolutions) are used for upsampling, ultimately outputting reconstructed high-resolution image signals (third and fourth image signals).
[0089] In this embodiment, a pre-trained fast super-resolution convolutional neural network is used to process the registered dual-channel image signals. By leveraging the powerful nonlinear representation capabilities of deep learning models, rich high-frequency details and realistic texture information can be learned and reconstructed from low-resolution input. At the same time, its optimized network structure ensures processing speed and meets the real-time requirements of clinical practice, thereby improving image resolution at the algorithm level and enhancing the clarity and realism of stereo vision.
[0090] In one possible implementation, step S222 may include steps S2221 to S2224:
[0091] Step S2221: Extract features from the dual-channel image signals to obtain a first number of initial feature maps;
[0092] Step S2222: Perform dimensionality reduction processing on the initial feature map, mapping the first number of initial feature maps to a second number of feature maps, wherein the second number is less than the first number;
[0093] Step S2223: Perform multi-layer nonlinear mapping on the dimension-reduced feature map, and perform dimension-up processing on the feature map after nonlinear mapping to restore the number of feature maps from the second number to the first number, thereby obtaining the dimension-upgraded feature map;
[0094] Step S2224: Upsample and reconstruct the upgraded feature map to generate the third image signal and the fourth image signal.
[0095] It should be noted that the first quantity of initial feature maps refers to the number of feature maps output after passing through the feature extraction layer (denoted as d), which represents various different types of primary features extracted from the input image. The dimensionality reduction process refers to operations such as 1x1 convolution, which, without changing the spatial dimensions of the feature maps, reduces the number of channels of the feature maps (from d to s, where s < d), in order to compress the feature representation and significantly reduce the number of parameters and computational complexity in subsequent calculations. The second quantity of feature maps is the feature maps after dimensionality reduction (with a quantity of s). The multi-layer non-linear mapping performs multiple convolution and non-linear activation operations in the low-dimensional feature space after dimensionality reduction, and can learn complex feature transformations. The dimensionality increase process is an operation opposite to dimensionality reduction, which restores the number of channels of the feature maps from s to d, preparing a suitable feature dimension for the final upsampling reconstruction. The upsampling reconstruction enlarges the spatial dimensions of the feature maps to the target high-resolution through operations such as deconvolution or sub-pixel convolution.
[0096] In a feasible implementation manner, the specific calculation model in step S2221 is as follows:
[0097]
[0098] Among them, F1 represents the first quantity of initial feature maps, d represents the first quantity, X represents the input dual-channel image signal, represents a convolutional layer using a 5×5 convolutional kernel with an output channel number of d, represents an element-wise activation function, H represents the number of row pixels of the dual-channel image signal, W represents the number of column pixels of the dual-channel image signal, and R represents the set of real numbers.
[0099] It should be noted that F1 represents the first quantity of initial feature maps, X represents the input dual-channel image signal, and its shape is [H×W×C in , where C in is the number of input channels, represents a two-dimensional convolution operation that uses d 5x5 convolutional kernels. Each kernel performs a sliding window calculation on all input channels of the input tensor X to generate a two-dimensional feature map, and a total of d feature maps are generated by d kernels. represents an element-wise non-linear activation function, which can be PReLU (Parametric RectifiedLinearUnit, linear rectified function) or ReLU (RectifiedLinearUnit, parametric rectified linear function), introducing non-linear transformation capabilities to the model so that it can fit more complex functions. F1 represents the output initial feature map tensor with a shape of [H×W×d]. R represents the set of real numbers, R H×W×dLet H represent a three-dimensional tensor of dimension H×W×d, where each element of the three-dimensional tensor belongs to the set of real numbers. By defining a feature extraction computation model using 5x5 convolutional kernels and nonlinear activation functions, the network is able to capture rich initial image features with a large receptive field in the first layer, providing a high-quality and comprehensive feature foundation for the entire super-resolution reconstruction process. The larger convolutional kernel size helps the network better understand the overall structure of the image in the early stages of reconstruction, which plays an important role in subsequent detail recovery and maintaining the authenticity of image structure.
[0100] In one feasible implementation, the specific calculation model for step S2222 is as follows:
[0101]
[0102] Where F2 represents the second number of feature maps, s represents the second quantity, and F1 represents the first number of initial feature maps. This indicates a convolutional layer using a 1×1 kernel and having s output channels. Let H represent the element-wise activation function, H represent the row pixels of the dual-channel image signal, W represent the column pixels of the dual-channel image signal, and R represent the set of real numbers.
[0103] It should be noted that F1 is the output of the previous layer, i.e., the initial feature map, with a shape of [H×W×d]. F2 represents the second number of feature maps. This represents a 2D convolution operation using a 1x1 kernel and s output channels. The special characteristic of a 1x1 convolution is that it does not change the spatial dimensions of the feature map (H and W remain unchanged), only the channel dimensions. Its function is to linearly combine (weighted sum) the feature information from d channels, thereby projecting the d-dimensional features onto a new s-dimensional feature space (sd). <d)。 F2 is an element-wise nonlinear activation function that introduces nonlinear transformation capabilities into the model, enabling it to fit more complex functions. F2 is the feature map tensor output after dimensionality reduction, with a shape of [H×W×s]. R represents the set of real numbers, R0... H×W×s Let H represent a three-dimensional tensor with dimensions H×W×s, where each element of the three-dimensional tensor belongs to the set of real numbers. By defining a specific computational model for dimensionality reduction using 1x1 convolutions, core features are extracted and the feature dimension is significantly reduced through linear combinations and nonlinear activations between channels without changing the spatial information of the feature map. This enables complex super-resolution reconstruction networks to run in real time with limited hardware resources, meeting the stringent real-time requirements of endoscopic surgery.
[0104] In one feasible implementation, the specific calculation model for step S2223 is as follows:
[0105] Mapping: Let G0 = F2, for j = 1, ..., m:
[0106] The mapping output is:
[0107] F3 = G m ∈R H×W×s
[0108] Expanding:
[0109] F4∈R H×W×d
[0110] It should be noted that mapping refers to the process of performing deep feature transformation on the dimensionality-reduced feature map, using m layers of 3×3 convolutional kernels, with each layer maintaining s feature maps. G0 = F2 indicates that the initial input to the mapping process is the dimensionality-reduced feature map F2. The calculation method for the j-th layer mapping is defined, indicating a convolutional layer using a 3×3 kernel. G represents the element-wise activation function (such as PReLU), m represents the number of mapping layers, and G... m This is the final output feature map after m layers of mapping. Expanding is the operation corresponding to the aforementioned dimensionality reduction process. It uses 1×1 convolutional kernels to restore the feature map size from s to d, preparing a suitable feature dimension for subsequent upsampling reconstruction. Although the dimensionality-reduced feature map is computationally efficient, its feature representation ability is relatively limited. Through multiple layers of 3×3 convolutions in the mapping stage, a deep nonlinear transformation is performed. Each convolutional layer expands the receptive field and learns more complex feature combinations. The stacking of m layers allows the network to establish a complex mapping relationship from low-resolution features to high-resolution features. The subsequent expansion stage projects the deeply processed low-dimensional features back to the high-dimensional space through 1×1 convolutions, restoring the richness of the features and matching the feature dimension with the initial stage of the network, providing a sufficient information foundation for the final high-resolution image reconstruction.
[0111] The specific calculation model for step S2224 is as follows:
[0112] Deconvolution:
[0113]
[0114] Where Y is the output image (tH×tW×Cout, t is the upsampling factor), This is an image with Cout output channels, generated using a 9x9 deconvolution.
[0115] It should be noted that deconvolution is an upsampling operation that enlarges the spatial size of a low-resolution feature map to the target high-resolution size. The 9×9 deconvolution kernel is the size of the upsampling filter; a larger kernel size helps generate a smoother, more detailed high-resolution image. Y represents the final output image, i.e., the third and fourth image signals, with dimensions tH×tW×Cout, where t is the upsampling factor (e.g., 4x) and Cout is the number of output channels. This represents an image with Cout output channels using a 9×9 transposed convolution kernel. σ is an element-wise nonlinear activation function, introducing nonlinear transformation capabilities to the model, enabling it to fit more complex functions. After feature extraction, dimensionality reduction, mapping, and dimensionality upscaling, the network has obtained feature maps containing rich high-resolution information, but these feature maps are still original low-resolution in spatial dimensions. By defining a computational model using a 9×9 deconvolution kernel for final upsampling reconstruction, the spatial transformation from low-resolution features to high-resolution pixels is completed. The larger deconvolution kernel size ensures that a wider range of contextual information can be used for pixel value prediction during the upsampling process, thereby generating high-quality images with rich details and sharp edges.
[0116] In this embodiment, the super-resolution reconstruction process includes feature extraction, dimensionality reduction, nonlinear mapping, dimensionality upscaling, and upsampling reconstruction operations. The dimensionality reduction and upscaling operations effectively control the complexity of the model and the computational cost, making it possible to run in real time on a low-computing-power platform. The intermediate multi-layer nonlinear mapping ensures that the model can learn the complex mapping relationship from low resolution to high resolution, thereby achieving high-quality reconstruction of image details while taking into account processing speed, and meeting the stringent requirements of endoscope systems for real-time high-definition imaging.
[0117] This application provides an image processing apparatus, which includes: at least one processor; and a memory communicatively connected to the at least one processor; wherein the memory stores instructions executable by the at least one processor, which are executed by the at least one processor to enable the at least one processor to perform the image processing method in Embodiment 2 described above.
[0118] The following is for reference. Figure 4This document illustrates a structural schematic diagram of an image processing device suitable for implementing embodiments of this application. The image processing device in the embodiments of this application may include, but is not limited to, mobile terminals such as mobile phones, laptops, digital broadcast receivers, PDAs (Personal Digital Assistants), PADs (Portable Application Description), PMPs (Portable Media Players), in-vehicle terminals (e.g., in-vehicle navigation terminals), and fixed terminals such as digital TVs and desktop computers. Figure 4 The image processing device shown is merely an example and should not impose any limitations on the functionality and scope of use of the embodiments of this application.
[0119] like Figure 4 As shown, the image processing device may include a processing unit 1001 (e.g., a central processing unit, a graphics processing unit, etc.) that can perform various appropriate actions and processes according to a program stored in a read-only memory (ROM) 1002 or a program loaded from a storage device 1003 into a random access memory (RAM) 1004. The RAM 1004 also stores various programs and data required for the operation of the image processing device. The processing unit 1001, ROM 1002, and RAM 1004 are interconnected via a bus 1005. An input / output (I / O) interface 1006 is also connected to the bus. Typically, the following systems can be connected to the I / O interface 1006: input devices 1007 including, for example, a touchscreen, touchpad, keyboard, mouse, image sensor, microphone, accelerometer, gyroscope, etc.; output devices 1008 including, for example, a liquid crystal display (LCD), speaker, vibrator, etc.; storage devices 1003 including, for example, magnetic tape, hard disk, etc.; and communication devices 1009. Communication device 1009 allows the image processing device to communicate wirelessly or wiredly with other devices to exchange data. Although the figure shows an image processing device with various systems, it should be understood that it is not required to implement or possess all of the systems shown. More or fewer systems may be implemented alternatively.
[0120] Specifically, according to the embodiments disclosed in this application, the processes described above with reference to the flowcharts can be implemented as computer software programs. For example, embodiments disclosed in this application include a computer program product comprising a computer program carried on a computer-readable medium, the computer program containing program code for performing the methods shown in the flowcharts. In such embodiments, the computer program can be downloaded and installed from a network via a communication device, or installed from storage device 1003, or installed from ROM 1002. When the computer program is executed by processing device 1001, it performs the functions defined in the methods of the embodiments disclosed in this application.
[0121] The image processing device provided in this application, employing the image processing method described in the above embodiments, can ensure the clarity of stereoscopic vision imaging while solving the technical problem of complex endoscope structures. Compared with the prior art, the beneficial effects of the image processing device provided in this application are the same as those of the image processing method described in the above embodiments, and other technical features of this image processing device are the same as those disclosed in the previous embodiment method, and will not be repeated here.
[0122] It should be understood that the various parts disclosed in this application can be implemented using hardware, software, firmware, or a combination thereof. In the description of the above embodiments, specific features, structures, materials, or characteristics can be combined in any suitable manner in one or more embodiments or examples.
[0123] The above description is merely a specific embodiment of this application, but the scope of protection of this application is not limited thereto. Any variations or substitutions that can be easily conceived by those skilled in the art within the scope of the technology disclosed in this application should be included within the scope of protection of this application. Therefore, the scope of protection of this application should be determined by the scope of the claims.
[0124] This application provides a computer-readable storage medium having computer-readable program instructions (i.e., a computer program) stored thereon, the computer-readable program instructions being used to execute the image processing method described in the above embodiments.
[0125] The computer-readable storage medium provided in this application may be, for example, a USB flash drive, but is not limited to, electrical, magnetic, optical, electromagnetic, infrared, or semiconductor systems, devices, or any combination thereof. More specific examples of computer-readable storage media may include, but are not limited to: electrical connections having one or more wires, portable computer disks, hard disks, random access memory (RAM), read-only memory (ROM), erasable programmable read-only memory (EPROM or flash memory), optical fiber, portable compact disk read-only memory (CD-ROM), optical storage devices, magnetic storage devices, or any suitable combination thereof. In this embodiment, the computer-readable storage medium may be any tangible medium containing or storing a program that can be used by or in conjunction with an instruction execution system, system, or device. The program code contained on the computer-readable storage medium may be transmitted using any suitable medium, including but not limited to: wires, optical cables, RF (Radio Frequency), etc., or any suitable combination thereof.
[0126] The aforementioned computer-readable storage medium may be included in an image processing device or may exist independently without being assembled into an image processing device.
[0127] The aforementioned computer-readable storage medium carries one or more programs that, when executed by an image processing device, cause the image processing device to: acquire the first image signal and the second image signal output by the stereoscopic endoscope device; perform super-resolution reconstruction on the first image signal and the second image signal to obtain a third image signal and a fourth image signal, wherein the resolution of the third image signal and the fourth image signal is greater than the resolution of the first image signal and the second image signal, respectively; and generate a stereoscopic image based on the third image signal and the fourth image signal.
[0128] Computer program code for performing the operations of this application can be written in one or more programming languages or a combination thereof, including object-oriented programming languages such as Java, Smalltalk, and C++, and conventional procedural programming languages such as the "C" language or similar programming languages. The program code can be executed entirely on the user's computer, partially on the user's computer, as a standalone software package, partially on the user's computer and partially on a remote computer, or entirely on a remote computer or server. In cases involving remote computers, the remote computer can be connected to the user's computer via any type of network—including a local area network (LAN) or a wide area network (WAN)—or can be connected to an external computer (e.g., via the Internet using an Internet service provider).
[0129] The flowcharts and block diagrams in the accompanying drawings illustrate the architecture, functionality, and operation of possible implementations of systems, methods, and computer program products according to various embodiments of this application. In this regard, each block in a flowchart or block diagram may represent a module, segment, or portion of code containing one or more executable instructions for implementing a specified logical function. It should also be noted that in some alternative implementations, the functions indicated in the blocks may occur in a different order than those indicated in the drawings. For example, two consecutively indicated blocks may actually be executed substantially in parallel, and they may sometimes be executed in reverse order, depending on the functions involved. It should also be noted that each block in the block diagrams and / or flowcharts, and combinations of blocks in the block diagrams and / or flowcharts, can be implemented using a dedicated hardware-based system that performs the specified function or operation, or using a combination of dedicated hardware and computer instructions.
[0130] The modules described in the embodiments of this application can be implemented in software or hardware. The names of the modules do not necessarily limit the functionality of the unit itself.
[0131] The readable storage medium provided in this application is a computer-readable storage medium that stores computer-readable program instructions (i.e., a computer program) for executing the above-described image processing method. This ensures the clarity of stereoscopic vision imaging while solving the technical problem of the complex structure of endoscopes. Compared with the prior art, the beneficial effects of the computer-readable storage medium provided in this application are the same as those of the image processing method provided in the above embodiments, and will not be repeated here.
[0132] The above are merely preferred embodiments of this application and do not limit the patent scope of this application. Any equivalent structural or procedural transformations made using the content of this application's specification and drawings, or direct or indirect applications in other related technical fields, are similarly included within the patent scope of this application.
Claims
1. A stereoscopic endoscope device, characterized by comprising: The stereoscopic endoscope device comprises an optical lens group, a light splitting prism, a first image sensor, a second image sensor and a wire interface; The optical lens group is used for collecting reflected light of a target object and processing it into a single light beam; The light splitting prism is used for splitting the incident single light beam into a first light path and a second light path; The first image sensor is arranged on the propagation path of the first light path and is used for receiving the first light path and converting it into a first image signal; The second image sensor is arranged on the propagation path of the second light path and is used for receiving the second light path and converting it into a second image signal; The wire interface is electrically connected to an image processing host at the rear end and is used for transmitting the first image signal and the second image signal to the image processing host, so that the image processing host generates a stereoscopic image according to the first image signal and the second image signal.
2. An image processing method, characterized by, The image processing method is applied to an image processing host which is electrically connected to the rear end of the stereoscopic endoscope device as claimed in claim 1, and the image processing method comprises: obtaining the first image signal and the second image signal output by the stereoscopic endoscope device; performing super-resolution reconstruction on the first image signal and the second image signal to obtain a third image signal and a fourth image signal respectively, wherein the resolutions of the third image signal and the fourth image signal are greater than the resolutions of the first image signal and the second image signal respectively; generating a stereoscopic image according to the third image signal and the fourth image signal.
3. The method of claim 2, wherein, The step of performing super-resolution reconstruction on the first image signal and the second image signal to obtain a third image signal and a fourth image signal respectively comprises: performing image alignment and disparity correction processing on the first image signal and the second image signal to obtain a registered dual-channel image signal; performing super-resolution reconstruction on the dual-channel image signal based on a preset super-resolution reconstruction algorithm to generate a third image signal and a fourth image signal.
4. The method of claim 3, wherein, The step of performing image alignment and disparity correction processing on the first image signal and the second image signal to obtain a registered dual-channel image signal comprises: extracting feature points of the first image signal and the second image signal; calculating a disparity map between the first image signal and the second image signal according to the feature points; performing image registration on the first image signal and the second image signal according to the disparity map to obtain a registered dual-channel image signal.
5. The method of claim 3, wherein, The step of performing super-resolution reconstruction on the dual-channel image signal based on a preset super-resolution reconstruction algorithm to generate a third image signal and a fourth image signal comprises: inputting the dual-channel image signal into a pre-trained fast super-resolution convolutional neural network; performing feature extraction, non-linear mapping and up-sampling operations on the dual-channel image signal through the fast super-resolution convolutional neural network to generate a third image signal and a fourth image signal.
6. The method of claim 5, wherein, The step of performing feature extraction, non-linear mapping and up-sampling operation on the dual-path image signal by the fast super-resolution convolutional neural network to generate the third image signal and the fourth image signal comprises: performing feature extraction on the dual-path image signal to obtain a first number of initial feature maps; performing dimension reduction processing on the initial feature maps to map the first number of initial feature maps to a second number of feature maps, wherein the second number is less than the first number; performing multi-layer non-linear mapping on the dimension-reduced feature maps, and performing dimension increasing processing on the non-linearly mapped feature maps to restore the number of feature maps from the second number to the first number to obtain dimension-increased feature maps; performing up-sampling reconstruction on the dimension-increased feature maps to generate the third image signal and the fourth image signal.
7. The method of claim 6, wherein, The step of performing feature extraction on the dual-path image signal to obtain a first number of initial feature maps comprises a specific calculation model: wherein F1 represents a first number of initial feature maps, d represents the first number, X represents an input dual-path image signal, represents a convolution layer using a 5x5 convolution kernel with an output channel number of d, represents an element-wise activation function, H represents a row pixel of the dual-path image signal, W represents a column pixel of the dual-path image signal, and R represents a set of real numbers.
8. The method of claim 6, wherein, The step of performing dimension reduction processing on the initial feature maps to map the first number of initial feature maps to a second number of feature maps, wherein the second number is less than the first number, comprises a specific calculation model: wherein F2 represents a second number of feature maps, s represents the second number, and F1 represents a first number of initial feature maps, represents a convolution layer using a 1x1 convolution kernel with an output channel number of s, represents an element-wise activation function, H represents a row pixel of the dual-path image signal, W represents a column pixel of the dual-path image signal, and R represents a set of real numbers.
9. An image processing apparatus characterized by comprising: The image processing device comprises a memory, a processor and a computer program stored on the memory and executable on the processor, and the computer program is configured to implement the steps of the image processing method according to any one of claims 2 to 8.
10. A storage medium, characterized by The storage medium is a computer readable storage medium, and the computer readable storage medium stores a computer program, and the computer program is executed by the processor to implement the steps of the image processing method according to any one of claims 2 to 8.