Binocular Disparity Estimation Method, Visual Prosthesis, and Computer-Readable Storage Medium

Through the combination of binocular parallax estimation method and visual prosthesis, the parallax map is accurately calculated and the electrical stimulation pulse signal is generated, which solves the problem of blind patients identifying and avoiding obstacles in complex environments and improves the safety of movement.

CN115998591BActive Publication Date: 2025-07-08INTELLIMICRO MEDICAL CO LTD
View PDF 2 Cites 0 Cited by

Patent Information

Application Number
CN202211521138.7
Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Filing Date
2022-11-30
Publication Date
2025-07-08
Estimated Expiration
2042-11-30

AI Technical Summary

Technical Problem

The prior art cannot effectively extract and provide obstacle information in the surrounding environment of blind patients, resulting in safety hazards.

Method used

The binocular parallax estimation method is used to acquire environmental images through a binocular camera, perform deep feature extraction and matching and fusion, and generate electrical stimulation pulse signals of visual prosthesis to assist blind people in identifying and avoiding obstacles.

Benefits of technology

It significantly improves the accuracy and effectiveness of obstacle identification and avoidance information transmitted by visual prosthesis to patients, and improves the mobility of blind patients.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN115998591B_ABST
    Figure CN115998591B_ABST
Patent Text Reader

Abstract

The present invention discloses a binocular disparity estimation method, a visual prosthesis and a computer-readable storage medium. The binocular disparity estimation method includes: acquiring a first image and a second image of the surrounding environment of a visual prosthesis wearer collected by a binocular camera; performing depth feature extraction and matching fusion on the first image and the second image to obtain a feature map; performing disparity estimation based on the feature map to obtain a target disparity map for generating an electrical stimulation pulse signal of the visual prosthesis. Combining the binocular disparity estimation method with the visual prosthesis can effectively extract obstacle information in the surrounding environment of the blind and provide it to blind patients, assisting the patients to identify and avoid various obstacles in the living scene and improving their mobility.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention relates to the technical field of medical devices, and in particular, to a binocular disparity estimation method, a visual prosthesis, and a computer-readable storage medium. Background Art

[0002] A visual prosthesis is a new type of medical device that induces phosphenes in blind patients by applying a stimulating current to the retina or visual cortex to form a visual perception. According to the different implantation positions of the stimulating electrodes, it can be divided into a retinal stimulation visual prosthesis (also known as an "implantable retinal stimulator") and a cerebral cortex stimulation visual prosthesis.

[0003] In the related art, after a retinal stimulation visual prosthesis obtains an image of the patient's surrounding environment through an external camera, it uses a conventional image processing algorithm to enhance the image and then converts it into an electrical stimulation signal, which is sent to an electrode array implanted on the patient's retina. The retinal cells are electrically stimulated to achieve visual perception and reconstruction.

[0004] However, in the actual daily life scenario, when a blind patient moves, there are often various obstacles around them at the same time. Currently, conventional image processing algorithms cannot effectively extract and display the obstacles in life and the distance between them and the obstacles to the patient in a complex scenario. Therefore, it cannot effectively assist the patient in identifying and avoiding various obstacles in the life scenario, and there are potential safety hazards. Summary of the Invention

[0005] The present invention aims to at least solve one of the technical problems existing in the prior art. For this reason, the first object of the present invention is to propose a binocular disparity estimation method, which can effectively extract the obstacle information in the surrounding environment of the blind and provide it to the blind patient, assist the patient in identifying and avoiding various obstacles in the life scenario, and improve their mobility.

[0006] The second object of the present invention is to propose a visual prosthesis.

[0007] The third object of the present invention is to propose a computer-readable storage medium.

[0008] To achieve the above object, the binocular disparity estimation method according to the first aspect embodiment of the present invention acquires a first image and a second image of the surrounding environment of the wearer of the visual prosthesis collected by a binocular camera; performs depth feature extraction and matching fusion on the first image and the second image to obtain a feature map; and performs disparity estimation according to the feature map to obtain a target disparity map for generating an electrical stimulation pulse signal of the visual prosthesis.

[0009] The binocular disparity estimation method according to the embodiments of the present invention can accurately calculate a disparity map, effectively extract obstacle information in a complex environment, and generate a target disparity map for generating an electrical stimulation pulse signal of a visual prosthesis. That is, by combining the binocular disparity estimation method with the electrical stimulation of the visual prosthesis, the accuracy and effectiveness of the obstacle recognition and avoidance information assistance transmitted by the visual prosthesis to the patient can be significantly improved, and the patient's mobility can be enhanced.

[0010] In some embodiments, performing depth feature extraction and matching fusion on the first image and the second image to obtain a feature map includes: using the first image and the second image as inputs, respectively extracting first image features in the first image and second image features in the second image through MobileNet in a manner of sharing weights; performing feature matching on the first image features and the second image features according to a set of preset disparity values to obtain a plurality of matching results; and fusing the plurality of matching results to obtain a set of feature maps.

[0011] In some embodiments, before using the first image and the second image as inputs, respectively extracting first image features in the first image and second image features in the second image through MobileNet in a manner of sharing weights, the control method further includes: performing distortion correction and stereo correction on the first image and the second image so that the matching points of the corrected first image and the second image are on the same pixel row.

[0012] In some embodiments, performing disparity estimation according to the feature map to obtain a target disparity map for generating an electrical stimulation pulse signal of a visual prosthesis includes: inputting the set of feature maps into a disparity estimation network to obtain an initial disparity map, and inputting the set of feature maps into a weight estimation network to obtain an attention weight; performing weighted fusion on the initial disparity map, the attention weight, and the feature map, and inputting the result into a global information optimization network; continuously optimizing the initial disparity map using the context time series relationship to obtain a final predicted disparity map; and downsampling the final predicted disparity map according to the resolution set by the visual prosthesis to obtain the target disparity map.

[0013] In some embodiments, inputting the set of feature maps into a disparity estimation network to obtain an initial disparity map includes: serializing the set of feature maps to obtain a serialized feature map; using a depth feature transformation network to model the dependencies of the input-output sequence to determine a feature mapping relationship; and aggregating the serialized feature map according to the feature mapping relationship through feature mapping to estimate the initial disparity map.

[0014] In some embodiments, inputting the set of feature maps into a weight estimation network to obtain attention weights includes: using the set of feature maps as input information; using a multi-scale dilated convolutional pyramid module to construct a matching cost by aggregating environmental information of different sizes and different positions; adjusting the matching cost by fusing and stacking multiple hourglass networks through 3D convolution; and outputting the attention weights through the SoftMax layer of the weight estimation network.

[0015] In some embodiments, using the context temporal relationship to optimize the initial disparity map for consecutive frames to obtain a final predicted disparity map includes: respectively performing feature extraction based on MobileNet and feature matching between consecutive frames on the first image and the second image of consecutive K frames; solving the camera pose transformation relationship of each frame image in the relative coordinate system based on the Perspective-n-Point algorithm; and using the projection between frames to complete the missing area to obtain the final predicted disparity map.

[0016] In some embodiments, using the projection between frames to complete the missing area to obtain the final predicted disparity map includes: determining the missing area of the image; performing superpixel segmentation of the missing area at a preset resolution with the nearest matching point in the missing area as the center, such that the missing area is included in the mask; and using the valid disparity within the adjacent frame masks based on the projection relationship of the nearest matching point between consecutive frames to perform local disparity completion of the missing area to obtain the final predicted disparity map.

[0017] To achieve the above object, the visual prosthesis according to the second aspect embodiment of the present invention includes: an implant device and a wireless signaler, the implant device being connected to the wireless signaler; a camera unit for collecting images of the wearer's surrounding environment; and an artificial intelligence image processing unit, the artificial intelligence image processing unit being connected to the camera unit and the wireless signaler, for obtaining a target disparity map according to the binocular disparity estimation method described above, and sending the target disparity map to the implant device through the wireless signaler, and the implant device generating an electrical stimulation pulse signal according to the target disparity map.

[0018] According to the visual prosthesis of the embodiment of the present invention, the artificial intelligence image processing unit executes the binocular disparity estimation method of the above embodiment to obtain a target disparity map, that is, applying the binocular disparity estimation method of the above embodiment to the visual prosthesis can effectively extract obstacle information in a complex environment, and can significantly improve the accuracy and effectiveness of the visual prosthesis in transmitting obstacle recognition and avoidance information assistance to patients, and improve the patient's mobility.

[0019] In some embodiments, the camera unit includes at least one set of binocular cameras.

[0020] To achieve the above object, an embodiment of the third aspect of the present invention further provides a computer-readable storage medium, on which a computer program is stored. When the computer program is executed by a processor, the binocular disparity estimation method described above is implemented.

[0021] Additional aspects and advantages of the present invention will be given in part in the following description, become apparent in part from the following description, or be learned through the practice of the present invention. Description of the Drawings

[0022] The above and / or additional aspects and advantages of the present invention will become apparent and be readily understood from the description of the embodiments in conjunction with the following drawings, where:

[0023] Figure 1 is a block diagram of a visual prosthesis according to an embodiment of the present invention;

[0024] Figure 2 is a flowchart of a binocular disparity estimation method according to an embodiment of the present invention;

[0025] Figure 3 is a flowchart of a binocular disparity estimation method according to another embodiment of the present invention;

[0026] Figure 4 is a schematic diagram of the working process of a visual prosthesis according to an embodiment of the present invention. Detailed Embodiments

[0027] The embodiments of the present invention will be described in detail below. The embodiments described with reference to the drawings are exemplary. The embodiments of the present invention will be described in detail below.

[0028] Combining the binocular disparity estimation method of the embodiments of the present invention with a visual prosthesis can significantly improve the accuracy and effectiveness of the visual prosthesis in transmitting obstacle recognition and avoidance information assistance to patients.

[0029] Figure 1 is a block diagram of a visual prosthesis according to an embodiment of the present invention, as Figure 1 shown, the visual prosthesis 100 includes an implant device 10, a wireless signaler 20, a camera unit 30, and an artificial intelligence image processing unit 40.

[0030] Among them, the imaging unit 30 is worn on the patient and can be used to collect image information of the wearer's surrounding environment. In some embodiments, the imaging unit 30 may include at least one set of binocular cameras. The artificial intelligence image processing unit 40 may be an intelligent computing terminal such as a smart phone, a tablet, etc., and is used to perform binocular disparity estimation on the images collected by the imaging unit 20 to obtain a target disparity map, and send the target disparity map to the implant device 10 through the wireless signaler 20. The electrode array of the implant device 10 is implanted into the visual cortex cells or retinal cells of the patient, and can generate an electrical stimulation pulse signal according to the target disparity map to stimulate the patient's visual cortex cells or retinal cells, so as to realize phosphene corresponding to the surrounding environment image, realize the recognition of objects in the surrounding environment, and then effectively avoid them, improving the safety of the patient's movement.

[0031] Next, a binocular disparity estimation method according to an embodiment of the first aspect of the present invention will be described with reference to the accompanying drawings. The binocular disparity estimation method of the embodiment of the present invention can be applied to the artificial intelligence image processing unit 40 of the visual prosthesis 100.

[0032] Figure 2 is a flowchart of a binocular disparity estimation method according to an embodiment of the present invention. As Figure 2 shown, the binocular disparity estimation method of the embodiment of the present invention includes steps S1-S3.

[0033] S1, obtain a first image and a second image of the surrounding environment of the visual prosthesis wearer collected by the binocular cameras.

[0034] Specifically, the visual prosthesis may adopt binocular cameras, and the binocular cameras are worn on the blind patient to collect image information of the wearer's surrounding environment.

[0035] For example, the binocular cameras may include two horizontally placed cameras on the left and right, or the binocular cameras include two cameras arranged up and down.

[0036] The images of the surrounding environment of the visual prosthesis wearer are respectively collected by the binocular cameras, that is, a first image and a second image are generated, and the image information is sent to the artificial intelligence image processing unit of the visual prosthesis. In the embodiment, the artificial intelligence image processing unit may be an intelligent computing terminal such as a smart phone, a tablet, etc., or the artificial intelligence image processing unit may be an image processing chip integrated with the binocular cameras.

[0037] S2, perform depth feature extraction and matching fusion on the first image and the second image to obtain a feature map.

[0038] Among them, feature extraction is the process of extracting characteristic information from an image. For example, the extraction of an edge feature map that only shows all the object edges in the input image. Deep feature extraction can be understood as feature extraction based on a neural network model. The neural network model is trained based on a large amount of data in the actual life scenario, so as to effectively extract relevant features such as obstacle features in a complex scenario, and provide more accurate obstacle information for the user to avoid obstacles more easily.

[0039] Specifically, deep feature extraction is respectively performed on the first image and the second image, and the extracted features are matched, and the matching results are fused to obtain a feature map.

[0040] S3. Perform disparity estimation based on the feature map to obtain a target disparity map for generating an electrical stimulation pulse signal of a visual prosthesis.

[0041] Among them, the disparity map can refer to the disparity map in the image pair collected by a binocular camera, which is the position deviation of the pixels of the same scene imaged by two cameras. For example, for a binocular camera composed of two horizontally placed cameras, this position deviation is generally reflected in the horizontal direction. In the depth of field technology center, the disparity map can be converted into a depth map.

[0042] Specifically, perform disparity estimation based on the feature map to obtain a target view. The target view can be converted into a depth map. The depth map can not only reflect the objects in the scene where the user is located, but also reflect the distance information between the objects and it. And an electrical stimulation pulse signal of a visual prosthesis is generated according to the depth map. The implantation device of the visual prosthesis stimulates the retina or visual cortex of the wearer according to the electrical stimulation pulse signal, so as to induce the wearer's phosphene, help the wearer understand the objects in the surrounding environment and the distance information between them, and facilitate avoidance.

[0043] According to the binocular disparity estimation method of the embodiment of the present invention, the disparity map can be accurately calculated, the obstacle information can be effectively extracted in a complex environment, and a target disparity map for generating an electrical stimulation pulse signal of a visual prosthesis can be generated. That is, by combining the binocular disparity estimation method with the electrical stimulation of the visual prosthesis, the accuracy and effectiveness of the visual prosthesis in transmitting obstacle recognition and avoidance information assistance to the patient can be significantly improved, and the patient's mobility can be improved.

[0044] The process of obtaining the feature map and the target disparity map in the embodiment of the present invention will be further described below.

[0045] In an embodiment, deep feature extraction can be performed based on a neural network model. The first image and the second image collected by the binocular camera are input into the neural network. The backbone network, such as MobileNet, extracts the first image features in the first image and the second image features in the second image respectively by using shared weights. Among them, MobileNet is a lightweight convolutional neural network, and the parameters for its convolution are much fewer than those of standard convolution, so it can greatly reduce the parameters and the amount of computation. The basic unit of MobileNet is the depthwise separable convolution, which can achieve deep feature extraction.

[0046] Further, the first image features and the second image features are respectively subjected to feature matching according to a set of preset disparity values to obtain a plurality of matching results. For example, the preset disparity values can be fixed disparity values, that is, the two feature maps are respectively matched according to a set of fixed disparity values. Furthermore, the plurality of matching structures are fused into a set of feature maps.

[0047] In some embodiments, before deep feature extraction, the images can also be subjected to distortion correction and stereo correction, that is, the first image and the second image are subjected to distortion correction and stereo correction. Among them, distortion correction refers to correcting the perspective distortion inherent in the optical lens of the camera on the camera by using a formula. Stereo correction refers to correcting the display effects of the two lenses on the binocular camera for the same target, so that the corresponding points in the two images can fall on the same reference line. That is, through distortion correction and stereo correction, the matching points of the corrected first image and the second image are on the same pixel row, which can improve the accuracy of subsequent deep feature extraction and matching fusion.

[0048] In an embodiment, after obtaining a set of feature maps, disparity estimation is performed according to the feature maps. The fused set of feature maps is processed in two paths. One path inputs the set of feature maps into a disparity estimation network to obtain an initial disparity map, and the other path inputs the set of feature maps into a weight estimation network to obtain attention weights.

[0049] Among them, when estimating the initial disparity map, the disparity estimation network serializes the input set of feature maps to obtain serialized feature maps. The deep feature transformation network is used to model the dependencies of the input-output sequence, that is, the deep feature transformation network is used to model the dependencies of the obtained serialized feature maps and the serialized feature maps output by the disparity estimation network to determine the feature mapping relationship between the two. Thus, the obtained serialized feature maps can be aggregated through feature mapping according to the feature mapping relationship to estimate and obtain the initial disparity map. Among them, the disparity estimation network and the deep feature transformation network can be obtained based on the basic network models in related technologies and using the feature maps of the images collected by the binocular camera in the actual life scenario as training data.

[0050] Among them, when obtaining the attention weights, a set of feature maps obtained are input into the weight estimation network. The weight estimation network uses a multi-scale dilated convolutional pyramid module to construct a matching cost by aggregating environmental information of different sizes and different positions, and adjusts the matching cost by fusing and stacking multiple hourglass networks through 3D convolution. The attention weights are output through the SoftMax layer of the weight estimation network. Among them, dilated convolution interpolates between the kernel sizes of the convolution kernel, and the parameter dilation is used to determine the number of interpolations between the kernel sizes.

[0051] For example, in some embodiments, the multi-scale dilated convolutional pyramid network can be based on the U-Net as the basic model. In the encoding-decoding stage, dilated convolution is used to replace ordinary convolution to expand the receptive field, so that the output of each convolutional layer contains feature information in a larger range than ordinary convolution, which is beneficial to obtaining the global information of obstacle features in remote sensing images. The pyramid convolution module combines the U-Net skip connection structure to integrate multi-scale features to obtain high-resolution global overall information and low-resolution local detail information, thereby improving the accuracy and comprehensiveness of obstacle recognition in complex scenarios.

[0052] Furthermore, after obtaining the initial disparity map and the attention weights, the initial disparity map, the attention weights, and the feature maps are weighted and fused, and then input into the global information optimization network. The initial disparity map is optimized using the context temporal relationship for consecutive frames to obtain the final predicted disparity map.

[0053] Among them, when optimizing the initial disparity map using the context information, in the feature matching process, if no valid features can be matched in the local area of a frame of image, it will lead to obvious missing parts in the disparity map. In the embodiments of the present invention, the continuous frame optimization using the temporal relationship is an effective way to complement the missing disparity in the local area.

[0054] Specifically, in some embodiments, feature extraction based on MobileNet is performed on the first image and the second image of consecutive K frames respectively, and feature matching is performed between consecutive frames. Among them, the value of K can be between 2 and 20 frames. For example, feature extraction based on the backbone network and feature matching between consecutive frames are performed on 10 consecutive left images. And the camera pose transformation relationship of each frame of image in the relative coordinate system is solved based on the Perspective-n-Point algorithm, and the projection between frames is obtained. The missing areas are complemented using the projection between frames to obtain the final predicted disparity map.

[0055] In an embodiment, to solve the problem that it is difficult to determine the projection / completion relationship due to the lack of rich texture in the missing area, in some embodiments of the present invention, the missing area of the image is determined, and local superpixel segmentation is performed with the nearest matching point in the missing area as the center and at a preset resolution, for example, a resolution of R, so that the missing area is included in the mask. Among them, in the field of computer vision, image segmentation refers to the process of subdividing a digital image into multiple image sub-regions, that is, a set of pixels, which is also called a superpixel. Then, based on the projection relationship of the nearest matching point between consecutive frames, local disparity completion is performed using the effective disparity within the adjacent frame masks to obtain the final predicted disparity map.

[0056] 0 Further, downsample the final predicted disparity map according to the resolution set by the visual prosthesis to obtain the

[0057] target disparity map.

[0058] Based on the description of the above embodiments, the binocular disparity estimation method of the embodiments of the present invention may include several main steps such as binocular image correction, depth feature extraction and matching fusion, dilated convolution and weight estimation, context temporal relationship continuous frame optimization, and

[0059] image downsampling. Figure 3 is a flowchart of a binocular disparity estimation method according to an embodiment of the present invention. As Figure 3 shown, it includes:

[0060] S11, binocular camera calibration, that is, performing distortion correction and stereo correction.

[0061] S12, collecting and correcting binocular images.

[0062] S13, performing depth feature extraction and matching fusion, and then respectively performing steps S14 and S15.

[0063] S14, performing disparity estimation to obtain an initial disparity map, and entering step S16.

[0064] 0 S15, performing weight estimation.

[0065] S16, performing weighted fusion with the initial correction map.

[0066] S17, using the context temporal relationship to optimize the initial disparity map for continuous frames.

[0067] S18, downsampling the image according to the set resolution.

[0068] S19, obtaining the target disparity map.

[0069] 5 Generally speaking, the binocular disparity estimation method of the embodiments of the present invention is for the first image and the second obtained by the binocular camera

[0070] The images are subjected to distortion correction and stereo correction so that the matching points of the two corrected images are on the same pixel row. On this basis, deep features are extracted from the corrected images, multiple sets of disparities are estimated synchronously using an efficient feature transformation module, and finally, dilated convolution is used to estimate the weight of each disparity. After jointly optimizing the two networks, the final target disparity map is obtained.

[0071] In the embodiment of the present invention, the binocular disparity image estimation method of the above embodiment is combined with the visual prosthesis technology. Figure 4 Figure showing the working principle of a visual prosthesis based on artificial intelligence video processing according to an embodiment of the present invention

[0072] As shown in Figure 4 shown, the camera unit collects images of the surrounding environment of the wearing patient and sends the image information to the artificial intelligence image processing unit. The artificial intelligence image processing unit obtains the target disparity map according to the binocular disparity estimation method of the above embodiment, and sends the target disparity map to the implant device through a wireless signaler. The implant device generates an electrical stimulation pulse signal according to the target disparity map to stimulate the retinal cells or visual cortex cells of the wearer, so that the patient can form a perception of the images and distances of surrounding obstacles.

[0073] Combining the artificial intelligence image processing technology of the binocular disparity estimation method based on deep feature matching and continuous frame optimization with the visual prosthesis can accurately extract the information of surrounding obstacles and distances of the patient, significantly improve the accuracy and effectiveness of the visual prosthesis in transmitting obstacle recognition and avoidance information assistance to the patient, and provide efficient and reliable assistance information for the patient to recognize and avoid obstacles when acting alone.

[0074] Among them, through deep feature extraction and matching, the accuracy of disparity map estimation can be improved. Incorporating the optimization algorithm of continuous frames can effectively reduce the flicker problem between continuous frames, and further significantly reduce the interference and discomfort of the patient during the image perception process. Using dilated convolution makes the receptive field calculation flexible and efficient, which can effectively improve the real-time calculation efficiency and provide real-time obstacle disparity information for blind patients. In addition, adopting the binocular stereo matching method to obtain depth depth information has advantages such as a long acquisition distance and low price compared with the prior art, and is more suitable for visual prosthesis products.

[0075] The embodiment of the present invention also proposes a computer-readable storage medium, on which a computer program is stored. When the computer program is executed by a processor, the binocular disparity estimation method of the above embodiment can be implemented. The specific implementation process of the binocular disparity estimation method can refer to the description of the above embodiment.

[0076] Among them, in the embodiment, the above binocular disparity estimation method can be embodied by a computer-readable storage medium storing a computer program, and the computer program can be embodied by computer-readable code, and the computer-readable code includes instructions executable by at least one computing device. The computer-readable storage medium can be associated with any data storage device capable of storing data, and the data can be read by a computer system. Examples of computer-readable storage media may include read-only memory, random access memory, CD-ROM, HDD, DVD, magnetic tape, and optical data storage devices, etc. The computer-readable storage medium can also be distributed in computer systems connected through a network, so that the computer-readable code can be stored and executed distributively.

[0077] Although the embodiments of the present invention have been shown and described, those of ordinary skill in the art can understand that various changes, modifications, substitutions, and variations can be made to these embodiments without departing from the principles and spirit of the present invention, and the scope of the present invention is defined by the claims and their equivalents.

Claims

1. A binocular parallax estimation method, characterized in that Including: Obtaining a first image and a second image of the surrounding environment of a visual prosthesis wearer collected by a binocular camera; Performing depth feature extraction, matching, and fusion on the first image and the second image to obtain a feature map, including: using the first image and the second image as inputs, respectively extracting first image features in the first image and second image features in the second image through MobileNet in a way of sharing weights, performing feature matching on the first image features and the second image features according to a set of preset disparity values to obtain multiple matching results, and fusing the multiple matching results to obtain a set of feature maps; Performing disparity estimation based on the feature map to obtain a target disparity map for generating an electrical stimulation pulse signal of a visual prosthesis, including: serializing the set of feature maps to obtain a serialized feature map, using a deep feature transformation network to model dependencies of input-output sequences to determine a feature mapping relationship, aggregating the serialized feature map according to the feature mapping relationship through feature mapping to estimate an initial disparity map, and using a multi-scale dilated convolutional pyramid module with the set of feature maps as input information to construct a matching cost by aggregating environmental information of different sizes and different positions; adjusting the matching cost by stacking multiple hourglass networks through 3D convolution, outputting attention weights through the SoftMax layer of a weight estimation network, performing weighted fusion on the initial disparity map, the attention weights, and the feature map and inputting the result into a global information optimization network, and optimizing the initial disparity map frame by frame using context time series relationships to obtain a final predicted disparity map, and downsampling the final predicted disparity map according to the resolution set by the visual prosthesis to obtain the target disparity map.

2. The binocular parallax estimation method according to claim 1, wherein Before using the first image and the second image as inputs and respectively extracting first image features in the first image and second image features in the second image through MobileNet in a way of sharing weights, the method further includes: Performing distortion correction and stereo correction on the first image and the second image so that matching points of the corrected first image and the second image are on the same pixel row.

3. The binocular parallax estimation method according to claim 1, characterized in that Optimizing the initial disparity map frame by frame using context time series relationships to obtain a final predicted disparity map, including: Performing feature extraction based on MobileNet and feature matching between consecutive frames on the first image and the second image of consecutive K frames; Solving the camera pose transformation relationship of each frame of image in a relative coordinate system based on the Perspective-n-Point algorithm; Using projection between frames to complete missing regions to obtain the final predicted disparity map.

4. The binocular parallax estimation method according to claim 3, wherein Using projection between frames to complete missing regions to obtain the final predicted disparity map, including: Determining missing regions of an image; Performing superpixel segmentation of the missing regions with the nearest matching point in the missing regions as the center and at a preset resolution so that the missing regions are covered by a mask; Based on the projection relationship of the nearest matching points between consecutive frames, and using the valid disparity within the adjacent frame masks, perform local disparity completion of the missing region to obtain the final predicted disparity map.

5. A visual prosthesis, characterized in that, Including: An implant device and a wireless signaler, where the implant device is connected to the wireless signaler; A camera unit for collecting images of the wearer's surrounding environment; An artificial intelligence image processing unit, which is connected to the camera unit and the wireless signaler, and is used to obtain a target disparity map according to the binocular disparity estimation method described in any one of claims 1-4, and send the target disparity map to the implant device through the wireless signaler, and the implant device generates an electrical stimulation pulse signal according to the target disparity map.

6. A computer-readable storage medium having a computer program stored thereon, characterized in that, When the computer program is executed by a processor, it implements the binocular disparity estimation method described in any one of claims 1-4.

Citation Information

Patent Citations

  • Robot binocular stereo vision obstacle sensing method and system

    CN114842340A

  • Methods and Systems for Detecting Obstacles for a Visual Prosthesis

    US20160317811A1