Visual reconstruction method, device, equipment, computer readable medium and program product
By acquiring and processing the field of view images, generating residual vision area and energy field maps, combined with the visual reconstruction model, the difficulty and cost of customized visual aid devices are solved, and visual reconstruction is achieved without optical accessories, improving the accuracy and user experience of visual reconstruction.
Patent Information
- Application Number
- CN202510827733.0
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-06-20
- Publication Date
- 2025-07-22
- Estimated Expiration
- 2045-06-20
AI Technical Summary
The customization of existing visual field defective visual aid equipment is difficult and expensive, and requires frequent replacement of optical accessories, which cannot adapt to user vision changes.
By acquiring the original image and the field of view of the target user, binarization, denoising, and contour detection are performed to generate residual vision area and energy field maps. Combined with the pre-trained visual reconstruction model, visual reconstruction maps are generated and displayed, avoiding the replacement of customized equipment and optical accessories.
This realizes visual reconstruction without the need for customized equipment and optical accessories, and users can observe relatively complete images, improving the accuracy and user experience of visual reconstruction.
Smart Images

Figure CN120355629A_ABST
Abstract
Description
Technical Field
[0001] Embodiments of the present disclosure relate to the field of computer technology, and more particularly, to a visual reconstruction method, apparatus, device, computer-readable medium, and program product. Background Art
[0002] Vision is one of the most important sensory functions of humans, and most external information is obtained through vision. With the rapid change of human lifestyle and the continuous increase of work and study pressure, the problem of visual impairment is serious globally. This kind of irreversible blindness eye disease with visual field defect will bring great life obstacles to eye disease users. Currently, the existing visual field defect assisting devices are mainly designed based on optical principles.
[0003] However, when adopting the above method, there are often the following technical problems: the customization of visual field defect assisting devices is difficult, the cost is high, and as the user's vision changes, optical accessories need to be frequently replaced.
[0004] The above information disclosed in this background art section is only used to enhance the understanding of the background of the inventive concept, and thus, it may include information that does not form the prior art known to those of ordinary skill in the art in this country. Summary of the Invention
[0005] This content section of the present disclosure is used to briefly introduce concepts that will be described in detail in the following detailed implementation section. This content section of the present disclosure is not intended to identify the key features or essential features of the claimed technical solution, nor is it intended to limit the scope of the claimed technical solution.
[0006] Some embodiments of the present disclosure provide a visual reconstruction method, apparatus, device, computer-readable medium, and computer program product to solve one or more of the technical problems mentioned in the above background art section.
[0007] In a first aspect, some embodiments of the present disclosure provide a visual reconstruction method, which includes: obtaining an original image; obtaining a visual field image of a target user; performing binarization processing on the visual field image to obtain a binarized visual field image; performing denoising processing on the binarized visual field image to obtain a denoised visual field image; performing contour detection on the denoised visual field image to obtain a contour detection result; generating a residual visual acuity area corresponding to the target user based on the contour detection result; generating an energy field map according to the original image; generating a visual reconstruction map corresponding to the original image according to the generated energy field map, the original image, and the residual visual acuity area; and displaying the visual reconstruction map.
[0008] Optionally, generating a visual reconstruction map corresponding to the original image based on the generated energy field map, the original image, and the residual visual field region includes: generating a visual reconstruction map corresponding to the original image according to the generated energy field map, the original image, the residual visual field region, and a pre-trained visual reconstruction model.
[0009] Optionally, generating the residual visual field region corresponding to the target user based on the contour detection result includes: determining each contour region included in the contour detection result; and based on each of the contour regions, extracting, within the image size range of the visual field image, an image region that satisfies a preset maximum inscribed condition as the residual visual field region corresponding to the target user.
[0010] Optionally, generating an energy field map according to the original image includes: inputting the original image into a pre-trained energy field map generation model to obtain an energy field map, where the energy field map generation model includes an input layer, a position encoding layer, each depthwise separable convolutional layer, each linear layer, and an output layer, the position encoding layer is residually connected to the first depthwise separable convolutional layer among each of the depthwise separable convolutional layers, and adjacent two depthwise separable convolutional layers among each of the depthwise separable convolutional layers are residually connected.
[0011] Optionally, generating an energy field map according to the original image includes: performing color channel segmentation processing on the original image to obtain a set of segmentation images corresponding to each color channel, where each segmentation image in the set of segmentation images corresponds to one color channel; for each segmentation image in the set of segmentation images, performing the following steps: generating gradient information in the horizontal direction corresponding to each pixel in the segmentation image as horizontal gradient information; generating gradient information in the vertical direction corresponding to each pixel in the segmentation image as vertical gradient information; for each pixel in the segmentation image, generating comprehensive gradient information of the pixel in the corresponding color channel as the energy value of the pixel in the corresponding color channel according to the horizontal gradient information and the vertical gradient information corresponding to the pixel; determining each pixel with a determined energy value as an initial energy field map corresponding to the color channel; and performing normalization processing on the energy values of each pixel in the initial energy field map to obtain the initial energy field map after normalization processing as the energy field map corresponding to the color channel.
[0012] In a second aspect, some embodiments of the present disclosure provide a visual reconstruction device, which includes: a first acquisition unit configured to acquire an original image; a second acquisition unit configured to acquire a visual field image of a target user; a binarization processing unit configured to perform binarization processing on the visual field image to obtain a binarized visual field image; a denoising unit configured to perform denoising processing on the binarized visual field image to obtain a denoised visual field image; a contour detection unit configured to perform contour detection on the denoised visual field image to obtain a contour detection result; a first generation unit configured to generate a residual visual acuity area corresponding to the target user based on the contour detection result; a second generation unit configured to generate an energy field map according to the original image; a third generation unit configured to generate a visual reconstruction map corresponding to the original image according to the generated energy field map, the original image, and the residual visual acuity area; and a display unit configured to display the visual reconstruction map.
[0013] Optionally, the third generation unit is further configured to: generate a visual reconstruction map corresponding to the original image according to the generated energy field map, the original image, the residual visual acuity area, and a pre-trained visual reconstruction model.
[0014] Optionally, the first generation unit is further configured to: determine each contour area included in the contour detection result; and based on each contour area, extract an image area that satisfies a preset maximum inscribed condition within the image size range of the visual field image as the residual visual acuity area corresponding to the target user.
[0015] Optionally, the second generation unit is further configured to: input the original image into a pre-trained energy field map generation model to obtain an energy field map, where the energy field map generation model includes an input layer, a position encoding layer, each depthwise separable convolutional layer, each linear layer, and an output layer, the position encoding layer is connected in a residual manner to the first depthwise separable convolutional layer among each depthwise separable convolutional layer, and two adjacent depthwise separable convolutional layers among each depthwise separable convolutional layer are connected in a residual manner.
[0016] Optionally, the second generating unit is further configured to: perform color channel segmentation processing on the original image to obtain a set of segmented images corresponding to each color channel, where each segmented image in the set of segmented images corresponds to a color channel; for each segmented image in the set of segmented images, perform the following steps: generate gradient information in the horizontal direction corresponding to each pixel in the segmented image as horizontal gradient information; generate gradient information in the vertical direction corresponding to each pixel in the segmented image as vertical gradient information; for each pixel in the segmented image, generate comprehensive gradient information of the pixel in the corresponding color channel as the energy value of the pixel in the corresponding color channel according to the horizontal gradient information and the vertical gradient information corresponding to the pixel; determine each pixel with a determined energy value as the initial energy field map corresponding to the color channel; perform normalization processing on the energy values of each pixel in the initial energy field map to obtain the normalized initial energy field map as the energy field map corresponding to the color channel.
[0017] In a third aspect, some embodiments of the present disclosure provide an electronic device, including: one or more processors; a storage device having one or more programs stored thereon, and when the one or more programs are executed by the one or more processors, the one or more processors implement the method described in any implementation manner of the first aspect above.
[0018] In a fourth aspect, some embodiments of the present disclosure provide a computer-readable medium having a computer program stored thereon, where the program, when executed by a processor, implements the method described in any implementation manner of the first aspect above.
[0019] In a fifth aspect, some embodiments of the present disclosure provide a computer program product, including a computer program that, when executed by a processor, implements the method described in any implementation manner of the first aspect above.
[0020] The above-mentioned various embodiments of the present disclosure have the following beneficial effects: Through the visual reconstruction method of some embodiments of the present disclosure, there is no need for customized equipment or replacement of optical accessories. Specifically, the reasons for the high cost of equipment and the frequent need to replace optical accessories are as follows: The customization of visual field defect assisting devices is difficult and costly, and as the user's vision changes, optical accessories need to be frequently replaced. Based on this, in the visual reconstruction method of some embodiments of the present disclosure, first, an original image is obtained. Thus, the original image can be used as the image to be visually reconstructed. Then, the visual field image of the target user is obtained. Secondly, the above-mentioned visual field image is binarized to obtain a binarized visual field image. Next, the above-mentioned binarized visual field image is denoised to obtain a denoised visual field image. Then, contour detection is performed on the above-mentioned denoised visual field image to obtain a contour detection result. Then, based on the above-mentioned contour detection result, the residual visual acuity area corresponding to the above-mentioned target user is generated. Thus, the residual visual acuity area can represent the area where the target user's vision is normal. Next, an energy field map is generated according to the above-mentioned original image. Thus, the energy value of each pixel point in the energy field map can identify the importance or significance of the pixel, which can play a guiding role in the visual reconstruction process. Then, according to the generated energy field map, the above-mentioned original image, and the above-mentioned residual visual acuity area, a visual reconstruction map corresponding to the above-mentioned original image is generated. Finally, the above-mentioned visual reconstruction map is displayed. Thus, the user can observe a relatively complete image. Also, because the image for the user to view has undergone visual reconstruction processing and does not need to be processed based on optical principles, there is no need for customized equipment or replacement of optical accessories. BRIEF DESCRIPTION OF THE DRAWINGS
[0021] In combination with the accompanying drawings and with reference to the following specific embodiments, the above and other features, advantages, and aspects of the various embodiments of the present disclosure will become more apparent. Throughout the drawings, the same or similar reference numerals represent the same or similar elements. It should be understood that the drawings are schematic, and the elements and elements are not necessarily drawn to scale.
[0022] Figure 1 is a flowchart according to some embodiments of the visual reconstruction method of the present disclosure; Figure 2 is a schematic diagram of an application scenario for obtaining the residual visual acuity area of the visual reconstruction method according to some embodiments of the present disclosure; Figure 3 is a schematic diagram of the model structure of the energy field map generation model of the visual reconstruction method according to some embodiments of the present disclosure; Figure 4 is a schematic diagram of an application scenario of the visual reconstruction method according to some embodiments of the present disclosure; Figure 5 is a schematic diagram of the structure of some embodiments of the visual reconstruction device according to the present disclosure; Figure 6 It is a schematic structural diagram of an electronic device suitable for implementing some embodiments of the present disclosure. Detailed implementation manners
[0023] Embodiments of the present disclosure will be described in more detail below with reference to the accompanying drawings. Although some embodiments of the present disclosure are shown in the drawings, it should be understood that the present disclosure can be implemented in various forms and should not be construed as limited to the embodiments set forth herein. On the contrary, these embodiments are provided to more thoroughly and completely understand the present disclosure. It should be understood that the drawings and embodiments of the present disclosure are only for exemplary purposes and are not used to limit the protection scope of the present disclosure.
[0024] In addition, it should be noted that for the sake of convenience of description, only parts related to the relevant invention are shown in the drawings. Without conflict, the embodiments in the present disclosure and the features in the embodiments can be combined with each other.
[0025] It should be noted that the concepts such as "first" and "second" mentioned in the present disclosure are only used to distinguish different devices, modules or units, and are not used to limit the order of functions performed by these devices, modules or units or their interdependent relationships.
[0026] It should be noted that the modifications of "one" and "multiple" mentioned in the present disclosure are illustrative rather than restrictive. Those skilled in the art should understand that unless otherwise clearly specified in the context, it should be understood as "one or more".
[0027] The names of the messages or information exchanged between multiple devices in the embodiments of the present disclosure are only for illustrative purposes and are not used to limit the scope of these messages or information.
[0028] Regarding the collection, storage, use and other operations of the user's personal information (such as ophthalmic perimetry reports, visual field images) involved in the present disclosure, before performing the corresponding operations, relevant organizations or individuals shall fulfill obligations including conducting personal information security impact assessments, fulfilling the obligation of informing the personal information subject, and obtaining the prior authorization and consent of the personal information subject.
[0029] The present disclosure will be described in detail below with reference to the drawings and in combination with embodiments.
[0030] Figure 1 Flow 100 of some embodiments of the visual reconstruction method according to the present disclosure is shown. The visual reconstruction method includes the following steps: Step 101, obtain an original image.
[0031] In some embodiments, the execution subject of the visual reconstruction method (such as a head-mounted display device) can obtain an original image. Among them, the above-mentioned head-mounted display device can include, but is not limited to: AR glasses, MR glasses, VR glasses. The above-mentioned execution subject can also be a terminal device with a display screen. For example, the terminal device can include, but is not limited to: mobile phones, tablets. In practice, the above-mentioned execution subject can capture an image of the real scene through a configured camera as the original image. The above-mentioned execution subject can also obtain an image to be visually reconstructed from local storage or a server as the original image.
[0032] Step 102, obtain the field of view image of the target user.
[0033] In some embodiments, the above-mentioned execution subject can obtain the field of view image of the target user. Among them, the above-mentioned target user can be a user with visual field defect who is currently using the above-mentioned execution subject. The field of view image can be a visual field analysis image, which can represent the visual field of the user. In practice, the field of view image can be extracted from the ophthalmic visual field meter report of the above-mentioned target user.
[0034] Step 103, perform binarization processing on the field of view image to obtain a binarized field of view image.
[0035] In some embodiments, the above-mentioned execution subject can perform binarization processing on the above-mentioned field of view image to obtain a binarized field of view image. In practice, in response to determining that the above-mentioned field of view image is a color image, the above-mentioned field of view image can be converted into a grayscale image. Then, binarization processing can be performed on the grayscale image to obtain a binarized field of view image. When the above-mentioned field of view image is a grayscale image, binarization processing can be directly performed. Here, the binarization processing method can include, but is not limited to: global threshold method, adaptive threshold method. In the binarized field of view image, the pixel value of a pixel is 0 or 255.
[0036] Step 104, perform denoising processing on the binarized field of view image to obtain a denoised field of view image.
[0037] In some embodiments, the above-mentioned execution subject can perform denoising processing on the above-mentioned binarized field of view image to obtain a denoised field of view image. In practice, the above-mentioned execution subject can remove isolated black dots in the above-mentioned binarized field of view image to obtain a denoised field of view image. For example, morphological opening operation can be performed on the binarized field of view image to remove isolated black dots in the above-mentioned binarized field of view image to obtain a denoised field of view image.
[0038] Step 105, perform contour detection on the denoised field of view image to obtain a contour detection result.
[0039] In some embodiments, the execution subject may perform contour detection on the denoised visual field image to obtain a contour detection result. In practice, a contour detection method may be used to perform contour detection on the denoised visual field image to obtain at least one contour area as the contour detection result. The contour area may be represented by an image coordinate group. For example, the contour detection method may be the findContours method provided by the OpenCV library.
[0040] Step 106: Generate a residual vision area corresponding to the target user based on the contour detection result.
[0041] In some embodiments, the execution subject may generate a residual vision area corresponding to the target user based on the contour detection result. The residual vision area may be an area representing normal vision. For example, the residual vision area may be represented by an image coordinate group in a field of view image. In practice, the contour areas included in the contour detection result may be first determined. Then, based on the contour areas, within the image size range of the field of view image, an image area that satisfies a preset maximum inscribed condition may be extracted as the residual vision area corresponding to the target user. The preset maximum inscribed condition may be that the image area is the maximum inscribed area that avoids each contour area within the image size range of the field of view image, and the shape of the image area is a preset shape. For example, the preset shape may be a rectangle. For another example, the preset shape may be a circle. For example, random sampling, a scan line algorithm, or a grid-based search method may be used to extract an image area that satisfies a preset maximum inscribed condition within the image size range of the field of view image as the residual vision area corresponding to the target user.
[0042] As an example, the perimetry report, visual field images, and residual vision areas can be found in Figure 2 First, a visual field image may be extracted from the eye disease perimeter report of the target user, and then the residual vision area corresponding to the target user may be generated using the extracted visual field image. Figure 2In the ophthalmic perimetry report, due to the limitation of the figure size, the relevant parameters are not included. For example, the relevant parameters may include: single visual field analysis, central 10-2 threshold test, fixation monitoring: gaze / blind spot, fixation target: central, fixation loss: 2 / 20, false positive error rate: 3%, false negative error rate: 15%, test duration: 10 minutes and 23 seconds, fovea: off, eye: left eye, stimulus: III, white, background brightness: 31.5 ASB, test strategy: SITA-standard, pupil diameter: 6.7 mm, refractive power: +3.25 DS, astigmatism: none, date: April 13, 2019, time: 3:53 PM, age: 67 years old, mean defect MD: -8.58 dB P < 1%, pattern standard deviation PSD: 6.33 dB P < 1%. The mean defect refers to the difference between the light sensitivity of the examined eye and that of normal people in the same age group. The larger the value, the worse the light sensitivity. The pattern standard deviation represents the difference in the smoothness of the light sensitivity change in the visual field compared with that of normal people.
[0043] Optionally, the residual vision area may be pre-annotated in the visual field image of the ophthalmic perimetry report of the above target user. In practice, the pre-annotated residual vision area can be obtained from the ophthalmic perimetry report of the above target user.
[0044] Step 107, generate an energy field map according to the original image.
[0045] In some embodiments, the above execution subject may generate an energy field map according to the above original image. The energy field map may be an image representing the energy values corresponding to each pixel in the original image. The larger the energy value, the more important and significant the pixel can be represented.
[0046] In some alternative implementation manners of some embodiments, the above-mentioned execution subject may input the above-mentioned original image into a pre-trained energy field map generation model to obtain an energy field map. Among them, the above-mentioned energy field map generation model may be a neural network model that takes the original image as input data and the corresponding energy field map as output data. The above-mentioned energy field map generation model may include an input layer, a position encoding layer, each depthwise separable convolution layer, each linear layer, and an output layer. The above-mentioned position encoding layer is residually connected to the first depthwise separable convolution layer among the above-mentioned each depthwise separable convolution layer. Two adjacent depthwise separable convolution layers among the above-mentioned each depthwise separable convolution layer are residually connected. The above-mentioned input layer, the above-mentioned position encoding layer, the above-mentioned each depthwise separable convolution layer, the above-mentioned each linear layer, and the output layer are connected in sequence. The activation function between the above-mentioned each depthwise separable convolution layer and the above-mentioned each linear layer may be LeakyReLu. The depthwise separable convolution layer can reduce the computational amount and model size by decomposing the standard convolution operation into two parts: depthwise convolution and pointwise convolution.
[0047] As an example, the model structure of the energy field map generation model can refer to Figure 3 . Figure 3 In, the energy field map generation model may include 4 depthwise separable convolution layers and 4 linear layers. The dotted line indicates the residual connection. The bold arrow indicates that LeakyReLu is used as the activation function.
[0048] In some alternative implementation manners of some embodiments, the above-mentioned execution subject may generate an energy field map according to the above-mentioned original image through the following steps: In the first step, perform color channel segmentation processing on the above-mentioned original image to obtain a set of segmented images corresponding to each color channel. Among them, each segmented image in the above-mentioned set of segmented images corresponds to a color channel. The above-mentioned each color channel may include: red channel, green channel, and blue channel. Thus, by segmenting the images of different color channels, the segmented images of each color channel can be processed separately to better capture the detail changes under different colors.
[0049] In the second step, for each segmented image in the above-mentioned set of segmented images, perform the following steps: In the first sub-step, generate the gradient information corresponding to the horizontal direction of each pixel in the above-mentioned segmented image as the horizontal gradient information. In practice, the Sobel operator may be used to generate the gradient information corresponding to the horizontal direction of each pixel in the above-mentioned segmented image as the horizontal gradient information. The gradient information may include a gradient value. The gradient value may represent the change rate of pixel intensity and can measure the edge intensity.
[0050] The second sub-step is to generate the gradient information in the vertical direction corresponding to each pixel in the above-mentioned segmented image as the vertical gradient information. In practice, the Sobel operator can be used to generate the gradient information in the vertical direction corresponding to each pixel in the above-mentioned segmented image as the vertical gradient information.
[0051] The third sub-step is that for each pixel in the above-mentioned segmented image, according to the horizontal gradient information and the vertical gradient information corresponding to the above-mentioned pixel, generate the comprehensive gradient information of the above-mentioned pixel in the corresponding color channel as the energy value of the above-mentioned pixel in the corresponding color channel. In practice, the sum of the square of the above-mentioned horizontal gradient information and the square of the above-mentioned vertical gradient information can be determined as the first value. Then, the square root of the above-mentioned first value can be determined as the comprehensive gradient information of the above-mentioned pixel in the corresponding color channel. The comprehensive gradient information can represent the comprehensive energy value of the pixel.
[0052] The fourth sub-step is to determine each pixel with the determined energy value as the initial energy field map corresponding to the above-mentioned color channel.
[0053] The fifth sub-step is to perform normalization processing on the energy values of each pixel in the above-mentioned initial energy field map to obtain the normalized initial energy field map as the energy field map corresponding to the above-mentioned color channel. Since the energy values in different regions may vary greatly, it is necessary to normalize the energy values of all pixels to the same range (such as 0 to 255) for easy comparison. For example, a linear transformation can be applied to perform normalization processing on the energy values of each pixel in the above-mentioned initial energy field map so that the energy values of the pixels in the normalized initial energy field map are within the above-mentioned range. Thus, a corresponding energy field map can be generated for each color channel to better capture the detailed changes under different colors.
[0054] In some optional implementation manners of some embodiments, the above-mentioned execution subject can generate an energy field map according to the above-mentioned original image through the following steps: The first step is to perform denoising processing on the above-mentioned original image to obtain a denoised image. In practice, a Gaussian filter can be used to perform denoising processing on the above-mentioned original image to obtain a denoised image. Thus, the accuracy of the subsequent determined local pixel energy values can be improved.
[0055] The second step is that for each pixel in the above-mentioned denoised image, the following steps are performed: The first sub-step is to determine the hue information of the above-mentioned pixel. Among them, the hue information can be the hue value of the above-mentioned denoised image in the HSL space.
[0056] The second sub-step is to determine the color system type corresponding to the above hue information according to the above hue information. In practice, the color system type corresponding to the above hue information can be found through a pre-configured hue-color system type comparison table. The hue-color system type comparison table can be a comparison table of hue values and color system types. The color system types can include a red color system, a green color system, and a blue color system.
[0057] The third sub-step is to determine the color emotion information corresponding to the above color system type according to the above color system type. Among them, the color emotion information can include an emotion value, which can represent the intensity of the emotional reaction that can be brought. In practice, the preset emotion value corresponding to the above color system type can be determined as the color emotion information. For example, the preset emotion value corresponding to the red color system can be 0.8, which can represent a strong emotional reaction. The preset emotion value corresponding to the green color system can be 0.5, which can represent a calm or natural emotional reaction. The preset emotion value corresponding to the blue color system can be 0.3, which can represent a calm or sad emotional reaction.
[0058] The fourth sub-step is to determine the pixel coordinates of the above pixel.
[0059] The fifth sub-step is to determine the spatial weight information corresponding to the above pixel according to the above pixel coordinates. In practice, the distance between the above pixel coordinates and the center coordinates of the above denoised image can be determined first. Then, the above distance can be input into a pre-constructed linear function to obtain a weight value as the spatial weight information corresponding to the above pixel. The above linear function can be a monotonically decreasing function with the distance between the pixel and the center coordinates as the independent variable and the weight value as the dependent variable, and the value range is [0, 1].
[0060] The sixth sub-step is to generate the local complexity information corresponding to the above pixel. In practice, the local window corresponding to the above pixel can be determined first. The local window can be a window of a preset size that contains the above pixel. For example, the preset size can be 8 8 pixels. Then, the standard deviation of each pixel value in the above local window can be determined as the local complexity information corresponding to the above pixel.
[0061] The seventh sub-step is to generate an initial energy value corresponding to the above pixel according to the above color emotion information, the above spatial weight information, and the above local complexity information. In practice, the product of the above color emotion information, the above spatial weight information, and the above local complexity information can be determined as the initial energy value of the above pixel.
[0062] The third step is to perform normalization processing on the initial energy values of the obtained pixels to obtain the energy values of the pixels. In practice, the ratio of the initial energy value of each pixel to the maximum initial energy value among the initial energy values of the above pixels can be determined as the normalized energy value.
[0063] In the fourth step, each pixel with a determined energy value is determined as an energy field map.
[0064] The above first step to fourth step are an inventive point of the embodiment of the present disclosure, which solves the technical problem of "when generating the energy field map of an image, only the basic color attributes of the image are often concerned, and the differences in the spatial importance of pixels and the richness of local area details are ignored, and information in multiple dimensions cannot be organically combined, resulting in poor accuracy of the generated energy field map". The factors that lead to poor accuracy of the generated energy field map are often as follows: when generating the energy field map of an image, only the basic color attributes of the image are often concerned, and the differences in the spatial importance of pixels and the richness of local area details are ignored, and information in multiple dimensions cannot be organically combined. If the above factors are solved, the effect of improving the accuracy of the generated energy field map can be achieved. To achieve this effect, on the one hand, the present disclosure introduces a color emotion value, so that the generated energy field map can not only display the basic content of the image, but also convey the emotional levels contained in the image. On the other hand, by using a spatial weight matrix and local complexity calculation, the main part of the image can be effectively emphasized, and the areas rich in details can be highlighted. Thus, information in multiple dimensions can be organically combined, improving the accuracy of the generated energy field map. Furthermore, more information input can be provided for subsequent visual reconstruction tasks, which helps to improve the accuracy of visual reconstruction tasks.
[0065] Step 108, generate a visual reconstruction map corresponding to the original image according to the generated energy field map, the original image, and the residual vision area.
[0066] In some embodiments, the above execution subject may generate a visual reconstruction map corresponding to the original image according to the generated energy field map, the original image, and the residual vision area. In practice, a visual reconstruction map corresponding to the original image may be generated according to the generated energy field map, the original image, the residual vision area, and a pre-trained visual reconstruction model. The visual reconstruction model may be a neural network model that takes the energy field map, the original image, and the residual vision area as inputs and outputs a visual reconstruction map corresponding to the original image. For example, the neural network model may be a U-Net, a conditional generative adversarial network (Conditional GANs, cGANs), or a variational autoencoder (Variational Autoencoders, VAEs). The training samples of the visual reconstruction model may include sample original images, energy field maps corresponding to the sample original images, the residual vision areas of users, and sample visual reconstruction maps. The sample visual reconstruction maps may be repaired from the sample original images by experts according to the residual vision areas of users.
[0067] As an example, the visual reconstruction map generated based on the energy field map, the original image, and the residual vision area can be referred to Figure 4 .
[0068] In some optional implementation manners of some embodiments, the above-mentioned execution subject can generate a visual reconstruction map corresponding to the original image according to the generated energy field map, the original image, the residual vision area, and a pre-trained visual reconstruction model through the following steps: First step, according to the image size of the energy field map, adjust the size of the residual vision area to obtain an adjusted residual vision area. In practice, in response to determining that the image size of the energy field map is the same as the image size of the visual field image corresponding to the residual vision area, the residual vision area can be directly determined as the adjusted residual vision area. In response to determining that the image size of the energy field map is different from the image size of the visual field image corresponding to the residual vision area, the image size of the visual field image can be adjusted to the image size of the energy field map. Then, the residual vision area in the adjusted visual field image can be determined as the adjusted residual vision area.
[0069] Second step, superimpose the adjusted residual vision area on the energy field map to obtain a superimposed image. In practice, the adjusted residual vision area can be marked on the energy field map to obtain a superimposed image.
[0070] Third step, perform energy value adjustment processing on the superimposed image to obtain an energy value adjusted image. In practice, for each pixel outside the adjusted residual vision area in the superimposed image, the product of the energy value corresponding to the pixel and a preset gain factor can be determined as the adjusted energy value. The preset gain factor can be a value greater than 1. For example, the preset gain factor can be 1.5. Finally, the superimposed image with adjusted energy values can be determined as the energy value adjusted image. Thus, the importance of the visually normal area is retained, and at the same time, the importance of the damaged area is increased for subsequent processing.
[0071] Fourth step, input the original image and the energy value adjusted image into the input layer included in the visual reconstruction model to obtain an original image tensor and an energy value adjusted image tensor. Among them, the visual reconstruction model can include an input layer, a feature extraction network, an attention mechanism layer, a visual repair layer, and an image optimization layer. Among them, the input layer can convert the input images into tensor forms suitable for subsequent processing respectively.
[0072] Fifth step, input the original image tensor into the feature extraction network to obtain a multi-scale feature map. The feature extraction network can be a convolutional neural network, and a series of 3 Convolution kernels of size 3 are used, and a ReLU activation function follows each convolutional layer to increase the non-linear ability. Max-pooling operations can be used every few layers to reduce the spatial dimension while maintaining feature richness. Through the feature extraction network, multi-level features such as color, texture, and edges in the image can be extracted. Different levels of abstract information are captured through multiple convolutional layers and pooling layers.
[0073] In the sixth step, the above multi-scale feature maps and the above energy value-adjusted image tensors are input into the above attention mechanism layer to obtain weighted feature maps. The attention mechanism layer can adopt a custom attention mechanism, such as soft attention or hard attention, which can dynamically adjust the contribution degree of each position in the multi-scale feature maps according to the importance weights of the energy field maps, so as to guide the model to focus on key regions based on the energy field maps, especially the parts that are crucial for users with visual impairments.
[0074] In the seventh step, the above weighted feature maps are input into the above visual restoration layer to obtain an initial restored image. The visual restoration layer can utilize the generator part in the generative adversarial network (GANs) to perform image restoration by learning the mapping relationship between a large number of real images and damaged images, so as to specifically process the visually impaired areas, such as filling in missing information and enhancing blurred areas.
[0075] In the eighth step, the above initial restored image is input into the above image optimization layer to obtain a visual reconstruction map. The image optimization layer can improve the image resolution and reduce noise through a super-resolution network or a finely tuned convolutional network, making the final output as close as possible to the original high-quality image to further enhance the image quality and ensure clear and natural details.
[0076] The above first step to eighth step, as an inventive point of the embodiment of the present disclosure, solves the technical problem that "for patients with complex visual field defect patterns, traditional image enhancement methods often fail to provide effective assistance, resulting in a poor user experience provided by the generated visual reconstruction images". The factors that lead to a poor user experience provided by the generated visual reconstruction images are often as follows: for patients with complex visual field defect patterns, traditional image enhancement methods often fail to provide effective assistance. If the above factors are solved, the effect of improving the user experience provided by the generated visual reconstruction images can be achieved. To achieve this effect, the present disclosure fuses the residual vision area with the energy field map, and specifically enhances the importance of key information, ensuring that the importance of the visually normal area is retained, while increasing the importance of the damaged area for subsequent processing, effectively improving the transmission efficiency of visual information and reducing the risk of information loss or misunderstanding caused by vision impairment. At the same time, through the energy field map-guided attention mechanism, the parts crucial for understanding and interpreting the image are emphasized, and even if these parts are near the patient's visual field blind spot, they can be appropriately enhanced. Finally, image optimization is performed to make the final output as close as possible to the original high-quality image to further improve the image quality and ensure clear and natural details. Thereby, the user experience provided by the generated visual reconstruction images is improved.
[0077] Step 109, display the visual reconstruction map.
[0078] In some embodiments, the above execution subject can display the above visual reconstruction map. In practice, the above visual reconstruction map can be displayed on the display screen of the above execution subject.
[0079] The above-mentioned various embodiments of the present disclosure have the following beneficial effects: Through the visual reconstruction method of some embodiments of the present disclosure, there is no need for customized equipment or optical accessory replacement. Specifically, the reasons for the high equipment cost and frequent optical accessory replacement are as follows: The customization of visual field defect assisting devices is difficult and costly, and as the user's vision changes, optical accessories need to be frequently replaced. Based on this, in the visual reconstruction method of some embodiments of the present disclosure, first, an original image is obtained. Thus, the original image can be used as the image to be visually reconstructed. Then, the visual field image of the target user is obtained. Secondly, the above visual field image is binarized to obtain a binarized visual field image. Next, the above binarized visual field image is denoised to obtain a denoised visual field image. Then, contour detection is performed on the above denoised visual field image to obtain a contour detection result. Then, based on the above contour detection result, the residual vision area corresponding to the above target user is generated. Thus, the residual vision area can represent the area where the target user's vision is normal. Next, an energy field map is generated according to the above original image. Thus, the energy value of each pixel point in the energy field map can identify the importance or significance of the pixel, which can play a guiding role in the visual reconstruction process. Then, according to the generated energy field map, the above original image, and the above residual vision area, a visual reconstruction map corresponding to the above original image is generated. Finally, the above visual reconstruction map is displayed. Thus, the user can observe a relatively complete image. Also, because the image for the user to view has undergone visual reconstruction processing and does not need to be processed based on optical principles, there is no need for customized equipment or optical accessory replacement.
[0080] Further referring to Figure 5 , as an implementation of the methods shown in the above figures, the present disclosure provides some embodiments of a visual reconstruction device. These device embodiments correspond to Figure 1 the method embodiments shown, and the device can be specifically applied to various electronic devices.
[0081] As shown in Figure 5As shown, the visual reconstruction device 500 of some embodiments includes: a first acquisition unit 501, a second acquisition unit 502, a binarization processing unit 503, a denoising unit 504, a contour detection unit 505, a first generation unit 506, a second generation unit 507, a third generation unit 508, and a display unit 509. Among them, the first acquisition unit 501 is configured to acquire an original image; the second acquisition unit 502 is configured to acquire a field of view image of a target user; the binarization processing unit 503 is configured to perform binarization processing on the above-mentioned field of view image to obtain a binarized field of view image; the denoising unit 504 is configured to perform denoising processing on the above-mentioned binarized field of view image to obtain a denoised field of view image; the contour detection unit 505 is configured to perform contour detection on the above-mentioned denoised field of view image to obtain a contour detection result; the first generation unit 506 is configured to generate a residual visual acuity area corresponding to the above-mentioned target user based on the above-mentioned contour detection result; the second generation unit 507 is configured to generate an energy field map according to the above-mentioned original image; the third generation unit 508 is configured to generate a visual reconstruction map corresponding to the above-mentioned original image according to the generated energy field map, the above-mentioned original image, and the above-mentioned residual visual acuity area; the display unit 509 is configured to display the above-mentioned visual reconstruction map.
[0082] Optionally, the third generation unit 508 may further be configured to: generate a visual reconstruction map corresponding to the above-mentioned original image according to the generated energy field map, the above-mentioned original image, the above-mentioned residual visual acuity area, and a pre-trained visual reconstruction model.
[0083] Optionally, the first generation unit 506 may further be configured to: determine each contour area included in the above-mentioned contour detection result; based on each of the above-mentioned contour areas, extract an image area that satisfies a preset maximum inscribed condition within the image size range of the above-mentioned field of view image as the residual visual acuity area corresponding to the above-mentioned target user.
[0084] Optionally, the second generation unit 507 may further be configured to: input the above-mentioned original image into a pre-trained energy field map generation model to obtain an energy field map, where the above-mentioned energy field map generation model includes an input layer, a position encoding layer, each depthwise separable convolution layer, each linear layer, and an output layer, the above-mentioned position encoding layer is connected in a residual manner to the first depthwise separable convolution layer among each of the above-mentioned depthwise separable convolution layers, and two adjacent depthwise separable convolution layers among each of the above-mentioned depthwise separable convolution layers are connected in a residual manner.
[0085] Optionally, the second generating unit 507 may further be configured to: perform color channel segmentation processing on the original image to obtain a set of segmented images corresponding to each color channel, where each segmented image in the set of segmented images corresponds to a color channel; for each segmented image in the set of segmented images, perform the following steps: generate horizontal gradient information corresponding to each pixel in the segmented image as horizontal gradient information; generate vertical gradient information corresponding to each pixel in the segmented image as vertical gradient information; for each pixel in the segmented image, generate comprehensive gradient information of the pixel in the corresponding color channel as the energy value of the pixel in the corresponding color channel according to the horizontal gradient information and the vertical gradient information corresponding to the pixel; determine each pixel with a determined energy value as the initial energy field map corresponding to the color channel; perform normalization processing on the energy values of each pixel in the initial energy field map to obtain the normalized initial energy field map as the energy field map corresponding to the color channel.
[0086] It can be understood that the units described in the apparatus 500 correspond to the steps in the method described with reference to Figure 1 Therefore, the operations, features, and beneficial effects described above for the method also apply to the apparatus 500 and the units included therein, and will not be elaborated here.
[0087] Next, with reference to Figure 6 , which shows a schematic structural diagram of an electronic device 600 (such as a head-mounted display device or a terminal device) suitable for implementing some embodiments of the present disclosure. Figure 6 The electronic device shown is only an example and should not impose any limitations on the functions and usage scopes of the embodiments of the present disclosure.
[0088] As Figure 6 shown, the electronic device 600 may include a processing device 601 (such as a central processing unit, a graphics processing unit, etc.), which can perform various appropriate actions and processes according to a program stored in a read-only memory (ROM) 602 or a program loaded from a storage device 608 into a random access memory (RAM) 603. In the RAM 603, various programs and data required for the operation of the electronic device 600 are also stored. The processing device 601, the ROM 602, and the RAM 603 are connected to each other through a bus 604. An input / output (I / O) interface 605 is also connected to the bus 604.
[0089] Typically, the following devices can be connected to the I / O interface 605: input devices 606 including, for example, a touch screen, a touchpad, a keyboard, a mouse, a camera, a microphone, an accelerometer, a gyroscope, etc.; output devices 607 including, for example, a liquid crystal display (LCD), a speaker, a vibrator, etc.; and a communication device 609. The communication device 609 can allow the electronic device 600 to communicate with other devices wirelessly or wiredly to exchange data. Although Figure 6 the electronic device 600 with various devices is shown, it should be understood that it is not required to implement or have all the shown devices. Instead, more or fewer devices can be implemented or had. Figure 6 Each block shown in
[0090]
[0091] Specifically, according to some embodiments of the present disclosure, the processes described above with reference to the flowcharts can be implemented as computer software programs. For example, some embodiments of the present disclosure include a computer program product that includes a computer program carried on a computer-readable medium, and the computer program contains program codes for performing the methods shown in the flowcharts. In such some embodiments, the computer program can be downloaded and installed from a network through the communication device 609, or installed from the storage device 608, or installed from the ROM 602. When the computer program is executed by the processing device 601, the above functions defined in the methods of some embodiments of the present disclosure are performed.It should be noted that the computer-readable media described in some embodiments of the present disclosure may be a computer-readable signal medium, a computer-readable storage medium, or any combination of the two. A computer-readable storage medium may be, for example, but not limited to, an electrical, magnetic, optical, electromagnetic, infrared, or semiconductor system, apparatus, or device, or any combination of the above. More specific examples of the computer-readable storage medium may include, but are not limited to: an electrical connection having one or more wires, a portable computer disk, a hard disk, a random access memory (RAM), a read-only memory (ROM), an erasable programmable read-only memory (EPROM or flash memory), an optical fiber, a portable compact disk read-only memory (CD-ROM), an optical storage device, a magnetic storage device, or any suitable combination of the above. In some embodiments of the present disclosure, the computer-readable storage medium may be any tangible medium that contains or stores a program, which can be used by or in conjunction with an instruction execution system, apparatus, or device. In some embodiments of the present disclosure, a computer-readable signal medium may include a data signal propagated in a baseband or as part of a carrier wave, which carries computer-readable program code. Such a propagated data signal may take various forms, including but not limited to electromagnetic signals, optical signals, or any suitable combination of the above. A computer-readable signal medium may also be any computer-readable medium other than a computer-readable storage medium, which can send, propagate, or transmit a program for use by or in conjunction with an instruction execution system, apparatus, or device. The program code contained on the computer-readable medium may be transmitted using any appropriate medium, including but not limited to: wires, optical cables, RF (radio frequency), etc., or any suitable combination of the above.
[0092] In some embodiments, the client and the server may communicate using any currently known or future-developed network protocol such as HTTP (HyperText Transfer Protocol), and may be interconnected with digital data communication in any form or medium (e.g., a communication network). Examples of communication networks include local area networks ("LANs"), wide area networks ("WANs"), the Internet (e.g., the Internet), and end-to-end networks (e.g., ad hoc end-to-end networks), as well as any currently known or future-developed networks.
[0093] The above computer-readable medium may be included in the above electronic device; or may exist separately without being assembled into the electronic device. The above computer-readable medium carries one or more programs, and when the above one or more programs are executed by the electronic device, the electronic device is caused to: obtain an original image; obtain a field of view image of a target user; perform binarization processing on the above field of view image to obtain a binarized field of view image; perform denoising processing on the above binarized field of view image to obtain a denoised field of view image; perform contour detection on the above denoised field of view image to obtain a contour detection result; generate a residual vision area corresponding to the above target user based on the above contour detection result; generate an energy field map according to the above original image; generate a visual reconstruction map corresponding to the above original image according to the generated energy field map, the above original image, and the above residual vision area; and display the above visual reconstruction map.
[0094] Computer program code for performing the operations of some embodiments of the present disclosure may be written in one or more programming languages or combinations thereof. The above programming languages include object-oriented programming languages such as Java, Smalltalk, C++, and also include conventional procedural programming languages such as the "C" language or similar programming languages. The program code may be executed entirely on the user's computer, partially on the user's computer, executed as a stand-alone software package, partially on the user's computer and partially on a remote computer, or entirely on a remote computer or server. In the case of a remote computer, the remote computer may be connected to the user's computer through any type of network, including a local area network (LAN) or a wide area network (WAN), or may be connected to an external computer (for example, by connecting through the Internet using an Internet service provider).
[0095] The flowcharts and block diagrams in the accompanying drawings illustrate the possible architectures, functions, and operations of systems, methods, and computer program products according to various embodiments of the present disclosure. In this regard, each block in the flowchart or block diagram may represent a module, a program segment, or a part of code that contains one or more executable instructions for implementing a specified logical function. It should also be noted that in some alternative implementations, the functions marked in the blocks may occur in a different order than marked in the accompanying drawings. For example, two consecutive blocks shown may actually be executed substantially in parallel, and they may sometimes be executed in the reverse order, depending on the functions involved. It should also be noted that each block in the block diagram and / or flowchart, and the combinations of blocks in the block diagram and / or flowchart, may be implemented by a dedicated hardware-based system for performing the specified functions or operations, or may be implemented by a combination of dedicated hardware and computer instructions.
[0096] The units described in some embodiments of the present disclosure can be implemented in software or in hardware. The described units can also be provided in a processor. For example, it can be described as: a processor includes a first acquisition unit, a second acquisition unit, a binarization processing unit, a denoising unit, a contour detection unit, a first generation unit, a second generation unit, a third generation unit, and a display unit. Among them, the names of these units do not constitute a limitation to the unit itself in some cases. For example, the first acquisition unit can also be described as "the unit for acquiring the original image".
[0097] The functions described above herein can be performed, at least in part, by one or more hardware logic components. For example, without limitation, exemplary types of hardware logic components that can be used include: Field Programmable Gate Array (FPGA), Application Specific Integrated Circuit (ASIC), Application Specific Standard Product (ASSP), System on Chip (SOC), Complex Programmable Logic Device (CPLD), and so on.
[0098] Some embodiments of the present disclosure also provide a computer program product, including a computer program which, when executed by a processor, implements any of the above visual reconstruction methods.
[0099] The above description is only some preferred embodiments of the present disclosure and an explanation of the applied technical principles. Those skilled in the art should understand that the scope of the invention involved in the embodiments of the present disclosure is not limited to the technical solutions formed by the specific combination of the above technical features, and should also cover other technical solutions formed by any combination of the above technical features or their equivalent features without departing from the above inventive concept. For example, the technical solutions formed by mutually replacing the above features with the (but not limited to) technical features having similar functions disclosed in the embodiments of the present disclosure.
Claims
1. A visual reconstruction method, comprising: Obtaining an original image; Obtaining a field of view image of a target user; Performing binarization processing on the field of view image to obtain a binarized field of view image; Performing denoising processing on the binarized field of view image to obtain a denoised field of view image; Performing contour detection on the denoised field of view image to obtain a contour detection result; Generating a residual visual acuity area corresponding to the target user based on the contour detection result; Generating an energy field map according to the original image; Generating a visual reconstruction map corresponding to the original image according to the generated energy field map, the original image, and the residual visual acuity area; Displaying the visual reconstruction map.
2. The method according to claim 1, wherein, The generating a visual reconstruction map corresponding to the original image according to the generated energy field map, the original image, and the residual visual acuity area includes: Generating a visual reconstruction map corresponding to the original image according to the generated energy field map, the original image, the residual visual acuity area, and a pre-trained visual reconstruction model.
3. The method according to claim 1, wherein, The generating a residual visual acuity area corresponding to the target user based on the contour detection result includes: Determining each contour area included in the contour detection result; Based on each of the contour areas, within the image size range of the field of view image, extracting an image area that satisfies a preset maximum inscribed condition as the residual visual acuity area corresponding to the target user.
4. The method according to claim 1, wherein The generating an energy field map according to the original image includes: Inputting the original image into a pre-trained energy field map generation model to obtain an energy field map, where the energy field map generation model includes an input layer, a position encoding layer, each depthwise separable convolutional layer, each linear layer, and an output layer, the position encoding layer is residually connected to the first depthwise separable convolutional layer among each depthwise separable convolutional layer, and two adjacent depthwise separable convolutional layers among each depthwise separable convolutional layer are residually connected.
5. The method according to claim 1, wherein The generating an energy field map according to the original image includes: Performing color channel segmentation processing on the original image to obtain a set of segmented images corresponding to each color channel, where each segmented image in the set of segmented images corresponds to one color channel; For each segmented image in the set of segmented images, perform the following steps: Generating horizontal gradient information corresponding to each pixel in the segmented image as horizontal gradient information; Generating vertical gradient information corresponding to each pixel in the segmented image as vertical gradient information; For each pixel in the segmented image, generating comprehensive gradient information of the pixel in the corresponding color channel as the energy value of the pixel in the corresponding color channel according to the horizontal gradient information and the vertical gradient information corresponding to the pixel; Determining each pixel with a determined energy value as the initial energy field map corresponding to the color channel; Performing normalization processing on the energy values of each pixel in the initial energy field map to obtain the initial energy field map after normalization processing as the energy field map corresponding to the color channel.
6. A visual reconstruction device, comprising: A first acquisition unit configured to acquire an original image; A second acquisition unit configured to acquire a field of view image of a target user; A binarization processing unit, configured to perform binarization processing on the field of view image to obtain a binarized field of view image; A denoising unit, configured to perform denoising processing on the binarized field of view image to obtain a denoised field of view image; A contour detection unit, configured to perform contour detection on the denoised field of view image to obtain a contour detection result; A first generation unit, configured to generate a residual visual acuity area corresponding to the target user based on the contour detection result; A second generation unit, configured to generate an energy field map according to the original image; A third generation unit, configured to generate a visual reconstruction map corresponding to the original image according to the generated energy field map, the original image, and the residual visual acuity area; A display unit, configured to display the visual reconstruction map.
7. An electronic device, comprising: One or more processors; A storage device having one or more programs stored thereon, When the one or more programs are executed by the one or more processors, the one or more processors implement the method according to any one of claims 1-5.
8. A computer-readable medium having a computer program stored thereon, wherein, The computer program, when executed by a processor, implements the method according to any one of claims 1-5.
9. A computer program product, comprising a computer program which, when executed by a processor, implements the method according to any one of claims 1-5.
Citation Information
Patent Citations
Digital therapeutic corrective spectacles
CN111511318A
Vision assistance method, storage medium and vision assistance device
CN119165947A
Method for compensating visual field defects, electronic device, smart glasses, computer readable storage medium
TW202141234A
Smart point cloud reconstruction of objects in visual scenes in computing environments
US20190325638A1
Cited By
Image reconstruction method and device for abnormal vision, equipment, medium and product
CN122244223A
Visual field function evaluation method and system based on adaptive steady-state visual evoked potential
CN122478446A