Visual reconstruction method, apparatus, device, computer-readable medium, and program product

By acquiring and processing the field of view images to generate energy field maps and visual reconstruction maps, the problem of customized visual aid equipment is solved, and a low-cost and flexible visual reconstruction method is realized to adapt to user vision changes.

CN120355629BActive Publication Date: 2025-08-22HANGZHOU LINGBAN TECH CO LTD
View PDF 2 Cites 0 Cited by

Patent Information

Application Number
CN202510827733.0
Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Filing Date
2025-06-20
Publication Date
2025-08-22
Estimated Expiration
2045-06-20

AI Technical Summary

Technical Problem

The customization of existing visual field defective visual aid equipment is difficult and expensive, and requires frequent replacement of optical accessories, which cannot adapt to user vision changes.

Method used

By acquiring the original image and the field of view of the target user, performing binarization and denoising processing, performing contour detection, generating residual vision areas, and generating visual reconstruction images based on the energy field map and field of view reconstruction model to display the reconstruction image.

Benefits of technology

There is no need to customize equipment and replace optical accessories, providing a relatively complete image observation experience, reducing equipment costs and adapting to user vision changes.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120355629B_ABST
    Figure CN120355629B_ABST
Patent Text Reader

Abstract

The embodiments of the present disclosure disclose a visual reconstruction method, apparatus, device, computer-readable medium, and program product. A specific implementation of the method includes: acquiring an original image; acquiring a field of view image of a target user; binarizing the field of view image to obtain a binarized field of view image; denoising the binarized field of view image to obtain a denoised field of view image; performing contour detection on the denoised field of view image to obtain a contour detection result; based on the contour detection result, generating a residual vision area corresponding to the target user; generating an energy field map according to the original image; generating a visual reconstruction map corresponding to the original image according to the generated energy field map, the original image, and the residual vision area; and displaying the visual reconstruction map. This implementation does not require customized visual field defect aids, nor does it require replacement of optical accessories.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] Embodiments of the present disclosure relate to the field of computer technology, and in particular to a visual reconstruction method, apparatus, device, computer-readable medium, and program product. Background Art

[0002] Vision is one of the most important human senses, providing the majority of external information. With the dramatic changes in lifestyles and increasing work and study pressures, visual impairment is a serious problem worldwide. Irreversible blinding eye diseases such as visual field loss can significantly hinder the lives of those with these conditions. Current visual aids for visual field loss are primarily designed based on optical principles.

[0003] However, when the above approach is adopted, the following technical problems often arise: customization of visual field defect aids is difficult and costly, and optical accessories need to be frequently replaced as the user's vision changes.

[0004] The above information disclosed in this Background section is only for enhancement of understanding of the background of the inventive concept and therefore it may contain information that does not form the prior art that is already known in this country to a person of ordinary skill in the art. Summary of the Invention

[0005] The content of this disclosure is used to briefly introduce concepts that will be described in detail in the detailed description section below. The content of this disclosure is not intended to identify key features or essential features of the claimed technical solution, nor is it intended to limit the scope of the claimed technical solution.

[0006] Some embodiments of the present disclosure provide visual reconstruction methods, devices, apparatuses, computer-readable media, and computer program products to solve one or more of the technical problems mentioned in the above background technology section.

[0007] In a first aspect, some embodiments of the present disclosure provide a visual reconstruction method, which includes: obtaining an original image; obtaining a field of view image of a target user; binarizing the field of view image to obtain a binarized field of view image; denoising the binarized field of view image to obtain a denoised field of view image; performing contour detection on the denoised field of view image to obtain a contour detection result; based on the contour detection result, generating a residual vision area corresponding to the target user; generating an energy field map according to the original image; generating a visual reconstruction map corresponding to the original image according to the generated energy field map, the original image and the residual vision area; and displaying the visual reconstruction map.

[0008] Optionally, the above-mentioned generating a visual reconstruction map corresponding to the above-mentioned original image based on the generated energy field map, the above-mentioned original image and the above-mentioned residual vision area includes: generating a visual reconstruction map corresponding to the above-mentioned original image based on the generated energy field map, the above-mentioned original image, the above-mentioned residual vision area and a pre-trained visual reconstruction model.

[0009] Optionally, the above-mentioned generating the residual vision area corresponding to the above-mentioned target user based on the above-mentioned contour detection result includes: determining each contour area included in the above-mentioned contour detection result; based on the above-mentioned each contour area, within the image size range of the above-mentioned field of view image, extracting the image area that meets the preset maximum inscribed condition as the residual vision area corresponding to the above-mentioned target user.

[0010] Optionally, the above-mentioned generation of an energy field map based on the above-mentioned original image includes: inputting the above-mentioned original image into a pre-trained energy field map generation model to obtain an energy field map, wherein the above-mentioned energy field map generation model includes an input layer, a position encoding layer, each depth-separable convolution layer, each linear layer and an output layer, the above-mentioned position encoding layer is residually connected to the first depth-separable convolution layer in each of the above-mentioned depth-separable convolution layers, and the two adjacent depth-separable convolution layers in each of the above-mentioned depth-separable convolution layers are residually connected.

[0011] Optionally, the above-mentioned generation of an energy field map based on the above-mentioned original image includes: performing color channel segmentation processing on the above-mentioned original image to obtain a set of segmented images corresponding to each color channel, wherein each segmented image in the above-mentioned segmented image set corresponds to a color channel; for each segmented image in the above-mentioned segmented image set, performing the following steps: generating gradient information in the horizontal direction corresponding to each pixel in the above-mentioned segmented image as horizontal gradient information; generating gradient information in the vertical direction corresponding to each pixel in the above-mentioned segmented image as vertical gradient information; for each pixel in the above-mentioned segmented image, generating comprehensive gradient information of the above-mentioned pixel in the corresponding color channel as the energy value of the above-mentioned pixel in the corresponding color channel based on the horizontal gradient information and vertical gradient information corresponding to the above-mentioned pixel; determining each pixel with determined energy value as the initial energy field map corresponding to the above-mentioned color channel; normalizing the energy value of each pixel in the above-mentioned initial energy field map to obtain the normalized initial energy field map as the energy field map corresponding to the above-mentioned color channel.

[0012] In a second aspect, some embodiments of the present disclosure provide a visual reconstruction device, which includes: a first acquisition unit, configured to acquire an original image; a second acquisition unit, configured to acquire a field of view image of a target user; a binarization processing unit, configured to perform binarization processing on the field of view image to obtain a binarized field of view image; a denoising unit, configured to perform denoising processing on the binarized field of view image to obtain a denoised field of view image; a contour detection unit, configured to perform contour detection on the denoised field of view image to obtain a contour detection result; a first generation unit, configured to generate a residual vision area corresponding to the target user based on the contour detection result; a second generation unit, configured to generate an energy field map according to the original image; a third generation unit, configured to generate a visual reconstruction map corresponding to the original image according to the generated energy field map, the original image and the residual vision area; and a display unit, configured to display the visual reconstruction map.

[0013] Optionally, the third generating unit is further configured to generate a visual reconstruction map corresponding to the original image based on the generated energy field map, the original image, the residual vision area and a pre-trained visual reconstruction model.

[0014] Optionally, the first generation unit is further configured to: determine the various contour areas included in the above-mentioned contour detection results; based on the above-mentioned various contour areas, within the image size range of the above-mentioned field of view image, extract the image area that meets the preset maximum inscribed condition as the residual vision area corresponding to the above-mentioned target user.

[0015] Optionally, the second generation unit is further configured to: input the above-mentioned original image into a pre-trained energy field map generation model to obtain an energy field map, wherein the above-mentioned energy field map generation model includes an input layer, a position encoding layer, each depth-separable convolution layer, each linear layer and an output layer, the above-mentioned position encoding layer is residually connected to the first depth-separable convolution layer in each of the above-mentioned depth-separable convolution layers, and the two adjacent depth-separable convolution layers in each of the above-mentioned depth-separable convolution layers are residually connected.

[0016] Optionally, the second generation unit is further configured to: perform color channel segmentation processing on the above-mentioned original image to obtain a set of segmented images corresponding to each color channel, wherein each segmented image in the above-mentioned segmented image set corresponds to a color channel; for each segmented image in the above-mentioned segmented image set, perform the following steps: generate the gradient information corresponding to the horizontal direction of each pixel in the above-mentioned segmented image as horizontal gradient information; generate the gradient information corresponding to the vertical direction of each pixel in the above-mentioned segmented image as vertical gradient information; for each pixel in the above-mentioned segmented image, generate the comprehensive gradient information of the above-mentioned pixel in the corresponding color channel as the energy value of the above-mentioned pixel in the corresponding color channel based on the horizontal gradient information and vertical gradient information corresponding to the above-mentioned pixel; determine each pixel with determined energy value as the initial energy field map corresponding to the above-mentioned color channel; normalize the energy value of each pixel in the above-mentioned initial energy field map to obtain the normalized initial energy field map as the energy field map corresponding to the above-mentioned color channel.

[0017] In a third aspect, some embodiments of the present disclosure provide an electronic device comprising: one or more processors; a storage device on which one or more programs are stored, and when the one or more programs are executed by one or more processors, the one or more processors implement the method described in any implementation of the first aspect above.

[0018] In a fourth aspect, some embodiments of the present disclosure provide a computer-readable medium having a computer program stored thereon, wherein when the program is executed by a processor, the method described in any implementation of the first aspect is implemented.

[0019] In a fifth aspect, some embodiments of the present disclosure provide a computer program product, including a computer program, which implements the method described in any implementation of the first aspect when executed by a processor.

[0020] The above-described embodiments of the present disclosure have the following beneficial effects: The visual reconstruction methods of some embodiments of the present disclosure eliminate the need for customized equipment and replacement of optical components. Specifically, the high equipment costs and frequent replacement of optical components are due to the difficulty and high cost of customizing visual acuity devices for visual impairment, and the frequent replacement of optical components as the user's vision changes. Based on this, the visual reconstruction methods of some embodiments of the present disclosure first acquire an original image. Thus, the original image can serve as the image to be visually reconstructed. Then, an image of the visual field of a target user is acquired. Next, the visual field image is binarized to obtain a binarized visual field image. Next, the binarized visual field image is denoised to obtain a denoised visual field image. Next, contour detection is performed on the denoised visual field image to obtain a contour detection result. Next, based on the contour detection result, a residual vision area corresponding to the target user is generated. Thus, the residual vision area can represent the target user's normal vision area. Next, an energy field map is generated based on the original image. The energy value of each pixel in the energy field map can identify the importance or significance of that pixel, thereby serving as a guide in the visual reconstruction process. Then, based on the generated energy field map, the original image, and the residual vision area, a visual reconstruction image corresponding to the original image is generated. Finally, the visual reconstruction image is displayed. This allows the user to observe a relatively complete image. Because the image presented to the user has undergone visual reconstruction, no optical processing is required, eliminating the need for customized equipment or replacement of optical components. BRIEF DESCRIPTION OF THE DRAWINGS

[0021] The above and other features, advantages, and aspects of the various embodiments of the present disclosure will become more apparent with reference to the following detailed description in conjunction with the accompanying drawings. Throughout the drawings, the same or similar reference numerals represent the same or similar elements. It should be understood that the drawings are schematic and that components and elements are not necessarily drawn to scale.

[0022] Figure 1 is a flow chart of some embodiments of the visual reconstruction method according to the present disclosure;

[0023] Figure 2 is a schematic diagram of an application scenario of obtaining a residual vision area according to a visual reconstruction method of some embodiments of the present disclosure;

[0024] Figure 3 is a schematic diagram of a model structure of an energy field map generation model of a visual reconstruction method according to some embodiments of the present disclosure;

[0025] Figure 4 is a schematic diagram of an application scenario of the visual reconstruction method according to some embodiments of the present disclosure;

[0026] Figure 5 is a schematic structural diagram of some embodiments of the visual reconstruction device according to the present disclosure;

[0027] Figure 6 It is a structural diagram of an electronic device suitable for implementing some embodiments of the present disclosure. DETAILED DESCRIPTION

[0028] Embodiments of the present disclosure will be described in more detail below with reference to the accompanying drawings. Although certain embodiments of the present disclosure are shown in the accompanying drawings, it should be understood that the present disclosure can be implemented in various forms and should not be construed as being limited to the embodiments described herein. On the contrary, these embodiments are provided to provide a more thorough and complete understanding of the present disclosure. It should be understood that the drawings and embodiments of the present disclosure are for illustrative purposes only and are not intended to limit the scope of protection of the present disclosure.

[0029] It should also be noted that, for ease of description, only the parts related to the invention are shown in the drawings. In the absence of conflict, the embodiments and features in the embodiments of the present disclosure may be combined with each other.

[0030] It should be noted that the concepts of "first" and "second" mentioned in this disclosure are only used to distinguish different devices, modules or units, and are not used to limit the order or interdependence of the functions performed by these devices, modules or units.

[0031] It should be noted that the modifications of "one" and "multiple" mentioned in the present disclosure are illustrative rather than restrictive, and those skilled in the art should understand that unless otherwise clearly indicated in the context, they should be understood as "one or more".

[0032] The names of the messages or information exchanged between multiple devices in the embodiments of the present disclosure are only used for illustrative purposes and are not used to limit the scope of these messages or information.

[0033] With regard to the collection, storage, and use of user personal information (such as perimetry reports and visual field images) involved in this disclosure, before performing the corresponding operations, the relevant organizations or individuals shall fulfill their obligations, including conducting personal information security impact assessments, fulfilling the obligation to inform the personal information subjects, and obtaining the authorization and consent of the personal information subjects in advance.

[0034] The present disclosure will be described in detail below with reference to the accompanying drawings and in conjunction with embodiments.

[0035] Figure 1 A process 100 of some embodiments of the visual reconstruction method according to the present disclosure is shown. The visual reconstruction method comprises the following steps:

[0036] Step 101: Acquire an original image.

[0037] In some embodiments, the execution entity of the visual reconstruction method (e.g., a head-mounted display device) can obtain the original image. The head-mounted display device may include, but is not limited to, AR glasses, MR glasses, and VR glasses. The execution entity may also be a terminal device with a display screen. For example, the terminal device may include, but is not limited to, a mobile phone or tablet computer. In practice, the execution entity may use a configured camera to capture an image of the real scene as the original image. The execution entity may also obtain the image to be visually reconstructed from local storage or a server as the original image.

[0038] Step 102: Acquire a field of view image of the target user.

[0039] In some embodiments, the execution entity may obtain a target user's visual field image. The target user may be a user currently using the execution entity who has a visual field loss. The visual field image may be a visual field analysis image capable of representing the user's visual field. In practice, the visual field image may be extracted from the target user's eye disease perimetry report.

[0040] Step 103 : binarize the visual field image to obtain a binarized visual field image.

[0041] In some embodiments, the execution subject may perform binarization processing on the field of view image to obtain a binary field of view image. In practice, in response to determining that the field of view image is a color image, the field of view image may be converted into a grayscale image. Then, the grayscale image may be binarized to obtain a binary field of view image. When the field of view image is a grayscale image, binarization processing may be performed directly. The binarization processing method here may include but is not limited to: a global threshold method and an adaptive threshold method. In the binary field of view image, the pixel value of a pixel is 0 or 255.

[0042] Step 104 , performing denoising processing on the binarized visual field image to obtain a denoised visual field image.

[0043] In some embodiments, the execution entity may perform denoising on the binarized field of view image to obtain a denoised field of view image. In practice, the execution entity may remove isolated black dots from the binarized field of view image to obtain the denoised field of view image. For example, an opening operation may be performed on the binarized field of view image to remove isolated black dots from the binarized field of view image to obtain the denoised field of view image.

[0044] Step 105: Perform contour detection on the denoised visual field image to obtain a contour detection result.

[0045] In some embodiments, the execution entity may perform contour detection on the denoised visual field image to obtain a contour detection result. In practice, a contour detection method may be used to perform contour detection on the denoised visual field image to obtain at least one contour region as the contour detection result. The contour region may be represented by a set of image coordinates. For example, the contour detection method may be the findContours method provided by the OpenCV library.

[0046] Step 106: Generate a residual vision area corresponding to the target user based on the contour detection result.

[0047] In some embodiments, the execution entity may generate a residual vision area corresponding to the target user based on the contour detection result. The residual vision area may be an area representing normal vision. For example, the residual vision area may be represented by an image coordinate group in the field of view image. In practice, the various contour areas included in the contour detection result may be determined first. Then, based on the various contour areas, within the image size range of the field of view image, an image area that meets a preset maximum inscribed condition may be extracted as the residual vision area corresponding to the target user. The preset maximum inscribed condition may be that the image area is the maximum inscribed area that avoids various contour areas within the image size range of the field of view image, and the shape of the image area is a preset shape. For example, the preset shape may be a rectangle. For another example, the preset shape may be a circle. For example, random sampling, a scan line algorithm, or a grid-based search method may be used to extract an image area that meets the preset maximum inscribed condition within the image size range of the field of view image as the residual vision area corresponding to the target user.

[0048] As an example, the perimetry report, visual field image, and residual vision area can be referred to Figure 2 First, a visual field image may be extracted from the target user's eye disease perimetry report, and then the extracted visual field image may be used to generate the residual vision area corresponding to the target user. Figure 2Due to image size limitations, relevant parameters are not included in the perimetry report. For example, relevant parameters may include: single visual field analysis, central 10-2 threshold test, gaze monitoring: gaze / blind spot, fixation target: central, fixation loss: 2 / 20, false positive error rate: 3%, false negative error rate: 15%, test duration: 10 minutes and 23 seconds, fovea: closed, eye: left eye, stimulus: III, white, background luminance: 31.5ASB, testing strategy: SITA-standard, pupil diameter: 6.7 mm, refractive power: +3.25DS, astigmatism: none, date: April 13, 2019, time: 3:53 PM, age: 67 years, mean defect (MD): -8.58 dB P < 1%, pattern standard deviation (PSD): 6.33 dB P < 1%. The mean defect refers to the difference between the light sensitivity of the subject's eye and the normal light sensitivity of a person of the same age group. A larger value indicates worse light sensitivity. The pattern standard deviation indicates how smooth the changes in light sensitivity in the visual field differ from those of normal people.

[0049] Optionally, the residual vision area may be pre-marked in the visual field image of the target user's eye disease perimeter report. In practice, the pre-marked residual vision area may be obtained from the target user's eye disease perimeter report.

[0050] Step 107: Generate an energy field map based on the original image.

[0051] In some embodiments, the execution entity may generate an energy field map based on the original image. The energy field map may be an image representing the energy values ​​corresponding to each pixel in the original image. A larger energy value may indicate a more important or significant pixel.

[0052] In some optional implementations of some embodiments, the execution entity may input the original image into a pre-trained energy field map generation model to obtain an energy field map. The energy field map generation model may be a neural network model that uses the original image as input data and the corresponding energy field map as output data. The energy field map generation model may include an input layer, a position encoding layer, depthwise separable convolutional layers, linear layers, and an output layer. The position encoding layer is residually connected to the first depthwise separable convolutional layer in each depthwise separable convolutional layer. Two adjacent depthwise separable convolutional layers in each depthwise separable convolutional layer are residually connected. The input layer, the position encoding layer, the depthwise separable convolutional layers, the linear layers, and the output layer are sequentially connected. The activation function between each depthwise separable convolutional layer and each linear layer may be LeakyReLu. The depthwise separable convolutional layer may reduce computational complexity and model size by decomposing the standard convolution operation into two parts: depthwise convolution and pointwise convolution.

[0053] As an example, the model structure of the energy field map generation model can be referred to Figure 3 . Figure 3 In

[15] , the energy field map generation model can include 4 depthwise separable convolutional layers and 4 linear layers. The dotted lines represent residual connections. The bold arrows indicate the use of LeakyReLu as the activation function.

[0054] In some optional implementations of some embodiments, the execution entity may generate an energy field map based on the original image through the following steps:

[0055] The first step is to segment the original image by color channel, obtaining a set of segmented images corresponding to each color channel. Each segmented image in the segmented image set corresponds to a color channel. These color channels may include red, green, and blue channels. Thus, by segmenting the image by color channel, the segmented images for each color channel can be processed separately to better capture the detailed variations in different colors.

[0056] In the second step, for each segmented image in the above segmented image set, perform the following steps:

[0057] The first sub-step is to generate horizontal gradient information corresponding to each pixel in the segmented image as horizontal gradient information. In practice, a Sobel operator can be used to generate horizontal gradient information corresponding to each pixel in the segmented image as horizontal gradient information. The gradient information can include a gradient value. The gradient value can represent the rate of change of pixel intensity and can measure edge strength.

[0058] The second sub-step is to generate vertical gradient information corresponding to each pixel in the segmented image as vertical gradient information. In practice, the Sobel operator can be used to generate vertical gradient information corresponding to each pixel in the segmented image as vertical gradient information.

[0059] In a third sub-step, for each pixel in the segmented image, based on the horizontal gradient information and vertical gradient information corresponding to the pixel, comprehensive gradient information of the pixel in the corresponding color channel is generated as the energy value of the pixel in the corresponding color channel. In practice, the sum of the square of the horizontal gradient information and the square of the vertical gradient information can be determined as a first value. Then, the first power of half of the first value can be determined as the comprehensive gradient information of the pixel in the corresponding color channel. The comprehensive gradient information can represent the comprehensive energy value of the pixel.

[0060] The fourth sub-step is to determine each pixel with a determined energy value as an initial energy field map corresponding to the above color channel.

[0061] The fifth sub-step is to normalize the energy values ​​of each pixel in the above-mentioned initial energy field map to obtain the normalized initial energy field map as the energy field map corresponding to the above-mentioned color channel. Since the energy values ​​in different areas may vary greatly, the energy values ​​of all pixels need to be normalized to the same range (such as 0 to 255) for easy comparison. For example, a linear transformation can be applied to normalize the energy values ​​of each pixel in the above-mentioned initial energy field map so that the energy values ​​of the pixels in the normalized initial energy field map are within the above-mentioned range. In this way, a corresponding energy field map can be generated for each color channel to better capture the detailed changes under different colors.

[0062] In some optional implementations of some embodiments, the execution entity may generate an energy field map based on the original image through the following steps:

[0063] The first step is to perform denoising on the original image to obtain a denoised image. In practice, a Gaussian filter can be used to perform denoising on the original image to obtain a denoised image. This can improve the accuracy of the local pixel energy values ​​subsequently determined.

[0064] In the second step, for each pixel in the above denoised image, perform the following steps:

[0065] The first sub-step is to determine the hue information of the pixel, wherein the hue information may be the hue value of the denoised image in the HSL space.

[0066] The second sub-step is to determine the color system type corresponding to the hue information based on the hue information. In practice, the color system type corresponding to the hue information can be found using a pre-configured hue and color system type comparison table. The hue and color system type comparison table can be a table that compares hue values ​​with color system types. Color system types can include red, green, and blue color systems.

[0067] The third sub-step is to determine the color emotion information corresponding to the above color type based on the above color type. The color emotion information may include an emotion value, which may indicate the intensity of the emotional response that can be brought about. In practice, the preset emotion value corresponding to the above color type may be determined as the color emotion information. For example, the preset emotion value corresponding to the red color may be 0.8, which may indicate a strong emotional response. The preset emotion value corresponding to the green color may be 0.5, which may indicate a calm or natural emotional response. The preset emotion value corresponding to the blue color may be 0.3, which may indicate a calm or sad emotional response.

[0068] The fourth sub-step is to determine the pixel coordinates of the above-mentioned pixels.

[0069] The fifth sub-step is to determine the spatial weight information corresponding to the above pixel based on the above pixel coordinates. In practice, the distance between the above pixel coordinates and the center coordinates of the above denoised image can be determined first. Then, the above distance can be input into a pre-constructed linear function to obtain the weight value as the spatial weight information corresponding to the above pixel. The above linear function can be a monotonically decreasing function with the distance between the pixel and the center coordinate as the independent variable and the weight value as the dependent variable, and the value range is [0, 1].

[0070] The sixth sub-step is to generate the local complexity information corresponding to the above pixel. In practice, the local window corresponding to the above pixel can be determined first. The local window can be a window of a preset size containing the above pixel. For example, the preset size can be 8 Then, the standard deviation of each pixel value in the local window can be determined as the local complexity information corresponding to the pixel.

[0071] The seventh sub-step is to generate an initial energy value corresponding to the pixel based on the color emotion information, the spatial weight information, and the local complexity information. In practice, the initial energy value of the pixel can be determined as the product of the color emotion information, the spatial weight information, and the local complexity information.

[0072] The third step is to normalize the initial energy values ​​of each pixel to obtain the energy value of each pixel. In practice, the normalized energy value can be determined as the ratio of the initial energy value of each pixel to the maximum initial energy value among the initial energy values ​​of each pixel.

[0073] The fourth step is to determine each pixel with a determined energy value as an energy field map.

[0074] The first through fourth steps described above, as an inventive feature of an embodiment of the present disclosure, address the technical issue that when generating an energy field map for an image, attention is often focused solely on the image's basic color properties, while differences in the spatial importance of pixels and the richness of detail in local regions are ignored. This results in a failure to organically integrate information from multiple dimensions, leading to poor accuracy in the generated energy field map. Factors contributing to the poor accuracy of the generated energy field map are often as follows: When generating an energy field map for an image, attention is often focused solely on the image's basic color properties, while differences in the spatial importance of pixels and the richness of detail in local regions are ignored, leading to a failure to organically integrate information from multiple dimensions. Addressing these factors can improve the accuracy of the generated energy field map. To achieve this, the present disclosure introduces color emotion values, enabling the generated energy field map to not only display the image's basic content but also convey the underlying emotional layers. Furthermore, by utilizing a spatial weight matrix and local complexity calculation, it effectively emphasizes the main image components and highlights areas rich in detail. This allows for the organic integration of information from multiple dimensions, improving the accuracy of the generated energy field map. This, in turn, provides more information input for subsequent visual reconstruction tasks, helping to improve their accuracy.

[0075] Step 108 : generating a visual reconstruction image corresponding to the original image based on the generated energy field map, the original image, and the residual vision area.

[0076] In some embodiments, the execution entity may generate a visual reconstruction map corresponding to the original image based on the generated energy field map, the original image, and the residual vision area. In practice, the visual reconstruction map corresponding to the original image may be generated based on the generated energy field map, the original image, the residual vision area, and a pre-trained visual reconstruction model. The visual reconstruction model may be a neural network model that takes the energy field map, the original image, and the residual vision area as input and outputs a visual reconstruction map corresponding to the original image. For example, the neural network model may be a U-Net, a conditional generative adversarial network (cGANs), or a variational autoencoder (VAEs). The training samples of the visual reconstruction model may include a sample original image, an energy field map corresponding to the sample original image, the user's residual vision area, and a sample visual reconstruction map. The sample visual reconstruction map may be obtained by an expert repairing the sample original image based on the user's residual vision area.

[0077] As an example, the visual reconstruction image generated based on the energy field map, the original image and the residual vision area can be referred to Figure 4 .

[0078] In some optional implementations of some embodiments, the execution entity may generate a visual reconstruction image corresponding to the original image according to the generated energy field map, the original image, the residual vision area, and a pre-trained visual reconstruction model through the following steps:

[0079] The first step is to resize the residual vision area based on the image size of the energy field map to obtain an adjusted residual vision area. In practice, in response to determining that the image size of the energy field map is the same as the image size of the visual field image corresponding to the residual vision area, the residual vision area can be directly determined as the adjusted residual vision area. In response to determining that the image size of the energy field map is different from the image size of the visual field image corresponding to the residual vision area, the image size of the visual field image can be adjusted to the image size of the energy field map. Then, the residual vision area in the adjusted visual field image can be determined as the adjusted residual vision area.

[0080] The second step is to superimpose the adjusted residual vision area on the energy field map to obtain a superimposed image. In practice, the adjusted residual vision area can be marked on the energy field map to obtain a superimposed image.

[0081] The third step is to perform energy value adjustment on the superimposed image to obtain an energy value-adjusted image. In practice, for each pixel in the superimposed image outside the adjusted residual vision area, the adjusted energy value can be determined by multiplying the energy value corresponding to the pixel by a preset gain factor. The preset gain factor can be a value greater than 1. For example, the preset gain factor can be 1.5. Finally, the superimposed image after energy value adjustment can be determined as the energy value-adjusted image. This ensures that the importance of the normal vision area is preserved while increasing the importance of the impaired area to facilitate subsequent processing.

[0082] In the fourth step, the original image and the energy-adjusted image are input into the input layer of the visual reconstruction model to obtain the original image tensor and the energy-adjusted image tensor. The visual reconstruction model may include an input layer, a feature extraction network, an attention mechanism layer, a visual restoration layer, and an image optimization layer. The input layer may convert the input images into tensor forms suitable for subsequent processing.

[0083] The fifth step is to input the original image tensor into the feature extraction network to obtain a multi-scale feature map. The feature extraction network can be a convolutional neural network, which can apply a series of 3 Convolutional kernels of size 3 are used, and each convolutional layer is followed by a ReLU activation function to increase nonlinearity. Max pooling can be used every few layers to reduce spatial dimensionality while maintaining feature richness. The feature extraction network can extract multi-level features such as color, texture, and edges from the image. Multiple convolutional and pooling layers are used to capture information at different levels of abstraction.

[0084] In the sixth step, the multi-scale feature map and the energy-adjusted image tensor are input into the attention mechanism layer to generate a weighted feature map. The attention mechanism layer can employ a custom attention mechanism, such as soft attention or hard attention, to dynamically adjust the contribution of each location in the multi-scale feature map based on the importance weight of the energy field map. This guides the model to focus on key areas based on the energy field map, particularly those crucial for users with impaired vision.

[0085] In the seventh step, the weighted feature map is fed into the visual restoration layer to generate the initial restored image. This restoration layer leverages the generator component of generative adversarial networks (GANs) to perform image restoration by learning the mapping relationship between a large number of real images and damaged images. This allows for specialized processing of damaged visual fields, such as filling in missing information and enhancing blurred areas.

[0086] In step 8, the initial inpainted image is fed into the image optimization layer to produce a visual reconstruction. This layer uses a super-resolution network or a finely tuned convolutional network to increase image resolution and reduce noise, ensuring the final output is as close to the original high-quality image as possible. This further enhances image quality and ensures clear and natural details.

[0087] The aforementioned steps 1-8, as an inventive feature of an embodiment of the present disclosure, address the technical issue that conventional image enhancement methods often struggle to provide effective assistance to patients with complex visual field loss patterns, resulting in a poor user experience for the generated visual reconstructed images. Factors contributing to the poor user experience of the generated visual reconstructed images are often as follows: Conventional image enhancement methods often struggle to provide effective assistance to patients with complex visual field loss patterns. Addressing these factors can improve the user experience of the generated visual reconstructed images. To achieve this, the present disclosure fuses the residual vision area with the energy field map and specifically enhances the importance of key information, ensuring that the importance of normal visual areas is preserved while simultaneously enhancing the importance of impaired areas for subsequent processing. This effectively improves the efficiency of visual information transmission and reduces the risk of information loss or misunderstanding due to visual impairment. Furthermore, the energy field map guides attention mechanisms, focusing on areas critical to understanding and interpreting the image, even those located near the patient's blind spot. Finally, image optimization is performed to ensure that the final output is as close as possible to the original high-quality image, further enhancing image quality and ensuring clear and natural details. This improves the user experience of the generated visual reconstructed images.

[0088] Step 109: Display the visual reconstruction image.

[0089] In some embodiments, the execution entity may display the visual reconstruction image. In practice, the visual reconstruction image may be displayed on a display screen of the execution entity.

[0090] The above-described embodiments of the present disclosure have the following beneficial effects: The visual reconstruction methods of some embodiments of the present disclosure eliminate the need for customized equipment and replacement of optical components. Specifically, the high equipment costs and frequent replacement of optical components are due to the difficulty and high cost of customizing visual acuity devices for visual impairment, and the frequent replacement of optical components as the user's vision changes. Based on this, the visual reconstruction methods of some embodiments of the present disclosure first acquire an original image. Thus, the original image can serve as the image to be visually reconstructed. Then, an image of the visual field of a target user is acquired. Next, the visual field image is binarized to obtain a binarized visual field image. Next, the binarized visual field image is denoised to obtain a denoised visual field image. Next, contour detection is performed on the denoised visual field image to obtain a contour detection result. Next, based on the contour detection result, a residual vision area corresponding to the target user is generated. Thus, the residual vision area can represent the target user's normal vision area. Next, an energy field map is generated based on the original image. The energy value of each pixel in the energy field map can identify the importance or significance of that pixel, thereby serving as a guide in the visual reconstruction process. Then, based on the generated energy field map, the original image, and the residual vision area, a visual reconstruction image corresponding to the original image is generated. Finally, the visual reconstruction image is displayed. This allows the user to observe a relatively complete image. Because the image presented to the user has undergone visual reconstruction, no optical processing is required, eliminating the need for customized equipment or replacement of optical components.

[0091] Further references Figure 5 As an implementation of the methods shown in the above figures, the present disclosure provides some embodiments of a visual reconstruction device. These device embodiments are similar to Figure 1 Corresponding to the method embodiments shown, the device can be specifically applied to various electronic devices.

[0092] like Figure 5As shown, the visual reconstruction device 500 of some embodiments includes: a first acquisition unit 501, a second acquisition unit 502, a binarization processing unit 503, a denoising unit 504, a contour detection unit 505, a first generation unit 506, a second generation unit 507, a third generation unit 508 and a display unit 509. Among them, the first acquisition unit 501 is configured to acquire the original image; the second acquisition unit 502 is configured to acquire the field of view image of the target user; the binarization processing unit 503 is configured to perform binarization processing on the above-mentioned field of view image to obtain a binarized field of view image; the denoising unit 504 is configured to perform denoising processing on the above-mentioned binarized field of view image to obtain a denoised field of view image; the contour detection unit 505 is configured to perform contour detection on the above-mentioned denoised field of view image to obtain a contour detection result; the first generation unit 506 is configured to generate the residual vision area corresponding to the above-mentioned target user based on the above-mentioned contour detection result; the second generation unit 507 is configured to generate an energy field map according to the above-mentioned original image; the third generation unit 508 is configured to generate a visual reconstruction map corresponding to the above-mentioned original image based on the generated energy field map, the above-mentioned original image and the above-mentioned residual vision area; the display unit 509 is configured to display the above-mentioned visual reconstruction map.

[0093] Optionally, the third generating unit 508 may be further configured to generate a visual reconstruction map corresponding to the original image based on the generated energy field map, the original image, the residual vision area and a pre-trained visual reconstruction model.

[0094] Optionally, the first generation unit 506 can be further configured to: determine the various contour areas included in the above-mentioned contour detection results; based on the above-mentioned various contour areas, within the image size range of the above-mentioned field of view image, extract the image area that meets the preset maximum inscribed condition as the residual vision area corresponding to the above-mentioned target user.

[0095] Optionally, the second generation unit 507 can be further configured to: input the above-mentioned original image into a pre-trained energy field map generation model to obtain an energy field map, wherein the above-mentioned energy field map generation model includes an input layer, a position encoding layer, each depth-separable convolution layer, each linear layer and an output layer, the above-mentioned position encoding layer is residually connected to the first depth-separable convolution layer in each of the above-mentioned depth-separable convolution layers, and the two adjacent depth-separable convolution layers in each of the above-mentioned depth-separable convolution layers are residually connected.

[0096] Optionally, the second generation unit 507 can be further configured to: perform color channel segmentation processing on the above-mentioned original image to obtain a set of segmented images corresponding to each color channel, wherein each segmented image in the above-mentioned segmented image set corresponds to a color channel; for each segmented image in the above-mentioned segmented image set, perform the following steps: generate the gradient information corresponding to the horizontal direction of each pixel in the above-mentioned segmented image as horizontal gradient information; generate the gradient information corresponding to the vertical direction of each pixel in the above-mentioned segmented image as vertical gradient information; for each pixel in the above-mentioned segmented image, generate the comprehensive gradient information of the above-mentioned pixel in the corresponding color channel as the energy value of the above-mentioned pixel in the corresponding color channel based on the horizontal gradient information and vertical gradient information corresponding to the above-mentioned pixel; determine each pixel with determined energy value as the initial energy field map corresponding to the above-mentioned color channel; normalize the energy value of each pixel in the above-mentioned initial energy field map to obtain the normalized initial energy field map as the energy field map corresponding to the above-mentioned color channel.

[0097] It is understood that the units described in the device 500 are similar to those in the reference Figure 1 Therefore, the operations, features and beneficial effects described above for the method are also applicable to the device 500 and the units included therein, and will not be repeated here.

[0098] Reference below Figure 6 , which shows a structural schematic diagram of an electronic device 600 (such as a head-mounted display device or a terminal device) suitable for implementing some embodiments of the present disclosure. Figure 6 The electronic device shown is only an example and should not bring any limitation to the functions and scope of use of the embodiments of the present disclosure.

[0099] like Figure 6 As shown, electronic device 600 may include a processing device 601 (e.g., a central processing unit, a graphics processing unit, etc.), which can perform various appropriate actions and processes according to programs stored in a read-only memory (ROM) 602 or programs loaded from a storage device 608 into a random access memory (RAM) 603. RAM 603 also stores various programs and data required for the operation of electronic device 600. Processing device 601, ROM 602, and RAM 603 are interconnected via a bus 604. An input / output (I / O) interface 605 is also connected to bus 604.

[0100] Typically, the following devices may be connected to the I / O interface 605: an input device 606 including, for example, a touch screen, a touchpad, a keyboard, a mouse, a camera, a microphone, an accelerometer, a gyroscope, etc.; an output device 607 including, for example, a liquid crystal display (LCD), a speaker, a vibrator, etc.; and a communication device 609. The communication device 609 may allow the electronic device 600 to communicate with other devices wirelessly or by wire to exchange data. Figure 6 The electronic device 600 is shown with various devices, but it should be understood that it is not required to implement or possess all of the devices shown. More or fewer devices may be implemented or possessed instead. Figure 6 Each block shown in the figure may represent one device, or may represent multiple devices as needed.

[0101] In particular, according to some embodiments of the present disclosure, the processes described above with reference to the flowcharts can be implemented as computer software programs. For example, some embodiments of the present disclosure include a computer program product comprising a computer program carried on a computer-readable medium, the computer program comprising program code for executing the method shown in the flowchart. In some such embodiments, the computer program can be downloaded and installed from a network via the communication device 609, or installed from the storage device 608, or installed from the ROM 602. When the computer program is executed by the processing device 601, the above-mentioned functions defined in the method of some embodiments of the present disclosure are performed.

[0102] It should be noted that the computer-readable medium described in some embodiments of the present disclosure may be a computer-readable signal medium or a computer-readable storage medium, or any combination thereof. Computer-readable storage media may include, for example, but not limited to, electrical, magnetic, optical, electromagnetic, infrared, or semiconductor systems, devices, or components, or any combination thereof. More specific examples of computer-readable storage media may include, but are not limited to, an electrical connection having one or more conductors, a portable computer disk, a hard disk, random access memory (RAM), read-only memory (ROM), erasable programmable read-only memory (EPROM or flash memory), optical fiber, a portable compact disk read-only memory (CD-ROM), an optical storage device, a magnetic storage device, or any suitable combination thereof. In some embodiments of the present disclosure, a computer-readable storage medium may be any tangible medium containing or storing a program that can be used by or in conjunction with an instruction execution system, device, or component. Furthermore, in some embodiments of the present disclosure, a computer-readable signal medium may include a data signal propagated in baseband or as part of a carrier wave, carrying computer-readable program code. This propagated data signal may take a variety of forms, including, but not limited to, electromagnetic signals, optical signals, or any suitable combination thereof. A computer-readable signal medium may also be any computer-readable medium other than a computer-readable storage medium that can transmit, propagate, or transport a program for use by or in connection with an instruction execution system, apparatus, or device. Program code embodied on a computer-readable medium may be transmitted using any suitable medium, including but not limited to wire, optical cable, RF (radio frequency), or any suitable combination thereof.

[0103] In some embodiments, the client and server can communicate using any currently known or later developed network protocol, such as HTTP (HyperText Transfer Protocol), and can be interconnected with any form or medium of digital data communication (e.g., a communication network). Examples of communication networks include a local area network ("LAN"), a wide area network ("WAN"), an internet (e.g., the Internet), and a peer-to-peer network (e.g., an ad hoc peer-to-peer network), as well as any currently known or later developed network.

[0104] The computer-readable medium may be included in the electronic device, or may exist independently without being incorporated into the electronic device. The computer-readable medium carries one or more programs. When the one or more programs are executed by the electronic device, the electronic device: acquires an original image; acquires a field of view image of a target user; binarizes the field of view image to obtain a binarized field of view image; denoises the binarized field of view image to obtain a denoised field of view image; performs contour detection on the denoised field of view image to obtain a contour detection result; generates a residual vision area corresponding to the target user based on the contour detection result; generates an energy field map according to the original image; generates a visual reconstruction map corresponding to the original image according to the generated energy field map, the original image, and the residual vision area; and displays the visual reconstruction map.

[0105] Computer program code for performing the operations of some embodiments of the present disclosure may be written in one or more programming languages, or a combination thereof, including object-oriented programming languages ​​such as Java, Smalltalk, C++, and conventional procedural programming languages ​​such as "C" or similar programming languages. The program code may be executed entirely on the user's computer, partially on the user's computer, as a stand-alone software package, partially on the user's computer and partially on a remote computer, or entirely on a remote computer or server. In cases involving a remote computer, the remote computer may be connected to the user's computer through any type of network, including a local area network (LAN) or a wide area network (WAN), or may be connected to an external computer (e.g., through the Internet using an Internet service provider).

[0106] The flowcharts and block diagrams in the accompanying drawings illustrate the possible implementation architecture, functions and operations of the systems, methods and computer program products according to various embodiments of the present disclosure. In this regard, each box in the flowchart or block diagram can represent a module, program segment, or a part of code, and the module, program segment, or a part of code contains one or more executable instructions for realizing the specified logical function. It should also be noted that in some alternative implementations, the functions marked in the box can also occur in a different order than that marked in the accompanying drawings. For example, two boxes represented in succession can actually be executed substantially in parallel, and they can sometimes be executed in the opposite order, depending on the functions involved. It should also be noted that each box in the block diagram and / or flowchart, and the combination of the boxes in the block diagram and / or flowchart, can be implemented with a dedicated hardware-based system that performs the specified function or operation, or can be implemented with a combination of dedicated hardware and computer instructions.

[0107] The units described in some embodiments of the present disclosure may be implemented by software or hardware. The units described may also be provided in a processor. For example, they may be described as follows: a processor including a first acquisition unit, a second acquisition unit, a binarization unit, a denoising unit, a contour detection unit, a first generation unit, a second generation unit, a third generation unit, and a display unit. The names of these units do not, in some cases, constitute limitations on the units themselves. For example, the first acquisition unit may also be described as a "unit for acquiring an original image."

[0108] The functions described above herein may be performed, at least in part, by one or more hardware logic components. For example, and without limitation, exemplary types of hardware logic components that may be used include: field programmable gate arrays (FPGAs), application specific integrated circuits (ASICs), application specific standard products (ASSPs), systems on chips (SOCs), complex programmable logic devices (CPLDs), and the like.

[0109] Some embodiments of the present disclosure further provide a computer program product, including a computer program, which implements any of the above-mentioned visual reconstruction methods when executed by a processor.

[0110] The above descriptions are merely some preferred embodiments of the present disclosure and illustrate the underlying technical principles. Those skilled in the art should understand that the scope of the invention encompassed by the embodiments of the present disclosure is not limited to technical solutions formed by specific combinations of the aforementioned technical features. It also encompasses other technical solutions formed by any combination of the aforementioned technical features or their equivalents, without departing from the aforementioned inventive concept. For example, a technical solution formed by replacing the aforementioned features with (but not limited to) technical features with similar functions disclosed in the embodiments of the present disclosure.

Claims

1. A visual reconstruction method, comprising: Get the original image; Acquire the target user's field of view image; performing binarization processing on the visual field image to obtain a binarized visual field image; performing denoising processing on the binarized visual field image to obtain a denoised visual field image; Performing contour detection on the denoised visual field image to obtain a contour detection result; Based on the contour detection result, generating a residual vision area corresponding to the target user; generating an energy field map according to the original image; generating a visual reconstruction image corresponding to the original image according to the generated energy field map, the original image, and the residual vision area; The visual reconstruction image is displayed.

2. The method according to claim 1, wherein Generating a visual reconstruction image corresponding to the original image based on the generated energy field map, the original image, and the residual vision area includes: A visual reconstruction map corresponding to the original image is generated according to the generated energy field map, the original image, the residual vision area and a pre-trained visual reconstruction model.

3. The method according to claim 1, wherein Generating the residual vision area corresponding to the target user based on the contour detection result includes: Determining each contour area included in the contour detection result; Based on the respective contour areas, within the image size range of the visual field image, an image area that meets a preset maximum inscribed condition is extracted as the residual vision area corresponding to the target user.

4. The method according to claim 1, wherein Generating an energy field map according to the original image includes: The original image is input into a pre-trained energy field map generation model to obtain an energy field map, wherein the energy field map generation model includes an input layer, a position encoding layer, each depth-wise separable convolution layer, each linear layer and an output layer, the position encoding layer is residually connected to the first depth-wise separable convolution layer in each depth-wise separable convolution layer, and the two adjacent depth-wise separable convolution layers in each depth-wise separable convolution layer are residually connected.

5. The method according to claim 1, wherein Generating an energy field map according to the original image includes: Performing color channel segmentation processing on the original image to obtain a set of segmented images corresponding to each color channel, wherein each segmented image in the set of segmented images corresponds to one color channel; For each segmented image in the segmented image set, performing the following steps: generating horizontal gradient information corresponding to each pixel in the segmented image as horizontal gradient information; Generating vertical gradient information corresponding to each pixel in the segmented image as vertical gradient information; For each pixel in the segmented image, generating, according to the horizontal gradient information and the vertical gradient information corresponding to the pixel, comprehensive gradient information of the pixel in the corresponding color channel as the energy value of the pixel in the corresponding color channel; Determining each pixel having the determined energy value as an initial energy field map corresponding to the color channel; The energy value of each pixel in the initial energy field map is normalized to obtain the normalized initial energy field map as the energy field map corresponding to the color channel.

6. A visual reconstruction device comprising: A first acquisition unit is configured to acquire an original image; a second acquiring unit, configured to acquire a visual field image of a target user; A binarization processing unit is configured to perform binarization processing on the field of view image to obtain a binarized field of view image; a denoising unit configured to perform denoising processing on the binarized visual field image to obtain a denoised visual field image; a contour detection unit, configured to perform contour detection on the denoised visual field image to obtain a contour detection result; a first generating unit configured to generate a residual vision area corresponding to the target user based on the contour detection result; A second generating unit is configured to generate an energy field map according to the original image; a third generating unit, configured to generate a visual reconstruction map corresponding to the original image based on the generated energy field map, the original image and the residual vision area; A display unit is configured to display the visual reconstruction image.

7. An electronic device comprising: one or more processors; a storage device having one or more programs stored thereon, When the one or more programs are executed by the one or more processors, the one or more processors implement the method according to any one of claims 1 to 5.

8. A computer-readable medium having a computer program stored thereon, wherein: When the computer program is executed by a processor, the method according to any one of claims 1 to 5 is implemented.

9. A computer program product comprising a computer program, which, when executed by a processor, implements the method according to any one of claims 1 to 5.

Citation Information

Patent Citations

  • Digital therapeutic corrective spectacles

    CN111511318A

  • Vision assistance method, storage medium and vision assistance device

    CN119165947A