Image device and image processing method
By using a fully convolutional neural network and labeling outline diagram technology in the image processing method, the problem of difficulty in identifying complex structural objects in the prior art is solved, and real-time application of blind people to recognize scene objects is realized.
Patent Information
- Application Number
- CN202011171651.9
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Filing Date
- 2020-10-28
- Publication Date
- 2025-06-13
- Estimated Expiration
- 2040-10-28
AI Technical Summary
Existing image processing methods are difficult to effectively identify and segment objects with complex structures and large internal differences, and are difficult to meet the requirements of real-time applications.
By marking different characters on the template outline diagrams corresponding to objects of different categories, and superimposedly generate a marked outline diagram. The fully convolutional neural network identification mechanism is used to identify scene objects and display their outline curves and classification characters.
It realizes that blind people recognize different objects in the scene through characters, improves their ability to recognize objects in complex structures, and is suitable for real-time applications.
Smart Images

Figure CN112348067B_ABST
Abstract
Description
Technical Field
[0001] The present invention relates to an image device and an image processing method, and particularly to an image device and an image processing method that are beneficial for blind people to identify scene objects. Background Art
[0002] Vision is the main perceptual organ for humans to obtain information. The number of blind people worldwide has exceeded 50 million, and nearly 5 million people suffer from damaged visual organs due to external factors every year. With the progress of technology, artificial vision prostheses have gradually been applied to the perception and processing of visual images. Blind people can perceive the visual prosthesis by electrically stimulating the visual cortex of the brain to see image light spots, and develop the remaining functional parts of their visual systems to restore some visual perceptions. Based on the research method of graph theory, the image segmentation problem is regarded as a vertex partitioning problem of a graph. The general method is to map the image to be segmented into a weighted undirected graph. In order to obtain better segmentation results, it is usually necessary to construct a complex cost function, and the algorithm has a high time complexity, making it difficult to meet the requirements of real-time applications. Generally speaking, if the past image processing method is adopted, it is first necessary to initialize a rough clustering using a pixel clustering-based method. Then, in an iterative manner, pixels with similar features such as color, brightness, and texture are clustered into the same superpixel by approaching in spatial distance and directly converging to obtain the final image segmentation result. However, such a processing method is not suitable for segmenting objects with relatively complex structures and large internal differences. Summary of the Invention
[0003] In view of this, the present invention proposes an image device and an image processing method. By labeling different characters on the template contours corresponding to different categories of objects and superimposing the template contours to generate a labeled contour map, blind people can identify different objects in the scene through different characters.
[0004] The present invention provides an image device suitable for assisting blind people to identify scenes. This image device includes an image processing module. The image processing module, through a fully convolutional neural network recognition mechanism, identifies which category of entity object at least one scene object carried on an input image corresponding to the scene is. The image processing module displays a contour curve relative to the at least one scene object according to the position of the at least one scene object in the scene. The image processing module classifies the at least one scene object using different characters and displays the characters representing different categories in the line area of the contour curve.
[0005] The present invention provides an image processing method applicable to assisting a blind person in recognizing a scene, including the following steps: through a fully convolutional neural network identification mechanism, identifying which category of an entity object is carried on at least one scene object in an input image corresponding to the scene; displaying a contour curve relative to the at least one scene object for the position of the at least one scene object in the scene; classifying the at least one scene object using different characters; and displaying different characters representing different category scene objects in a line area of the contour curve corresponding to the scene object. Description of the Drawings
[0006] Figure 1A is a schematic structural diagram of an image device according to an embodiment of the present invention.
[0007] Figure 1B is a flowchart of an image processing method according to an embodiment of the present invention.
[0008] Figure 1C is a flowchart of an image processing method according to an embodiment of the present invention.
[0009] Figure 2 is an input image corresponding to a scene.
[0010] Figure 3 is a schematic diagram of a microelectrode array.
[0011] Figure 4 is an object area map corresponding to a scene.
[0012] Figure 5A and Figure 5B is a template area map corresponding to a scene.
[0013] Figure 6A and Figure 6B is a template contour map corresponding to a scene.
[0014] Figure 7A and Figure 7B is a marked contour map corresponding to a scene.
[0015] Figure 8 is a marked contour map corresponding to a scene.
[0016]
Symbol Description
[0017] 110: Image device
[0018] 120: Camera module
[0019] 130: Microelectrode array
[0020] 140: Image processing module
[0021] 200: Input image
[0022] 210: Object
[0023] 220: Object
[0024] 230: Object
[0025] S110 - S130: Process Steps
[0026] S132 - S138: Process Steps Detailed Embodiment
[0027] To make the above and other objects, features, and advantages of the present invention more obvious and understandable, the following preferred embodiments are specifically presented, and in conjunction with the accompanying drawings, the detailed description is as follows. It should be noted that the description in this section is the best way to implement the present invention, aiming to illustrate the spirit of the present invention rather than to limit the protection scope of the present invention. It should be understood that the following embodiments can be implemented through software, hardware, firmware, or any combination of the above.
[0028] The present invention provides an image device suitable for assisting a blind person in identifying physical objects in indoor or outdoor scenes (in this specification, an object is an object). Each template contour corresponding to different categories of physical objects is marked with different characters, and a marked contour diagram is superimposed to enable the blind person to identify different objects in the scene through different characters.
[0029] Figure 1A is a schematic structural diagram of the image device according to an embodiment of the present invention. The image device 110 includes a camera module 120, a microelectrode array 130, and an image processing module 140. The camera module 120 can be a camera, which is configured on the blind person, for example, using a head-mounted mechanical structure to configure the camera on the forehead of the blind person to capture the surrounding environment. The camera module 120 captures an indoor or outdoor scene to obtain an original image corresponding to the scene.
[0030] A visual prosthesis is an implantable medical electronic device whose function is to restore the vision of severely blind patients to a certain extent. The visual prosthesis technology takes advantage of the fact that most blind people often have only a part of the visual pathway affected, while the structure and function of the remaining nerve tissues are still intact. By applying specific artificial electrical signal stimuli to the intact part of the blind person's visual pathway, nerve cells are stimulated to simulate the effect of natural light stimulation, enabling the blind person to have visual sensations.
[0031] The visual prosthesis according to an embodiment of the present invention can be represented by a microelectrode array 130, as Figure 3 shown. The microelectrode array 130 according to an embodiment of the present invention is used to prompt the blind person about the configuration of physical objects in the indoor or outdoor scene space. As Figure 3The microelectrode array 130, each black dot in the figure is a microelectrode, and these microelectrodes directly stimulate the intact part of the visual pathway in the occipital visual cortex of the blind person's brain to help the blind person restore partial visual perception. The blind person can understand the current indoor or outdoor environmental furnishings through the image display of the microelectrode array 130.
[0032] The image processing module 140 of the embodiment of the present invention is coupled to the camera module 120 and the microelectrode array 130. Among them, the image processing module 140 receives and analyzes the input image 200 captured by the camera module 120. The image processing module 140 can be implemented by a Field Programmable Gate Array (FPGA). If the product is embodied, it can also be presented in the form of a Graphics Processing Unit (GPU) or in the form of a Vision Processing Unit (VPU).
[0033] Figure 1B , 1C is a flowchart of the image processing method of the embodiment of the present invention.
[0034] As Figure 1B shown, the image processing module 140 first adjusts the size of the captured original image to generate an input image (step S110). Then, the image processing module 140 displays different types of scene objects with different pixel values in the object area map (S120). Finally, the image processing module 140 displays different characters representing different types of scene objects in the line area corresponding to the contour curve of the scene object in the object area map (S130). The image processing module 140 first executes step S110.
[0035] In step S110, the image processing module 140 first adjusts the size of the captured original image to generate an input image. Specifically, the image processing module 140 first adjusts the size of the original image captured by the camera module 120 for convenient processing. For example, it first adjusts the original image captured by the camera module 120 to generate an input image with a length of 256 pixels and a width of 256 pixels. In other embodiments of the present invention, the size of the original image may not be changed. For example, the original image may be directly used as the input image for subsequent image processing. Figure 2 Shown is an input image 200 corresponding to an indoor scene in the embodiment of the present invention. From Figure 2 the input image 200, it can be seen that there are two beds (210 and 220) and a cabinet 230 in the room at this time. Then, the image processing module 140 executes step S120.
[0036] In step S120, the image processing module 140 displays scene objects of different categories with different pixel values in the object area map. Specifically, the image processing module 140 uses a fully convolutional neural network (CNN) recognition mechanism to identify which category of entity object at least one scene object carried on the input image 200 corresponding to the indoor scene is. Taking Figure 2 the input image 200 as an example, the image processing module 140 analyzes the input image 200. The input image 200 has input features such as the shape, size, color, and texture of the beds 210 and 220, and also has input features such as the shape, size, color, and texture of the cabinet 230. The image processing module 140 uses the fully convolutional neural network recognition mechanism to analyze these input features and identifies that there are beds 210, 220 and a cabinet 230 in the room at this time. In one embodiment, the image processing module 140 uses the fully convolutional neural network to extract the pixel semantic information of the input image 200, and based on the above pixel semantic information, determines which category of entity object at least one scene object carried on the input image 200 is.
[0037] The fully convolutional neural network recognition mechanism automatically extracts features on the image or the image by constructing multiple convolutional layers. Generally, the shallower convolutional layers at the front use smaller receptive fields and can learn some specific features of the image (such as color, shape, texture features), while the deeper convolutional layers at the back use larger receptive fields and can learn more abstract features (such as physical size, position, direction, etc.). Therefore, the fully convolutional neural network has been widely used in the fields of image classification and image detection.
[0038] Figure 4 is an object area map according to an embodiment of the present invention. When the image processing module 140 identifies which category of entity object one or more scene objects carried on the input image 200 are, the image processing module 140 displays various scene objects of different categories with different pixel values in the object area map, as Figure 4 shown. Among them, one pixel value represents one category of entity object. The pixel value refers to the brightness of a single pixel point, and the larger the pixel value, the brighter it is. In one embodiment, the range of the pixel value is 0 to 255.
[0039] For example, after the image processing module 140 identifies that the objects 210 and 220 belong to the category of beds and the object 230 belongs to the category of cabinets, in Figure 4In the object area map, the image processing module 140 renders the objects 210 and 220 with the same pixel value, renders the object 230 with another pixel value, and in this object area map, the positions where the objects 210, 220, and 230 are displayed are the positions of the objects relative to the indoor scene.
[0040] Then, in step S130, the image processing module 140 displays different characters representing different category scene objects in the line area corresponding to the contour curve of the scene object in the object area map. The following will be combined with Figure 1C Step S130 will be described in detail.
[0041] As Figure 1C shown, the image processing module 140 lists the scene objects of the same category in the object area map in the same template area map (S132). Then, the image processing module 140 performs edge detection on the display areas of the scene objects in each template area map to generate corresponding template contour maps respectively (S134). Then, the image processing module 140 displays the characters representing different categories at the central positions in the line areas of the contour curves of each template contour map (S136). Finally, the image processing module 140 superimposes each template contour map with different characters to generate a marked contour map (S138). The image processing module 140 first executes step S132.
[0042] In step S132, the image processing module 140 lists the scene objects of the same category in the object area map in the same template area map. Specifically, the image processing module 140 performs a template binarization process to separate the scene objects with different pixel values in the object area map, and lists the scene objects with the same pixel value in the same template area map. Figure 5A And Figure 5B are the template area maps corresponding to an indoor scene, representing different category scene objects respectively. Figure 5A Present the same category of objects: bed 210 and bed 220. Figure 5B Present the same category of objects: cabinet 230.
[0043] The image processing module 140 controls the display areas of the scene objects in Figure 5A And Figure 5B The template area map is equal to the coverage area of the scene object in the input image. That is to say, the position of the tile in the template area map is the relative position of the scene object in the indoor scene. Among them, in the same template area map, the image processing module 140 sets the pixel value of its background to the first pixel value, and sets the pixel value of the display area of the scene object in the template area map to the second pixel value. In one embodiment, the first pixel value is 0 and the second pixel value is 255. As Figure 5A And Figure 5BAs shown in the template area diagram, the background is black, that is, the pixel value is the first pixel value. The object is white, that is, the pixel value is the second pixel value.
[0044] Then, in step S134, the image processing module 140 performs edge detection on the display areas of the scene objects in each template area diagram, and respectively generates corresponding template contour diagrams. Specifically, the image processing module 140 performs edge detection on the display areas of the scene objects in each template area diagram, that is, for Figure 5A objects 210 and 220, and for Figure 5B object 230 to perform edge detection. The image processing module 140 controls the pixel value of the contour curve of the tile display area to be the second pixel value, that is, the contour of the tile is displayed in white, and the width of the contour curve is the first pixel width. The image processing module 140 also sets the pixel values of the parts in the display area other than the contour curve to the first pixel value, that is, the parts other than the contour are displayed in black, so as to generate a template contour diagram. Each template area diagram respectively has a corresponding template contour diagram, which are respectively corresponding to Figure 5A of Figure 6A , and corresponding to Figure 5B of Figure 6B . In one embodiment, the first pixel width is 2 pixel widths.
[0045] Then, in step S136, the image processing module 140 displays characters representing different categories at the central positions within the line areas of the contour curves of each template contour diagram. Specifically, for the template contour diagrams corresponding to different categories of entity objects, such as Figure 6A and Figure 6B representing different categories of entity objects, the image processing module 140 displays characters representing different categories within the line areas of the contour curves, such as Figure 7A and Figure 7B of the marked contour diagrams. As Figure 7A shown, at the central position within the line area of the contour curve, that is, at the central positions of the object 210 tile and the object 220 tile, the character 3 representing the category "bed" is displayed. As Figure 7B shown, at the central position of the object 230 tile, the character 4 representing the category "cabinet" is displayed.
[0046] In one embodiment of the present invention, the calculation method of the central position is that the image processing module 140 first finds a minimum circumscribed rectangle of the tile contour curve, and the central position is defined as the diagonal center of the minimum circumscribed rectangle.
[0047] Finally, in step S138, the image processing module 140 superimposes each template contour map with different characters to generate a marked contour map. Specifically, the image processing module 140 superimposes each template contour map with different characters to generate a marked contour map. In an embodiment of the present invention, the template contour maps of the 7A and 7B with characters are superimposed, as Figure 8 shown.
[0048] The image processing module 140 displays the template contour map on the microelectrode array 130 for the blind to identify the indoor scene through characters. That is to say, the blind will first understand the meaning represented by the characters through other means in advance, such as being guided by others or feeling the Braille description by hand, to understand that the character 3 represents "bed" and the character 4 represents "cabinet". Then, by using the image device 110 provided by the present invention, the template contour map as Figure 8 shown is displayed on the microelectrode array 130 to assist the blind in knowing what objects are currently in the room and their locations from this template contour map. In one embodiment, the image processing module 140 first adjusts the size of the template contour map according to the display range of the microelectrode array 130. Then, the resized template contour map is displayed on the microelectrode array 130. For example, when the display range of the microelectrode array 130 is 32 pixels in length and 32 pixels in width, the image processing module 140 first adjusts the template contour map to 32 pixels in length and 32 pixels in width. Then, the resized template contour map is displayed on the microelectrode array 130.
[0049] Overall, the image processing module 140 displays a contour curve relative to the object on the microelectrode array 130 according to the position of the object in the scene. The image processing module 140 also classifies multiple scene objects with different characters. On the microelectrode array 130, the image processing module 140 displays the characters representing different categories of objects in the line area of the contour curve. Therefore, the blind can know what objects are in the currently contacted environment from the characters representing each object carried on the microelectrode array 130.
[0050] Although the present invention has been disclosed above with preferred embodiments, it is not intended to limit the present invention. Those skilled in the art can make some changes and modifications without departing from the spirit and scope of the present invention. For example, the systems and methods described in the embodiments of the present invention can be implemented by physical embodiments of hardware, software, or a combination of hardware and software. Therefore, the protection scope of the present invention shall be subject to the scope defined by the appended claims.
Claims
1. An image device is applicable to assist the blind in identifying scene objects, comprising: an image processing module, which, through a fully convolutional neural network recognition mechanism, identifies which category of entity object at least one scene object carried on the input image corresponding to the scene is. The image processing module displays a contour curve relative to the at least one scene object according to the position of the at least one scene object in the scene. The image processing module classifies the at least one scene object using different characters, and displays the characters representing different categories in the line area of the contour curve, and a microelectrode array, coupled to the image processing module, configured on the visual cortex of the blind person's brain, for directly stimulating the intact part of the visual pathway of the blind person's said visual cortex of the brain to help the blind person restore partial visual perception, wherein the contour curve and the characters are displayed on the microelectrode array, wherein the image processing module superimposes template contour maps with different such characters to generate a marked contour map, which is displayed on the microelectrode array for the blind person to identify the scene, wherein, after the image processing module identifies which category of the entity object the at least one scene object carried on the input image is, the image processing module displays the at least one scene object of different categories with different pixel values in an object area map, wherein one pixel value represents one category of the entity object, wherein the image processing module performs a template binarization process to separate each of the scene objects with different pixel values in the object area map, and the scene objects with the same pixel value are listed in the same template area map, wherein the image processing module controls the display area of the scene object in the template area map to be equal to the coverage area of the scene object in the input image. In the same template area map, the image processing module sets the pixel value of its background to a first pixel value, and sets the pixel value of the display area of the scene object in the template area map to a second pixel value, wherein the image processing module performs edge detection on the display area of the at least one scene object in each template area map to generate a corresponding contour curve. The image processing module controls the pixel value of each contour curve to be the second pixel value, and sets the pixel value of the part in the display area that is not the contour curve to the first pixel value to generate the template contour map, and each template area map respectively has a corresponding template contour map.
2. The image device according to claim 1, further comprising: a camera module, coupled to the image processing module, which captures the scene to obtain an original image corresponding to the scene.
3. The image device according to claim 2, wherein, the image processing module first adjusts the size of the original image to generate the input image.
4. The image device according to claim 1, wherein, the image processing module extracts the pixel semantic information of the input image, and identifies which category of the entity object the at least one scene object carried on the input image is according to the pixel semantic information.
5. The image device according to claim 1, wherein, For the template contour diagram corresponding to the entity object of different categories, the image processing module will display the characters representing different categories at the central position within the line area of the contour curve.
6. An image processing method, applicable to assisting the blind in recognizing scene objects, the method comprises: Through a fully convolutional neural network recognition mechanism, identifying which category of entity object at least one scene object carried on the input image corresponding to the scene is; For the position of the at least one scene object in the scene, displaying a contour curve relative to the at least one scene object; Using different characters to classify the at least one scene object; Displaying different characters representing different category scene objects within the line area of the contour curve corresponding to the scene object, wherein the contour curve and the characters are displayed on a microelectrode array, and the microelectrode array is arranged on the intact part of the visual pathway of the visual cortex of the blind person's brain for directly stimulating the visual cortex of the blind person to help the blind person restore partial visual perception; Superimposing the template contour diagrams with different characters to generate a marked contour diagram, which is displayed on the microelectrode array for the blind person to recognize the scene; After identifying which category of entity object the at least one scene object carried on the input image is, displaying the at least one scene object of different categories with different pixel values in the object area diagram, wherein one pixel value represents one category of the entity object; Performing a template binarization process to separate the scene objects with different pixel values in the object area diagram, and listing the scene objects with the same pixel value in the same template area diagram; Controlling the display area of the scene object in the template area diagram to be equal to the covering area of the scene object in the input image, wherein, in the same template area diagram, the image processing module sets the pixel value of its background to the first pixel value, and sets the pixel value of the display area of the scene object in the template area diagram to the second pixel value; and Performing edge detection on the display area of the at least one scene object in each template area diagram to generate the corresponding contour curve, and the image processing module controls the pixel value of the contour curve of the display area to be the second pixel value, and sets the pixel value of the part of the display area that is not the contour curve to the first pixel value to generate the template contour diagram, and each template area diagram respectively has the corresponding template contour diagram.
7. The image processing method according to claim 6, further comprises: Taking a picture of the scene to obtain an original image corresponding to the scene.
8. The image processing method according to claim 7, further comprises: First adjusting the size of the original image to generate the input image.
9. The image processing method according to claim 6, further comprises: Extracting the pixel semantic information of the input image, and identifying which category of entity object the at least one scene object carried on the input image is according to the pixel semantic information.
10. The image processing method according to claim 6, further comprises: For the template contour diagram corresponding to the entity object of different categories, the characters representing different categories are displayed at the central position within the line area of the contour curve.
Citation Information
Patent Citations
Cross-domain large-range scene generation method
CN110147733A
Crop segmentation method based on unmanned aerial vehicle aerial images
CN111259898A
Visual compensation method based on neural network and tactile dot matrix
CN111428583A
A neuroprosthetic system and method for substituting a sensory modality of a mammal by high-density electrical stimulation of a region of the cerebral cortex
WO2020043790A1