Visual prosthesis control method and system and storage medium
By identifying targets and extracting regions of interest from grayscale images, grouping and sorting stimulation current values, and controlling the electrode array to release stimulation current, the problems of redundant information transmission and electrode channel interference in visual prostheses are solved, thereby improving the brain's perception quality of targets.
Patent Information
- Application Number
- CN202511479448.0
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-10-16
- Publication Date
- 2026-01-16
AI Technical Summary
Existing visual prostheses, due to their use of large-channel electrode arrays in full-image stimulation mode, result in redundant information transmission and electrode channel interference, which reduces the brain's perception quality of the target object.
By identifying targets and extracting regions of interest from grayscale images, the stimulation current values are grouped and sorted. The stimulation current matrix is then mapped to control the release of stimulation current from the electrode array, prioritizing the release of high-brightness regions, reducing the number of electrode channels participating in the release at the same time, and minimizing interference.
It improves the brain's perception of the target object, ensures the reception of complete image stimuli, reduces interference between electrodes, and enhances the perception effect.
Smart Images

Figure CN121338239A_ABST
Abstract
Description
Technical Field
[0001] This application relates to the field of medical device technology, and in particular to a method, system and storage medium for controlling a visual prosthesis. Background Technology
[0002] Visual prostheses are neural interface technologies that transmit image information to the visual system through artificial electrical stimulation. They convert the recognized image signal into electrical pulse signals, which are then sent to an electrode array deployed in the image recognition region of the user's brain. The electrode array then releases current based on these pulse signals to stimulate the brain's image recognition region, allowing the user to perceive the image. To improve spatial resolution, current visual prostheses typically employ electrode arrays with a large number of channels (thousands or even tens of thousands) to increase the stimulation density of the brain's image recognition region, thereby finely depicting image contours and details.
[0003] However, existing visual prostheses utilize large-channel electrode arrays that operate in a full-image stimulation mode, leading to a decline in perceptual quality. Specifically, full-image stimulation involves converting all information from the acquired image into electrical pulse signals to control the electrode array to release stimulating currents. However, images contain a large amount of redundant information, including background, in addition to target object information. This results in redundant information and target object information being transmitted to the brain simultaneously during full-image stimulation, leading to a decrease in the brain's perceptual quality of the target object. Furthermore, because full-image stimulation triggers a large number of electrode channels simultaneously, when the electrode channels outputting target object information and those outputting redundant information are adjacent, the output stimulation signals interfere with each other, further reducing the brain's perceptual quality of the target object. Summary of the Invention
[0004] In view of the above problems, this application provides a method, system, and storage medium for controlling a visual prosthesis, in order to improve the quality of the brain's perception of a target object. The specific solution is as follows:
[0005] The first aspect of this application provides a method for controlling a visual prosthesis, comprising:
[0006] The target object recognition results based on the obtained grayscale image are broadcast aloud, and the target object image is cropped from the grayscale image based on the target object selection results fed back by the user, to obtain at least one initial target image;
[0007] The region of interest is extracted from the initial target image to obtain a target object contour image, which includes multiple contour pixels.
[0008] Based on the brightness value of each contour pixel, the contour pixels are grouped to obtain multiple data groups. The stimulation order of each data group is set according to the average brightness value of the data groups from high to low. Each data group includes the spatial coordinates of multiple contour pixels. Based on the stimulation current value corresponding to the smallest non-zero gray value among the gray values of each contour pixel, and the ratio between the gray value of each contour pixel and the smallest non-zero gray value, the stimulation current value of each contour pixel is obtained.
[0009] Based on the spatial coordinates and the stimulation current value, a stimulation current matrix corresponding to each data group is mapped, and based on the stimulation sequence of each data group, stimulation current is released in the electrode array in sequence based on each stimulation current matrix. The stimulation current matrix includes electrode channel identifiers corresponding to the spatial coordinates of multiple contour pixels, and stimulation current values corresponding to the contour pixels.
[0010] In one possible implementation, prior to the release of stimulation currents in the electrode arrays sequentially based on the stimulation order of each of the data sets, the method further includes:
[0011] For each of the described stimulation current matrices:
[0012] Obtain the total current value of each of the stimulation current values in the stimulation current matrix;
[0013] If the total current value is greater than a preset safety threshold, based on the shearing coefficient corresponding to the preset safety threshold, the stimulation current value with the largest value in the stimulation current matrix that has not been sheared is sheared to obtain an updated stimulation current matrix, and the operation step of obtaining the total current value of each stimulation current value in the stimulation current matrix is performed.
[0014] If the total current value is not greater than the preset safety threshold, the operation steps of releasing stimulation current in the control electrode array based on the stimulation sequence of each data group and the stimulation current matrix are executed.
[0015] In one possible implementation, grouping the contour pixels based on their brightness values to obtain multiple data groups includes:
[0016] The brightness distribution is determined based on the brightness values of each of the contour pixels;
[0017] Based on the brightness distribution, each contour pixel is grouped to obtain multiple data groups. The minimum brightness of the contour pixel in the first data group is higher than the maximum brightness of the contour pixel in the second data group. The first data group is timed before the second data group. The first data group and the second data group are data groups among multiple data groups.
[0018] In one possible implementation, the step of extracting the region of interest from the initial target image to obtain the target object contour image includes:
[0019] The initial target image is downsampled and Gaussian filtered to obtain multiple secondary target images of different resolutions.
[0020] For any two of the secondary target images: unify the size of the two secondary target images, and subtract the pixel values of two corresponding pixels in the two secondary target images to obtain a feature map;
[0021] The obtained feature maps are weighted and summed to obtain the region of interest image;
[0022] The contour image of the target object is obtained by using a preset boundary extraction algorithm to extract the contour of the region of interest image.
[0023] In one possible implementation, before performing contour extraction on the region of interest image using a preset boundary extraction algorithm, the method further includes:
[0024] The region of interest image is downsampled to obtain a region of interest image with the same size as the electrode array.
[0025] In one possible implementation, the voice broadcasting based on the target object recognition result of the obtained grayscale image includes:
[0026] Obtain the original image and preprocess the original image to obtain the grayscale image;
[0027] The grayscale image is input into a preset target recognition model to obtain the type, confidence level, and bounding box of each target object in the grayscale image;
[0028] The voice broadcast is performed sequentially according to the order of confidence level from high to low for each of the target objects.
[0029] In one possible implementation, the step of cropping the grayscale image based on the target selection result from user feedback to obtain at least one initial target image includes:
[0030] The target selection results are analyzed to obtain the identifier of at least one target.
[0031] A preset image cropping algorithm is invoked to crop the image of the grayscale image that is located within the coordinate range of the bounding box corresponding to the identifier of the target object, thereby obtaining at least one initial target image.
[0032] A second aspect of this application provides a control system for a visual prosthesis, comprising:
[0033] The first image extraction module is used to broadcast the target object recognition results based on the obtained grayscale image via voice, and to crop the target object image from the grayscale image based on the target object selection results fed back by the user, so as to obtain at least one initial target image.
[0034] The second image extraction module is used to extract the region of interest from the initial target image to obtain a target object contour image, wherein the target object contour image includes multiple contour pixels.
[0035] The parameter acquisition module is used to group the contour pixels based on their brightness values to obtain multiple data groups, and to set the stimulation order of each data group based on the average brightness values of the data groups from high to low. The data groups include the spatial coordinates of multiple contour pixels. The module also obtains the stimulation current value of each contour pixel based on the stimulation current value corresponding to the smallest non-zero gray value among the gray values of each contour pixel, and the ratio between the gray values of each contour pixel and the smallest non-zero gray value.
[0036] An array control module is used to map a stimulation current matrix corresponding to each data group based on the spatial coordinates and the stimulation current value, and to control the release of stimulation current in the electrode array according to the stimulation sequence of each data group based on the stimulation current matrix of each stimulation current matrix. The stimulation current matrix includes electrode channel identifiers corresponding to the spatial coordinates of multiple contour pixels and stimulation current values corresponding to the contour pixels.
[0037] In one possible implementation, the array control module is further configured as follows:
[0038] Before releasing stimulation currents in the electrode arrays sequentially based on the stimulation order of each of the data sets, according to each stimulation current matrix:
[0039] Obtain the total current value of each of the stimulation current values in the stimulation current matrix;
[0040] If the total current value is greater than a preset safety threshold, based on the shearing coefficient corresponding to the preset safety threshold, the stimulation current value with the largest value in the stimulation current matrix that has not been sheared is sheared to obtain an updated stimulation current matrix, and the operation step of obtaining the total current value of each stimulation current value in the stimulation current matrix is performed.
[0041] If the total current value is not greater than the preset safety threshold, the operation steps of releasing stimulation current in the control electrode array based on the stimulation sequence of each data group and the stimulation current matrix are executed.
[0042] In one possible implementation, the parameter acquisition module is configured to group the contour pixels based on their brightness values to obtain multiple data groups as follows:
[0043] The brightness distribution is determined based on the brightness values of each of the contour pixels;
[0044] Based on the brightness distribution, each contour pixel is grouped to obtain multiple data groups. The minimum brightness of the contour pixel in the first data group is higher than the maximum brightness of the contour pixel in the second data group. The first data group is timed before the second data group. The first data group and the second data group are data groups among multiple data groups.
[0045] In one possible implementation, the second image extraction module is configured as follows:
[0046] The initial target image is downsampled and Gaussian filtered to obtain multiple secondary target images of different resolutions.
[0047] For any two of the secondary target images: unify the size of the two secondary target images, and subtract the pixel values of two corresponding pixels in the two secondary target images to obtain a feature map;
[0048] The obtained feature maps are weighted and summed to obtain the region of interest image;
[0049] The contour image of the target object is obtained by using a preset boundary extraction algorithm to extract the contour of the region of interest image.
[0050] In one possible implementation, the second image extraction module is further configured as follows:
[0051] Before extracting the contour of the region of interest image using a preset boundary extraction algorithm, the region of interest image is downsampled to obtain a region of interest image with the same size as the electrode array.
[0052] In one possible implementation, the first image extraction module is configured to perform voice broadcasting based on the target object recognition result of the obtained grayscale image:
[0053] Obtain the original image and preprocess the original image to obtain the grayscale image;
[0054] The grayscale image is input into a preset target recognition model to obtain the type, confidence level, and bounding box of each target object in the grayscale image;
[0055] The voice broadcast is performed sequentially according to the order of confidence level from high to low for each of the target objects.
[0056] In one possible implementation, the first image extraction module is configured to perform target image cropping on the grayscale image based on the target selection result provided by user feedback to obtain at least one initial target image:
[0057] The target selection results are analyzed to obtain the identifier of at least one target.
[0058] A preset image cropping algorithm is invoked to crop the image of the grayscale image that is located within the coordinate range of the bounding box corresponding to the identifier of the target object, thereby obtaining at least one initial target image.
[0059] A third aspect of this application provides a computer storage medium carrying one or more computer programs that, when executed by an electronic device, enable the electronic device to control a visual prosthesis according to the first aspect or any implementation thereof.
[0060] By employing the above technical solutions, this application provides a control method, system, and storage medium for a visual prosthesis. This involves configuring voice broadcasting based on target object recognition results from an acquired grayscale image, and cropping the grayscale image to obtain at least one initial target image based on user feedback regarding target object selection. This process effectively removes redundant data from the grayscale image that is of no interest to the user. Subsequently, by configuring region of interest extraction on the initial target image, a target object contour image is obtained, further reducing the amount of redundant data. Furthermore, by configuring the brightness values of each contour pixel to group the contour pixels, multiple data groups containing the spatial coordinates of multiple contour pixels are obtained. The stimulation current value corresponding to the smallest non-zero gray value among the gray values of each contour pixel, as well as the ratio between the gray values of each contour pixel and the smallest non-zero gray value, are configured to obtain the stimulation current value of each contour pixel. Based on the spatial coordinates and stimulation current values, the stimulation current matrix corresponding to each data group is mapped. Thus, in the subsequent control of the electrode array, the stimulation current in the electrode array is released sequentially based on the stimulation order of each data group and the stimulation current matrix. Since the stimulation current matrix includes the electrode channel identifiers corresponding to the spatial coordinates of multiple contour pixels and the stimulation current values corresponding to the contour pixels, rather than the data of all contour pixels that constitute the contour of the target object, the number of electrode channels participating in the release of stimulation current at the same time is reduced, the interference between electrodes is reduced, and the perception quality of the brain of the target object is improved. Finally, because the stimulation is controlled sequentially in the electrode array based on the stimulation sequence of each data group, and the stimulation sequence is set from high to low based on the average brightness value of the data groups, the brain preferentially receives the secondary stimulation with higher overall brightness. Higher brightness results in a longer duration of stimulation. Therefore, this application, by configuring the stimulation sequence based on each data group and controlling the release of stimulation current in the electrode array sequentially based on the stimulation current matrix, ensures that the brain receives complete image stimulation, thus ensuring perceptual quality. Attached Figure Description
[0061] The above and other features, advantages, and aspects of the embodiments of this disclosure will become more apparent from the accompanying drawings and the following detailed description. Throughout the drawings, the same or similar reference numerals denote the same or similar elements. It should be understood that the drawings are schematic, and the originals and elements are not necessarily drawn to scale.
[0062] Figure 1 A flowchart of a method for controlling a visual prosthesis provided in this application;
[0063] Figure 2 This application provides a schematic diagram of the structure of a visual prosthesis system;
[0064] Figure 3This application provides a schematic diagram of a cutting operation process;
[0065] Figure 4 A flowchart of another method for controlling a visual prosthesis provided in this application;
[0066] Figure 5 A block diagram of a control system for a visual prosthesis provided in this application;
[0067] Figure 6 This is a schematic diagram of the structure of an electronic device provided in this application. Detailed Implementation
[0068] The embodiments of this application are described below with reference to the accompanying drawings. The terminology used in the implementation section of this application is for explaining specific embodiments only and is not intended to limit the scope of this application.
[0069] The embodiments of this application will now be described with reference to the accompanying drawings. Those skilled in the art will recognize that, with technological advancements and the emergence of new scenarios, the technical solutions provided in the embodiments of this application are equally applicable to similar technical problems.
[0070] The terms "first," "second," etc., used in the specification, claims, and accompanying drawings of this application are used to distinguish similar objects and are not necessarily used to describe a specific order or sequence. It should be understood that such terms are interchangeable where appropriate; this is merely a way of distinguishing objects with the same attributes in the embodiments of this application. Furthermore, the terms "comprising" and "having," and any variations thereof, are intended to cover non-exclusive inclusion, so that a process, method, system, product, or apparatus that comprises a series of elements is not necessarily limited to those elements, but may include other elements not explicitly listed or inherent to those processes, methods, products, or apparatuses.
[0071] The first aspect of this application provides a method for controlling a visual prosthesis, such as... Figure 1 As shown, it includes:
[0072] S101. The target object recognition result based on the obtained grayscale image is broadcast by voice, and the target object image is cropped from the grayscale image based on the target object selection result fed back by the user, so as to obtain at least one initial target image.
[0073] It should be noted that, in practical application scenarios, the visual prosthesis control method provided in the first aspect of this application can be configured as follows: Figure 2 The visual prosthetic system shown. Specifically, such as... Figure 2This is a schematic diagram of a visual prosthesis system. The system includes a camera 1, a control terminal 2, a transmitter 3, a receiving coil 4, a controller 5, and an electrode array 6. The camera 1 is used to acquire images, the control terminal 2 is used to execute the control method for the visual prosthesis provided in the first aspect of this application, and the transmitter 3 is used to transmit the stimulation current matrix sent by the control terminal 2 to the controller 5 via the receiving coil 4. The controller 5 is used to control the electrode array 6 to release secondary current according to the stimulation current matrix.
[0074] It should be noted that, in practical applications, the aforementioned grayscale image can be an image obtained by converting a color image captured by an image acquisition device (such as an external camera) to grayscale. This application improves image contrast and thus enhances the accuracy of the obtained target recognition results by configuring the color image captured by the image acquisition device to be converted into a grayscale image.
[0075] Those skilled in the art will understand that, in practical applications, to further improve the quality of grayscale images, further processing including image denoising, contrast enhancement, and image normalization can be applied. Image denoising can be achieved through mean filtering, contrast enhancement through histogram equalization, and image normalization through min-max normalization. This application does not impose excessive limitations or elaborate on the specific types and implementation methods of the above processing methods.
[0076] It should be noted that in practical applications, grayscale images contain both target objects and distractions. For example, in obstacle avoidance scenarios, the grayscale image of the road in front of the user may include safety posts, background buildings, pedestrian crossings, and roadside shop billboards and their content. For the user, safety posts and pedestrian crossings are target objects, while background buildings, roadside shop billboards and their content are distractions. If these are not distinguished and processed in their entirety, not only will redundant information be introduced, but the single-recognition response rate will also decrease. Therefore, this application configures the target object recognition results based on the obtained grayscale image to be broadcast aloud, and based on the target object selection results fed back by the user, the grayscale image is cropped to obtain at least one initial target image. This reduces the amount of redundant information and the response rate in subsequent processing, reduces the number of electrode channels in the electrode array occupied by redundant information, and thus reduces the risk of decreased perception quality caused by interference due to a large number of electrode channels being called at the same time.
[0077] S102. Extract the region of interest from the initial target image to obtain the target object contour image, which includes multiple contour pixels.
[0078] It should be noted that in practical applications, the aforementioned Region of Interest (ROI) is a specific area in the image that the user is interested in or that requires further processing. In the application scenario of visual prosthetics, the aforementioned ROI is the outline of the target object (object or content) in the initial target image, while information such as the color of the target object and background objects is redundant. Therefore, this application extracts the ROI from the initial target image by configuring it to obtain the target object outline information including multiple outline pixels of the target object, thereby further reducing redundant information, reducing the number of electrode channels occupied by redundant information in the electrode array, and thus reducing the risk of decreased perception quality caused by interference due to a large number of electrode channels being called at the same time.
[0079] S103. Group the contour pixels based on their brightness values to obtain multiple data groups, and set the stimulation order of each data group according to the average brightness value of the data group from high to low. The data group includes the spatial coordinates of multiple contour pixels. Based on the stimulation current value corresponding to the smallest non-zero gray value in the gray values of each contour pixel, and the ratio between the gray values of each contour pixel and the smallest non-zero gray value, obtain the stimulation current value of each contour pixel.
[0080] It should be noted that, in practical applications, the researchers of this application have discovered that the human visual system prioritizes high-contrast and high-brightness areas. High-brightness pixels in the outline image of a target object typically correspond to the object's edge contour. Therefore, this application groups contour pixels based on their brightness values to obtain multiple data groups. The stimulation sequence of each data group is then set according to its average brightness value from highest to lowest. This ensures that when a stimulation current is subsequently released, the stimulation current corresponding to the high-brightness contour pixels is released first, allowing the brain to quickly perceive the approximate outline of the target object. Furthermore, due to the perceptual stagnation of the visual system, by sequentially releasing stimulation currents to the contour pixels in each data group with gradually decreasing average brightness values, the brain can still perceive the outline and details of the target object, thus improving the quality of perception.
[0081] It should be noted that, in practical application scenarios, there are multiple ways to obtain the stimulation current value of each contour pixel in step S102 based on the stimulation current value corresponding to the smallest non-zero gray value among the gray values of each contour pixel, and the ratio between the gray value of each contour pixel and the smallest non-zero gray value. Here, one is provided as an example, including the following steps A1 to A4.
[0082] Step A1: Traverse the grayscale values of each contour pixel in the target object contour image, extract the smallest non-zero grayscale value, and mark the contour pixel corresponding to the smallest non-zero grayscale value as the reference pixel. Then trigger step A2.
[0083] Step A2: Based on the preset mapping table, find the stimulation current value corresponding to the smallest non-zero grayscale value obtained in Step A1, and determine it as the reference current value. Then, trigger Step A3.
[0084] In one possible implementation, the preset mapping table in step A2 above stores each grayscale value and its corresponding stimulation current value.
[0085] Step A3: For each contour pixel, calculate the ratio of its grayscale value to that of the reference pixel. Then trigger step A4.
[0086] Step A4: For each contour pixel: the product of the proportional relationship of the contour pixel and the reference current value in step A2 is determined as the stimulation current value of the contour pixel.
[0087] S104. Based on spatial coordinates and stimulation current values, map stimulation current matrices corresponding to each data group, and based on the stimulation sequence of each data group, control the release of stimulation currents in the electrode array according to each stimulation current matrix. The stimulation current matrix includes electrode channel identifiers corresponding to the spatial coordinates of multiple contour pixels, and stimulation current values corresponding to the contour pixels.
[0088] It should be noted that in practical application scenarios, there are multiple implementation methods for mapping the stimulation current matrix corresponding to each data group based on spatial coordinates and stimulation current values. Here, one is provided as an example, including the following steps B1 to B2.
[0089] Step B1: For each data group, an initial stimulation current matrix is mapped based on the spatial coordinates of multiple contour pixels in the data group. The position of each element in the initial stimulation current matrix represents the electrode channel in the electrode matrix. Then, step B2 is triggered.
[0090] Step B2: For each initial stimulation current matrix: Based on the spatial coordinates represented by each element in the initial stimulation current matrix, determine the corresponding contour pixel point of each element, and update the element value to the stimulation current value corresponding to the contour pixel point.
[0091] It should be noted that in practical applications, since the aforementioned stimulation current matrix includes only a portion of the contour pixels that constitute the target object's outline, rather than all of them, this application reduces the number of electrode channels simultaneously releasing stimulation current by configuring the stimulation sequence based on each data group and controlling the release of stimulation current in the electrode array sequentially based on each stimulation current matrix. This reduces the risk of interference between electrodes and improves the brain's perception quality of the target object. Furthermore, because the stimulation sequence is based on the stimulation sequence of each data group, and the stimulation sequence is set according to the average brightness value of the data group from high to low, it ensures that the brain can receive complete image stimulation, thus ensuring perception quality.
[0092] It should be noted that in practical applications, during the process of releasing stimulation current in the electrode array based on each stimulation current matrix, the amplitude of the pulse current output by the electrode channel can be determined based on the gray value of the contour pixel, and the frequency of the pulse current output by the motor channel can be determined based on the brightness value of the contour pixel.
[0093] This application configures voice broadcasting based on the target object recognition results of the obtained grayscale image, and performs target object image cropping on the grayscale image based on the target object selection results fed back by the user, thereby obtaining at least one initial target image, thus eliminating redundant data in the grayscale image that is not of interest to the user. Subsequently, by configuring the extraction of the region of interest from the initial target image, the target object contour image is obtained, thereby further reducing the amount of redundant data. Furthermore, by configuring the brightness values of each contour pixel to group the contour pixels, multiple data groups containing the spatial coordinates of multiple contour pixels are obtained. The stimulation current value corresponding to the smallest non-zero gray value among the gray values of each contour pixel, as well as the ratio between the gray values of each contour pixel and the smallest non-zero gray value, are configured to obtain the stimulation current value of each contour pixel. Based on the spatial coordinates and stimulation current values, the stimulation current matrix corresponding to each data group is mapped. Thus, in the subsequent control of the electrode array, the stimulation current in the electrode array is released sequentially based on the stimulation order of each data group and the stimulation current matrix. Since the stimulation current matrix includes the electrode channel identifiers corresponding to the spatial coordinates of multiple contour pixels and the stimulation current values corresponding to the contour pixels, rather than the data of all contour pixels that constitute the contour of the target object, the number of electrode channels participating in the release of stimulation current at the same time is reduced, the interference between electrodes is reduced, and the perception quality of the brain of the target object is improved. Finally, because the stimulation is controlled sequentially in the electrode array based on the stimulation sequence of each data group, and the stimulation sequence is set from high to low based on the average brightness value of the data groups, the brain preferentially receives the secondary stimulation with higher overall brightness. Higher brightness results in a longer duration of stimulation. Therefore, this application, by configuring the stimulation sequence based on each data group and controlling the release of stimulation current in the electrode array sequentially based on the stimulation current matrix, ensures that the brain receives complete image stimulation, thus ensuring perceptual quality.
[0094] In one possible implementation, before releasing stimulation currents sequentially from the electrode array based on the stimulation sequence of each data set and according to each stimulation current matrix, the following is also included:
[0095] For each stimulus current matrix:
[0096] Obtain the total current value of each stimulus current value in the stimulus current matrix;
[0097] If the total current value is greater than the preset safety threshold, based on the shearing coefficient corresponding to the preset safety threshold, the stimulus current value with the largest value in the stimulus current matrix that has not been sheared is sheared to obtain the updated stimulus current matrix, and the operation step of obtaining the total current value of each stimulus current value in the stimulus current matrix is executed.
[0098] If the total current value is not greater than the preset safety threshold, the operation steps of releasing the stimulation current in the control electrode array are performed according to the stimulation sequence of each data group and the stimulation current matrix of each stimulation current matrix.
[0099] It should be noted that in practical applications, excessively high stimulation currents can cause user discomfort and affect the stimulation effect of subsequent stimulation current matrices, reducing the brain's perception of the image quality. Therefore, this application configures the total current value of each stimulation current value in the stimulation current matrix, and configures the method to cut off the largest uncut stimulation current value in the stimulation current matrix based on a shearing coefficient corresponding to the preset safety threshold when the total current value exceeds a preset safety threshold. This ensures that the total current value of the stimulation current matrix used for electrode matrix control does not exceed the preset safety threshold, thereby avoiding user discomfort and perception quality issues.
[0100] It should be noted that, in practical applications, the above flowchart of the cutting operation can be illustrated as follows: Figure 3 As shown, it includes the following steps S301 to S305.
[0101] Step S301: According to the order of stimulation, the stimulation current matrix with the highest priority and without added detection tags is determined as the current stimulation current matrix. Step S302 is then triggered.
[0102] Step S302: Calculate the total current value of the current stimulation current matrix. Then trigger step S303.
[0103] Step S303: Determine whether the total current value of the current stimulation current matrix is greater than a preset safety threshold. If yes, trigger step S304; otherwise, trigger step S305.
[0104] Step S304: Based on the shearing coefficient corresponding to the preset safety threshold, the stimulation current value with the largest value in the stimulation current matrix that has not been sheared is sheared to obtain an updated stimulation current matrix. Step S302 is then triggered.
[0105] Step S305: Add a detected tag to the current stimulation current matrix. And trigger step S301.
[0106] In one possible implementation, after step S305 is completed, an operation step can be triggered to release stimulation current in the electrode array based on the stimulation sequence of each data group and the stimulation current matrix of each stimulation current matrix. Each stimulation current matrix is a stimulation current matrix with a detected tag added.
[0107] It should be noted that, in practical applications, the aforementioned preset safety threshold can be the maximum value of the stimulation current that will not cause discomfort to the user, determined based on calibration tests.
[0108] In one possible implementation, the contour pixels are grouped based on their brightness values to obtain multiple data groups, including:
[0109] The brightness distribution is determined based on the brightness values of each contour pixel.
[0110] Based on the brightness distribution, each contour pixel is grouped to obtain multiple data groups. The minimum brightness of the contour pixels in the first data group is higher than the maximum brightness of the contour pixels in the second data group. The first data group is time-series prior to the second data group. The first and second data groups are data groups among multiple data groups.
[0111] It should be noted that in practical application scenarios, there are multiple ways to determine the brightness distribution based on the brightness value of each contour pixel, and to group each contour pixel based on the brightness distribution to obtain multiple data groups. Here, we provide an example, which includes the following steps C1 to C3.
[0112] Step C1: Traverse and extract the spatial coordinates and brightness values of each contour pixel in the target object contour image. Then trigger C2.
[0113] Step C2: Construct a distribution histogram based on the brightness values. This triggers step C3.
[0114] Step C3: Extract the spatial coordinates of multiple contour pixels in a region of the distribution histogram into a data set.
[0115] In one possible implementation, the region of interest is extracted from the initial target image to obtain the target object contour image, including:
[0116] The initial target image is downsampled and Gaussian filtered to obtain multiple secondary target images of different resolutions.
[0117] For any two secondary target images in each secondary target image: unify the size of the two secondary target images, and subtract the pixel values of corresponding two pixels in the two secondary target images to obtain the feature map;
[0118] The obtained feature maps are weighted and summed to obtain the region of interest image.
[0119] A pre-defined boundary extraction algorithm is used to extract the contour of the region of interest image to obtain the contour image of the target object.
[0120] It should be noted that in practical applications, the aforementioned downsampling is an image processing method that reduces image resolution by decreasing the number of pixels. Gaussian filtering is an image processing method used to remove image noise while preserving the main features of the image. Since the initial target image contains a lot of redundant information (such as the background and ambient light of the target object), and the resolution of this redundant information is usually lower than the resolution of the target object's contour, this application obtains multiple secondary target images of different resolutions by performing downsampling and Gaussian filtering on the initial target image. This achieves the removal of redundant information and image noise, improving the representation accuracy of contour pixels in the secondary target images.
[0121] It should be noted that in practical applications, the representation of the target object and background differs across secondary target images of different resolutions. High-resolution images emphasize detailed features, while low-resolution images emphasize overall structural features. Therefore, this application configures any two secondary target images in each secondary target image to have their dimensions unified, and calculates the difference between the pixel values of corresponding two pixels in the two secondary target images to obtain a feature map. This feature map is then weighted and summed to increase the contrast between the target object and the background, thereby improving the feature display effect of the contour pixels in the obtained region of interest image. Furthermore, a preset boundary extraction algorithm is used to extract the contour of the region of interest image, ensuring that the obtained target object contour image includes each contour pixel that clearly represents the target object's contour, thus improving the generation quality of the target object contour image.
[0122] Those skilled in the art will understand that, in practical applications, there are various types of preset boundary extraction algorithms, including but not limited to: Iterative Boundary Extraction (IBE) algorithm, Otsu's Method (Otsu), and Geodesic Active Contour (GAC) algorithm. This application does not impose excessive limitations or elaborate on the types and usage processes of the aforementioned preset boundary extraction algorithms.
[0123] In one possible implementation, before extracting the contour of the region of interest image using a predefined boundary extraction algorithm, the following is also included:
[0124] The region of interest image is downsampled to obtain a region of interest image with the same size as the electrode array.
[0125] It should be noted that in practical application scenarios, this application configures the downsampling operation of the region of interest image to obtain a region of interest image with the same size as the electrode array, thereby making the final stimulation current matrix compatible with the electrode array.
[0126] In one possible implementation, voice broadcasting is performed based on the target object recognition results of the obtained grayscale image, including:
[0127] Obtain the original image and preprocess it to obtain a grayscale image;
[0128] Input the grayscale image into the preset target recognition model to obtain the type, confidence level and bounding box of each target object in the grayscale image;
[0129] The types of each target object are announced in order of confidence level from high to low.
[0130] It should be noted that, in practical applications, the aforementioned preset target recognition model can be a model obtained by training an existing target detection model. The construction process of the aforementioned preset target recognition model may include the following steps D1 to D3.
[0131] Step D1: Obtain multiple training grayscale images and add labels to them. This triggers step D2.
[0132] In one possible implementation, the content of the above label is the target object in the training grayscale image.
[0133] Step D2: Train the initial target recognition model using the labeled training grayscale image added in step D1 to obtain a preset target recognition model. The input of the preset target recognition model is the grayscale image, and the output is the type, confidence level, and bounding box of each target object in the grayscale image.
[0134] In one possible implementation, since the control method for the visual prosthesis provided by the first aspect of this application and any implementation thereof is applied to a visual prosthesis system, and the visual prosthesis system requires human wearing, it has high requirements for weight and volume. Therefore, to avoid the problem of the control terminal becoming too bulky due to meeting the computational power requirements of the model, the initial target recognition model in step D2 above can adopt a lightweight target detection model, such as the YOLOv11 algorithm.
[0135] In one possible implementation, based on the target object selection result provided by user feedback, the grayscale image is cropped to obtain at least one initial target image, including:
[0136] The target selection results are analyzed to obtain the identifier of at least one target.
[0137] The preset image cropping algorithm is invoked to crop the image of the grayscale image within the coordinate range of the bounding box corresponding to the target object's identifier, thereby obtaining at least one initial target image.
[0138] To facilitate understanding of the visual prosthesis control method provided by the first aspect and any implementation thereof of this application, an example of a possible implementation of this application is described below:
[0139] like Figure 4 The diagram shows a flowchart of a method for controlling a visual prosthesis. The specific operation steps are as follows:
[0140] Step S401: Obtain the original image and preprocess it to obtain a grayscale image. Then, step S402 is triggered.
[0141] Step S402: Input the grayscale image into the preset target recognition model to obtain the type, confidence level, and bounding box of each target object in the grayscale image. Then trigger step S403.
[0142] Step S403: The types of each target object are announced via voice in descending order of confidence level. This triggers step S404.
[0143] Step S404: Based on the target selection result provided by the user, the grayscale image is cropped to obtain at least one initial target image. Step S405 is then triggered.
[0144] Step S405: Extract the region of interest from each initial target image to obtain the target object contour image. This triggers steps S406 and S407.
[0145] Step S406: For each target object contour image: group the contour pixels based on the brightness distribution of each contour pixel in the target object contour image to obtain multiple data groups. Then trigger step S408.
[0146] Step S407: Based on the stimulation current value corresponding to the smallest non-zero gray value among the gray values of each contour pixel, and the ratio between the gray values of each contour pixel and the smallest non-zero gray value, obtain the stimulation current value of each contour pixel. Then trigger step S408.
[0147] Step S408: Based on the spatial coordinates and stimulation current values, map the stimulation current matrix corresponding to each data group. Then, step S409 is triggered.
[0148] Step S409: Based on the stimulation sequence of each data group, release stimulation currents sequentially in the control electrode array based on each stimulation current matrix.
[0149] It should be noted that, in practical application scenarios, steps S401 to S404 are as described above. Figure 1 One possible implementation of step S101 shown above. Step S405 is as described above. Figure 1 One possible implementation of step S102 shown. Steps S406 and S407 described above are as follows: Figure 1 One possible implementation of step S103 is shown. Steps S408 and S409 described above are as follows: Figure 1 One possible implementation of step S104 shown.
[0150] A second aspect of this application provides a control system for a visual prosthesis, such as... Figure 5 As shown, the control system of this visual prosthesis includes:
[0151] The first image extraction module 501 is used to perform voice broadcast on the target object recognition result based on the obtained grayscale image, and to perform target object image cropping on the grayscale image based on the target object selection result fed back by the user, so as to obtain at least one initial target image.
[0152] The second image extraction module 502 is used to extract the region of interest from the initial target image to obtain the target object contour image, which includes multiple contour pixels.
[0153] The parameter acquisition module 503 is used to group the contour pixels based on the brightness value of each contour pixel to obtain multiple data groups, and to set the stimulation order of each data group based on the average brightness value of the data group from high to low. The data group includes the spatial coordinates of multiple contour pixels. Based on the stimulation current value corresponding to the smallest non-zero gray value among the gray values of each contour pixel, and the ratio between the gray value of each contour pixel and the smallest non-zero gray value, the stimulation current value of each contour pixel is obtained.
[0154] The array control module 504 is used to map the stimulation current matrix corresponding to each data group based on spatial coordinates and stimulation current values, and to control the release of stimulation current in the electrode array according to the stimulation sequence of each data group based on each stimulation current matrix. The stimulation current matrix includes the electrode channel identifier corresponding to the spatial coordinates of multiple contour pixels, and the stimulation current value corresponding to the contour pixels.
[0155] In one possible implementation, the array control module 504 described above is further configured as follows:
[0156] Before releasing stimulation currents in the electrode arrays based on the stimulation sequence of each data set and according to each stimulation current matrix, the stimulation current matrices are:
[0157] Obtain the total current value of each stimulus current value in the stimulus current matrix;
[0158] If the total current value is greater than the preset safety threshold, based on the shearing coefficient corresponding to the preset safety threshold, the stimulus current value with the largest value in the stimulus current matrix that has not been sheared is sheared to obtain the updated stimulus current matrix, and the operation step of obtaining the total current value of each stimulus current value in the stimulus current matrix is executed.
[0159] If the total current value is not greater than the preset safety threshold, the operation steps of releasing the stimulation current in the control electrode array are performed according to the stimulation sequence of each data group and the stimulation current matrix of each stimulation current matrix.
[0160] In one possible implementation, the parameter acquisition module 503 is configured to group the contour pixels based on their brightness values to obtain multiple data groups as follows:
[0161] The brightness distribution is determined based on the brightness values of each contour pixel.
[0162] Based on the brightness distribution, each contour pixel is grouped to obtain multiple data groups. The minimum brightness of the contour pixels in the first data group is higher than the maximum brightness of the contour pixels in the second data group. The first data group is time-series prior to the second data group. The first and second data groups are data groups among multiple data groups.
[0163] In one possible implementation, the second image extraction module 502 described above is configured as follows:
[0164] The initial target image is downsampled and Gaussian filtered to obtain multiple secondary target images of different resolutions.
[0165] For any two secondary target images in each secondary target image: unify the size of the two secondary target images, and subtract the pixel values of corresponding two pixels in the two secondary target images to obtain the feature map;
[0166] The obtained feature maps are weighted and summed to obtain the region of interest image.
[0167] A pre-defined boundary extraction algorithm is used to extract the contour of the region of interest image to obtain the contour image of the target object.
[0168] In one possible implementation, the second image extraction module 502 described above is further configured as follows:
[0169] Before using a preset boundary extraction algorithm to extract the contour of the region of interest image, the region of interest image is downsampled to obtain a region of interest image with the same size as the electrode array.
[0170] In one possible implementation, the first image extraction module 501 described above is configured to perform voice broadcasting based on the target object recognition result of the obtained grayscale image:
[0171] Obtain the original image and preprocess it.
[0172] The preprocessed original image is input into a preset target recognition model to obtain the type, confidence level and bounding box of each target object in the preprocessed original image;
[0173] The types of each target object are announced in order of confidence level from high to low.
[0174] In one possible implementation, the first image extraction module 501 described above is configured to perform target image cropping on the grayscale image based on the target selection result provided by the user feedback to obtain at least one initial target image:
[0175] The target selection results are analyzed to obtain the identifier of at least one target.
[0176] The preset image cropping algorithm is invoked to crop the image of the grayscale image within the coordinate range of the bounding box corresponding to the target object's identifier, thereby obtaining at least one initial target image.
[0177] A third aspect of this application provides a computer storage medium carrying one or more computer programs, which, when executed by an electronic device, enable the electronic device to control a visual prosthesis according to the first aspect or any implementation thereof.
[0178] This application also provides an electronic device in its embodiments. (See reference...) Figure 6 The diagram illustrates a structural schematic suitable for implementing the electronic device in the embodiments of this application. The electronic device in the embodiments of this application may include, but is not limited to, fixed terminals such as mobile phones, laptops, PDAs (personal digital assistants), PADs (tablet computers), desktop computers, etc. Figure 6 The electronic device shown is merely an example and should not impose any limitation on the functionality and scope of use of the embodiments of this application.
[0179] like Figure 6As shown, the electronic device may include a processing unit (e.g., a central processing unit, a graphics processing unit, etc.) 601, which can perform various appropriate actions and processes according to a program stored in a read-only memory (ROM) 602 or a program loaded from a storage device 608 into a random access memory (RAM) 603. When the electronic device is powered on, the RAM 603 also stores various programs and data required for the operation of the electronic device. The processing unit 601, ROM 602, and RAM 603 are interconnected via a bus 604. An input / output (I / O) interface 605 is also connected to the bus 604.
[0180] Typically, the following devices can be connected to I / O interface 605: input devices 606 including, for example, touchscreens, touchpads, keyboards, mice, cameras, microphones, accelerometers, gyroscopes, etc.; output devices 607 including, for example, liquid crystal displays (LCDs), speakers, vibrators, etc.; storage devices 608 including, for example, memory cards, hard drives, etc.; and communication devices 609. Communication device 609 allows electronic devices to communicate wirelessly or wiredly with other devices to exchange data. Although Figure 6 Electronic devices with various devices are shown, but it should be understood that it is not required to implement or have all of the devices shown. More or fewer devices may be implemented or have instead.
[0181] This application also provides a computer program product including computer-readable instructions, which, when executed on an electronic device, cause the electronic device to implement any of the visual prostheses control methods provided in this application.
[0182] It should also be noted that the device embodiments described above are merely illustrative. The units described as separate components may or may not be physically separate, and the components shown as units may or may not be physical units; that is, they may be located in one place or distributed across multiple network units. Some or all of the modules can be selected to achieve the purpose of this embodiment according to actual needs. In addition, in the device embodiment drawings provided in this application, the connection relationship between modules indicates that they have a communication connection, which can be implemented as one or more communication buses or signal lines.
[0183] Through the above description of the embodiments, those skilled in the art can clearly understand that this application can be implemented by means of software plus necessary general-purpose hardware, or it can be implemented by special-purpose hardware including application-specific integrated circuits, special-purpose CPUs, special-purpose memory, special-purpose components, etc. Generally, any function performed by a computer program can be easily implemented by corresponding hardware, and the specific hardware structure used to implement the same function can also be diverse, such as analog circuits, digital circuits, or special-purpose circuits. However, for this application, software program implementation is more often the preferred implementation method. Based on this understanding, the technical solution of this application, in essence, or the part that contributes to the prior art, can be embodied in the form of a software product. This computer software product is stored in a readable storage medium, such as a computer floppy disk, USB flash drive, mobile hard disk, ROM, RAM, magnetic disk, or optical disk, etc., and includes several instructions to cause a computer device (which may be a personal computer, training equipment, or network device, etc.) to execute the methods described in the various embodiments of this application.
[0184] In the above embodiments, implementation can be achieved, in whole or in part, through software, hardware, firmware, or any combination thereof. When implemented in software, it can be implemented, in whole or in part, as a computer program product.
[0185] The computer program product includes one or more computer instructions. When the computer program instructions are loaded and executed on a computer, all or part of the processes or functions described in the embodiments of this application are generated. The computer may be a general-purpose computer, a special-purpose computer, a computer network, or other programmable device. The computer instructions may be stored in a computer-readable storage medium or transmitted from one computer-readable storage medium to another. For example, the computer instructions may be transmitted from one website, computer, training device, or data center to another website, computer, training device, or data center via wired (e.g., coaxial cable, fiber optic, digital subscriber line (DSL)) or wireless (e.g., infrared, wireless, microwave, etc.) means. The computer-readable storage medium may be any available medium that a computer can store or a data storage device such as a training device or data center that integrates one or more available media. The available media may be magnetic media (e.g., floppy disks, hard disks, magnetic tapes), optical media (e.g., DVDs), or semiconductor media (e.g., solid-state drives (SSDs)).
Claims
1. A control method of a visual prosthesis, characterized by, The method comprises the following steps: voice broadcast is performed on the target object recognition result based on the obtained gray image, and target object image cropping is performed on the gray image based on the target object selection result fed back by the user, to obtain at least one initial target image; a region of interest is extracted from the initial target image to obtain a target object contour image, and the target object contour image comprises a plurality of contour pixel points; a plurality of data groups are obtained by grouping the contour pixel points based on the brightness values of the contour pixel points, and a stimulation sequence of the data groups is set from high to low based on the average brightness values of the data groups, the data groups comprising the spatial coordinates of a plurality of contour pixel points; a stimulation current value corresponding to the minimum non-zero gray value in the gray value of each contour pixel point is obtained based on the gray value of each contour pixel point, and a stimulation current value of each contour pixel point is obtained based on the proportion relationship between the gray value of each contour pixel point and the minimum non-zero gray value. Based on the spatial coordinates and the stimulation current value, a stimulation current matrix corresponding to each data group is mapped, and the stimulation current matrix is used to control the release of stimulation current in the electrode array in sequence based on the stimulation sequence of each data group, the stimulation current matrix comprising electrode channel identifiers corresponding to the spatial coordinates of a plurality of contour pixel points and the stimulation current values corresponding to the contour pixel points.
2. The control method of a visual prosthesis according to claim 1, characterized in that, Before the stimulation current matrix is used to control the release of stimulation current in the electrode array in sequence based on the stimulation sequence of each data group, the method further comprises the following steps: for each stimulation current matrix: obtain the total current value of each stimulation current value in the stimulation current matrix; if the total current value is greater than a preset safety threshold, clip the stimulation current value with the maximum value in the stimulation current matrix and not subjected to clipping based on a clipping coefficient corresponding to the preset safety threshold, to obtain an updated stimulation current matrix, and perform the operation of obtaining the total current value of each stimulation current value in the stimulation current matrix; if the total current value is not greater than the preset safety threshold, perform the operation of using the stimulation current matrix to control the release of stimulation current in the electrode array in sequence based on the stimulation sequence of each data group.
3. The control method of a visual prosthesis according to claim 1, characterized in that, The grouping of the contour pixel points based on the brightness values of the contour pixel points to obtain a plurality of data groups comprises the following steps: determine the brightness distribution based on the brightness values of the contour pixel points; group the contour pixel points based on the brightness distribution to obtain a plurality of data groups, wherein the minimum brightness of the contour pixel points in a first data group is higher than the maximum brightness of the contour pixel points in a second data group, the time sequence of the first data group is earlier than that of the second data group, and the first data group and the second data group are the data groups in the plurality of data groups.
4. The control method of a visual prosthesis according to claim 1, characterized in that, The extraction of the region of interest from the initial target image to obtain the target object contour image comprises the following steps: perform downsampling and Gaussian filtering operations on the initial target image to obtain a plurality of secondary target images with different resolutions of the initial target image; For any two of the secondary target images: unify the size of the two secondary target images, and obtain a feature map by differencing the pixel values of the corresponding two pixel points in the two secondary target images; perform weighted summation on the obtained feature map to obtain a region of interest image; perform contour extraction on the region of interest image using a preset boundary extraction algorithm to obtain the target object contour image.
5. The control method of a visual prosthesis according to claim 4, characterized in that, Before the contour extraction on the region of interest image using the preset boundary extraction algorithm, the method further includes: perform downsampling operation on the region of interest image to obtain a region of interest image with the same size as the electrode array.
6. The control method of a visual prosthesis according to claim 1, characterized in that, The voice broadcast based on the target object recognition result of the obtained gray-scale image includes: obtain an original image and pre-process the original image to obtain the gray-scale image; input the gray-scale image into a preset target recognition model to obtain the type, confidence and bounding box of each target object in the gray-scale image; in the order from high to low of the confidence, sequentially perform the voice broadcast on the type of each target object.
7. The control method of a visual prosthesis according to claim 6, characterized in that, The target object selection result based on the user feedback is used to perform target object image cropping on the gray-scale image to obtain at least one initial target image, including: analyze the target object selection result to obtain the identification of at least one target object; call a preset image cropping algorithm to perform the target object image cropping on the image in the gray-scale image located within the coordinates of the bounding box corresponding to the identification of the target object based on the coordinates of the bounding box of the target object to obtain at least one initial target image.
8. A control system of a visual prosthesis, characterized in that including: a first image extraction module configured to perform voice broadcast based on the target object recognition result of the obtained gray-scale image, and perform target object image cropping on the gray-scale image based on the target object selection result of the user feedback to obtain at least one initial target image; a second image extraction module configured to perform region of interest extraction on the initial target image to obtain a target object contour image, the target object contour image including a plurality of contour pixel points; a parameter obtaining module configured to group the contour pixel points based on the brightness values of the contour pixel points to obtain a plurality of data groups, set the stimulation order of the data groups in the order from high to low of the average brightness values of the data groups, the data group including the spatial coordinates of a plurality of contour pixel points, obtain the stimulation current values of the contour pixel points based on the minimum non-zero gray-scale value in the gray-scale values of the contour pixel points and the proportion relationship between the gray-scale values of the contour pixel points and the minimum non-zero gray-scale value; an array control module configured to map the stimulation current matrix corresponding to each data group based on the spatial coordinates and the stimulation current values, and sequentially control the release of stimulation current in the electrode array based on each stimulation current matrix based on the stimulation order of each data group, the stimulation current matrix including the electrode channel identification corresponding to the spatial coordinates of a plurality of contour pixel points and the stimulation current values corresponding to the contour pixel points.
9. The control system of a visual prosthesis according to claim 8, characterized in that The array control module is further configured to: Before each of the stimulation current matrices is used to control the release of stimulation current in the electrode array in the order of the stimulation sequence based on each of the data sets, each of the stimulation current matrices is: A total current value of each of the stimulation current values in the stimulation current matrix is obtained; If the total current value is greater than a preset safety threshold, one of the stimulation current values in the stimulation current matrix that has the maximum value and has not been clipped is clipped based on a clipping coefficient corresponding to the preset safety threshold, an updated stimulation current matrix is obtained, and the operation of obtaining the total current value of each of the stimulation current values in the stimulation current matrix is performed; If the total current value is not greater than the preset safety threshold, the operation of using each of the stimulation current matrices to control the release of stimulation current in the electrode array in the order of the stimulation sequence based on each of the data sets is performed.
10. A computer storage medium, characterized in that, The storage medium carries one or more computer programs, and when the one or more computer programs are executed by the electronic device, the electronic device can implement the control method of the visual prosthesis according to any one of claims 1 to 7.
Citation Information
Cited By
Intelligent visual prosthesis system and coding method
CN122075932A