Cortical visual prosthesis control method and system based on reinforcement learning and behavior feedback
By using a reinforcement learning and behavioral feedback-based approach, external image feature vectors are acquired and electrical stimulation pulse sequences are generated. Closed-loop optimization is then performed in conjunction with patient behavioral feedback, which solves the encoding failure problem of cortical visual prostheses in long-term blind patients and achieves improved adaptive visual perception stability and recognition accuracy.
Patent Information
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- MINGSHI BRAIN MACHINERY TECHNOLOGY (SUZHOU) CO LTD
- Filing Date
- 2026-02-10
- Publication Date
- 2026-05-01
AI Technical Summary
Existing cortical visual prosthesis technology suffers from coding failure and inaccurate perception in long-term blind patients due to the lack of biological reference standards, making it difficult to achieve continuous and effective visual reconstruction.
A reinforcement learning and behavioral feedback-based approach is adopted to generate an electrical stimulation pulse sequence for a multi-channel microelectrode array by acquiring external image feature vectors, and to perform closed-loop optimization using patient behavioral feedback, thus forming an adaptive control of stimulation-perception-feedback-optimization.
It achieves improved adaptive visual perception stability and recognition accuracy in long-term blind patients, and can adapt online to individual cortical functional reorganization status and cope with photopharma morphological drift.
Smart Images

Figure CN121944384A_ABST
Abstract
Description
A Cortical Visual Prosthesis Control Method and System Based on Reinforcement Learning and Behavioral Feedback Technical Field
[0001] This invention relates to the field of visual prosthesis control technology, and in particular to a cortical visual prosthesis control method and system based on reinforcement learning and behavioral feedback. Background Technology
[0002] Cortical visual prostheses aim to reconstruct partial artificial vision for patients who have lost visual function due to optic nerve damage or long-term blindness by implanting a microelectrode array into the primary visual cortex (V1) to apply electrical stimulation to the brain to induce photic hallucinations. The core of these prostheses lies in the coding strategy: how to efficiently and accurately convert image information acquired by an external camera into an electrical stimulation pattern suitable for the cortical neural response.
[0003] Current mainstream technologies generally employ static coding methods based on retinal topological mapping: First, during the initial implantation phase, patients verbally describe the location of light spots by stimulating each electrode individually, establishing a lookup table of "image pixels – electrode locations." Then, during the usage phase, the processed image is directly mapped to the corresponding electrode and stimulated at a fixed frequency. While this method can achieve basic shape recognition under ideal conditions, it faces fundamental obstacles in clinical practice. Long-term blindness leads to functional reorganization of the visual cortex, and patients cannot provide "standard" neural activity corresponding to specific images as supervisory signals. This makes coding strategies relying on prior biological mapping or healthy neural response patterns difficult to establish or rapidly fail under pathological conditions. Furthermore, even after initial calibration, the induced optic hallucinations often experience morphological drift due to dynamic changes in the nervous system, causing the fixed coding strategy to mismatch over time and lacking continuous and effective closed-loop optimization capabilities. Therefore, there is an urgent need for a control method that does not rely on prior knowledge of damaged pathways and can autonomously learn and continuously optimize stimulation strategies based on available clinical signals to overcome coding failure and perceptual inaccuracies caused by the lack of biological reference standards. Summary of the Invention
[0004] In view of this, the present invention proposes a method and system for controlling cortical visual prostheses based on reinforcement learning and behavioral feedback, which can significantly improve the stability and recognition accuracy of artificial visual perception. The present invention provides the following technical solution: a method for controlling cortical visual prostheses based on reinforcement learning and behavioral feedback, the method comprising: acquiring a visual image of the external environment and extracting a feature vector from the visual image; inputting the feature vector into a reinforcement learning environment to generate an electrical stimulation pulse sequence corresponding to multiple microelectrode channels; applying the electrical stimulation pulse sequence to a microelectrode array implanted in the patient's visual cortex to induce photophasic visual perception; acquiring behavioral feedback generated by the patient performing a visual recognition task based on the photophasic visual perception; and updating the parameters of the reinforcement learning environment based on the behavioral feedback to form a closed-loop optimization.
[0005] Optionally, extracting features from the visual image includes: performing semantic-level feature extraction on the visual image using a deep neural network to obtain feature vectors representing key visual content of the image, wherein the key visual content includes at least one of object edge structure, texture distribution, or semantic category information.
[0006] Optionally, inputting the feature vector into the reinforcement learning environment to generate an electrical stimulation pulse sequence corresponding to multiple microelectrode channels includes: inputting the feature vector and historical pulse sequences as joint inputs into a multivariable point process model; calculating the stimulation probability of each microelectrode channel at a discrete time step through the multivariable point process model; and generating a binarized pulse output for each channel by random sampling based on the stimulation probability to form a multi-channel pulse sequence.
[0007] Optionally, during the application of the electrical stimulation pulse sequence, the cumulative charge of each microelectrode channel is monitored in real time, and when the cumulative charge of any channel reaches a preset charge safety threshold, the pulse output of that channel is forcibly suppressed until the cumulative charge decays to below the safe range.
[0008] Optionally, the behavioral feedback includes external rewards and internal rewards, wherein: the external rewards are generated based on the degree of matching between the patient's judgment result on the visual recognition task and the real image labels, with a positive reward value assigned when the match is successful and a zero or negative reward value assigned when the match fails; the internal rewards are calculated based on the perceptual category distribution induced by historical stimuli, and when the perceptual category distribution deviates from a preset uniformity threshold, additional internal reward weights are assigned to perceptual categories that occur infrequently.
[0009] Optionally, updating the parameters of the reinforcement learning environment based on the behavioral feedback includes: fusing the external reward and the internal reward into a total reward signal using a weighted summation method; calculating the parameter update amount using a policy gradient algorithm based on the total reward signal; and updating the parameters of the reinforcement learning environment according to the parameter update amount, so that the subsequently generated electrical stimulation pulse sequence tends to improve the success rate of the visual recognition task.
[0010] This invention further discloses a cortical visual prosthesis control system based on reinforcement learning and behavioral feedback, comprising: an image acquisition module for acquiring visual images of the external environment; a feature extraction module for extracting feature vectors from the visual images; a pulse generation module for inputting the feature vectors into a reinforcement learning environment to generate an electrical stimulation pulse sequence corresponding to multiple microelectrode channels; a stimulation driving module for converting the electrical stimulation pulse sequence into electrical stimulation commands and applying them to the microelectrode array to induce photophany perception; a feedback acquisition module for acquiring behavioral feedback generated by the patient performing a visual recognition task based on the photophany perception; and a parameter update module for updating the parameters of the reinforcement learning environment based on the behavioral feedback to form a closed-loop optimization.
[0011] The present invention further discloses a computer-readable storage medium storing a computer program that, when executed by a processor, implements the above-described method.
[0012] The present invention further discloses an electronic device, including a memory, a processor, and a computer program stored in the memory and executable on the processor, wherein the processor executes the program to implement the above-described method.
[0013] The present invention further discloses a computer program product, including a computer program, characterized in that the computer program implements the above-described method when executed by a processor.
[0014] According to the technical solution of this invention, a closed-loop control based on reinforcement learning and behavioral feedback effectively solves the coding failure problem caused by the lack of biological reference standards in the clinical application of cortical visual prostheses. It directly generates an electrical stimulation pulse sequence adapted to a multi-channel microelectrode array by using the semantic features of external images as the state input of the reinforcement learning environment. Crucially, it continuously updates the reinforcement learning strategy parameters by acquiring the behavioral feedback generated by the patient performing visual recognition tasks, forming an adaptive closed loop of stimulation-perception-feedback-optimization. This enables the stimulation strategy to evolve online, automatically adapting not only to the cortical functional reorganization state caused by long-term blindness but also effectively addressing the morphological drift of photopsia over time, thereby significantly improving the stability and long-term usability of artificial visual perception. Attached Figure Description
[0015] For illustrative purposes and not for limitation, the present invention will now be described in conjunction with embodiments and accompanying drawings, wherein: Figure 1 is a flowchart illustrating the cortical visual prosthesis control method based on reinforcement learning and behavioral feedback according to an embodiment of the present invention; Figure 2 is a schematic diagram of the constituent modules of the cortical visual prosthesis control system based on reinforcement learning and behavioral feedback according to an embodiment of the present invention; Figure 3 is a schematic diagram of the constituent structure of the electronic device according to an embodiment of the present invention; Figure 4 is a schematic diagram of the architecture of the cortical visual prosthesis control system based on reinforcement learning and behavioral feedback according to an embodiment of the present invention; and Figure 5 is a comparison diagram of visual perception effects according to an embodiment of the present invention. Detailed Implementation
[0016] To enable those skilled in the art to better understand the present application, the technical solutions in the embodiments of the present application will be clearly and completely described below with reference to the accompanying drawings. Obviously, the described embodiments are only a part of the embodiments of the present application, and not all of the embodiments. Based on the embodiments in the present application, all other embodiments obtained by those skilled in the art without creative effort should fall within the scope of protection of the present application.
[0017] It should be noted that, where there is no conflict, the embodiments and features of the embodiments in this application can be combined with each other. The embodiments of this application will be described in detail below with reference to the accompanying drawings.
[0018] Referring to Figures 1 and 5, this embodiment discloses a cortical visual prosthesis control method based on reinforcement learning and behavioral feedback. The method includes the following steps: S100: Acquiring visual images of the external environment and extracting feature vectors from the visual images. Specifically, for acquiring the visual images, a miniature camera worn on the patient's head is used to collect visual images of the external environment in real time. The acquired raw images are first preprocessed, including converting the RGB three-channel image to a grayscale image to reduce computational complexity, followed by contrast adaptive histogram equalization to enhance local details, and normalizing pixel values to the [0,1] interval to form a standardized input image. For feature vector extraction, the preprocessed standardized input image is input into a convolutional neural network (CNN) for semantic-level feature extraction to obtain the feature vectors of the visual image. , the feature vector As the state input for the reinforcement learning environment, this implementation example uses an 18-layer ResNet-18 residual network as the feature extraction backbone. During the forward propagation of the network, the original image sequentially passes through a 7×7 convolutional layer, a max-pooling layer, and four residual block groups, finally outputting a 512-dimensional feature vector at the global average pooling layer. This feature vector can effectively characterize the key visual content of the image, including but not limited to: object edge structure: contour and boundary information captured by the activation response of shallow convolutional kernels; texture distribution: surface texture features obtained through spatial frequency analysis of mid-layer feature maps; semantic category information: object category probability distribution (such as letters, simple geometric shapes, etc.) output by deep fully connected layers.
[0019] S200: Input the feature vector into the reinforcement learning environment to generate an electrical stimulation pulse sequence corresponding to multiple microelectrode channels.
[0020] In this embodiment, the reinforcement learning environment is a reinforcement learning point process model (RLPP) designed specifically for cortical electrical stimulation. It is a parameterized policy model used to determine the optimal action (i.e., multi-channel electrical stimulation pulse sequence) based on the currently perceived environmental state (i.e., image feature vector) in order to maximize the long-term accumulated reward (i.e., behavioral feedback).
[0021] Specifically, the input to the point process model includes the current state and historical information, where the current state is a feature vector representing the current visual image. The historical information consists of the historical pulse sequences output by each channel within a preset time window, enabling the point process model to learn the temporal patterns of stimuli. Subsequently, the feature vectors... The historical pulse sequence is fed as a joint input into the multivariable point process model. Based on this joint input, the multivariable point process model calculates the stimulation probability of each microelectrode channel at the current discrete time step. The calculation formula is: ,in, and These are the network parameters. The output layer dimension of the model is strictly consistent with the number of physical channels in the microelectrode array, ensuring that each output dimension directly corresponds to one physical electrode channel, thereby establishing a direct mapping relationship between algorithmic decisions and physical stimuli.
[0022] The probability of obtaining stimulation from all channels Then, a specific stimulus pulse sequence is generated through a non-homogeneous Bernoulli process. For each channel, based on its stimulus probability The system independently determines whether to output a stimulus pulse at the current time step ("1" indicates an output pulse, "0" indicates no pulse). This process is applied in parallel to all channels, ultimately forming a multi-channel synchronized binary pulse sequence. This pulse sequence has a fine temporal structure, which can simulate the temporal characteristics of biological neuron firing, significantly different from the constant frequency stimulation mode used in traditional schemes.
[0023] S300: Apply the electrical stimulation pulse sequence to the microelectrode array implanted in the patient's visual cortex to induce photophantom perception.
[0024] Specifically, based on the output stimulus pulse sequence in step S200 The signal is a time-series signal consisting of binary variables (0 or 1), where each bit corresponds to a stimulation or non-stimulation instruction for a microelectrode channel at a specific time step. For each stimulation event marked as "1" in the pulse sequence, it is converted into a set of specific electrical stimulation parameters. The electrical stimulation parameters include at least: current amplitude, pulse width, and pulse waveform. The electrical stimulation parameters are synchronously applied to the microelectrode array implanted in the patient's primary visual cortex (V1 area). Each microelectrode in the array corresponds strictly to a specific channel in the pulse sequence. When the instruction for a channel in the sequence is 1 at time step k, the corresponding electrode tip applies an electrical pulse defined by the above parameters to the surrounding neuronal group.
[0025] Furthermore, during stimulation, the accumulated charge of each microelectrode channel is monitored in real time. Specifically, for each channel, the total charge of the applied pulse is continuously accumulated. When the accumulated charge of any channel reaches the preset charge safety threshold CICmax, the pulse output of that channel is immediately and forcibly suppressed, keeping it in a "0" state in subsequent time steps until the accumulated charge decreases below the safe range through natural tissue decay. This safety constraint mechanism is directly embedded in the stimulation generation process as a hard limit, ensuring biocompatibility for long-term use.
[0026] Finally, when electrical stimulation is applied to cortical neurons, patients subjectively perceive discrete points or spots of light, a phenomenon known as phosphene. Coordinated stimulation of multiple electrodes can combine to form phosphene patterns with spatial structures, enabling patients to recognize basic visual elements such as line directions, simple geometric shapes, or letter outlines.
[0027] S400: Obtain behavioral feedback generated by the patient performing a visual recognition task based on the aforementioned optical illusion perception.
[0028] Specifically, patients are presented with photohallucinations induced by electrical stimulation, along with corresponding real-world image labels. Based on their subjective perception of the photohallucinations, patients perform visual recognition tasks through an interactive interface, such as identifying the direction of lines (horizontal / vertical / diagonal), geometric shapes (circles / squares / triangles), or letter categories represented by the photohallucinations. Patients select their identification results via buttons, voice input, or a touchscreen, and their behavioral responses are recorded in real time.
[0029] The patient's discrimination result is compared with the real image label to generate an external reward signal. When the discrimination result matches the real label, a positive reward value (+1) is assigned; when the discrimination fails, a zero reward is assigned. This external reward directly represents the quality of visual perception induced by the current stimulus strategy.
[0030] To further address the environmental non-stationarity issue caused by photophany drift, an internal reward signal is calculated. Specifically, the frequency of various perceptual outcomes reported by patients within a preset time window is statistically analyzed to form a perceptual category distribution. When this distribution exhibits significant bias (i.e., the model excessively generates a certain type of stimulus pattern, causing patients to continuously report the same or similar perceptual categories), it is determined that the distribution deviates from a uniform state. In this case, additional internal reward weights are assigned to perceptual categories that appear less frequently or not at all. This mechanism forces the model to maintain its ability to continuously explore the stimulus space, avoiding the model falling into a locally optimal solution that fails due to the time-varying nature of photophany patterns.
[0031] Finally, the external and internal rewards are weighted and summed to form the total reward signal, calculated using the following formula: ,in As an external reward, As an internal reward, This is a balancing factor (which gradually decays with increasing training rounds). This is the total reward signal. It serves as feedback input to the reinforcement learning environment and is used for subsequent policy parameter updates.
[0032] This implementation introduces an internal reward mechanism. By statistically analyzing the distribution of patient-reported perceptual categories within a preset time window, when the model is detected to be excessively biased towards outputting a single or a few perceptual categories (i.e., the distribution deviates from uniformity), additional exploratory reward weights are automatically assigned to perceptual categories that appear infrequently or not at all. This forces the model to maintain its ability to continuously explore the stimulus space. The core function of this mechanism is to effectively address the environmental non-stationarity problem caused by photophany drift. Due to the dynamic changes in the nervous system, the photophany patterns produced by the same electrode stimulation will drift over time. Without an exploration mechanism, the model will adhere to the old, ineffective stimulus patterns and will be unable to adapt. Through the continuous drive of internal rewards, the model can actively try new stimulus combinations when photophany drift occurs, quickly discover and converge to new effective encoding strategies, thereby significantly improving the long-term stability and adaptability of artificial visual perception and avoiding the degradation of perceptual quality caused by changes in neural plasticity.
[0033] S500: Update the parameters of the reinforcement learning environment based on the behavioral feedback to form a closed-loop optimization.
[0034] In this embodiment, the policy gradient algorithm is used to calculate the update amount of the policy model parameters in order to update the model parameters. The gradient calculation formula is as follows: ,in, Indicates in image features Take stimulating actions The probability, Represents the cumulative return value. For the learnable parameters of the point process model, Indicates the parameter The gradient represents the rate of change of the parameters. Based on this gradient calculation, the policy model parameters are updated to make subsequent generated electrical stimulation pulse sequences more likely to improve the success rate of visual recognition tasks.
[0035] This parameter update process is executed immediately after each visual recognition task, forming a complete closed loop of image input, feature extraction, impulse generation, cortical stimulation, behavioral feedback, and parameter update. This closed-loop optimization mechanism enables continuous adaptation; that is, when a patient's photopsia undergoes morphological drift due to neural plasticity, the behavioral feedback signal changes accordingly, driving the strategy model to automatically adjust the stimulation pattern, quickly discovering and converging to a new effective encoding strategy. This effectively overcomes the long-term failure problem caused by the lack of perceptual stability in traditional fixed encoding schemes.
[0036] It should be noted that the closed-loop optimization in this embodiment relies on behavioral feedback generated by the visual recognition tasks that the patient can perform, without the need to collect any neural signals as supervisory labels. This effectively bypasses the fundamental clinical limitation that blind patients cannot provide standard V1 cortical neural response data due to optic nerve atrophy or functional impairment, enabling adaptive coding learning under closed-loop optimization to be achieved in pathological conditions.
[0037] In summary, this embodiment extracts feature vectors from external images as the state input to the reinforcement learning environment. It directly generates millisecond-level temporal pulse sequences adapted to a multi-channel microelectrode array using a multivariable point process model, and applies these sequences to the visual cortex to induce photophanosis. This embodiment does not rely on prior knowledge of healthy neural activity; instead, it uses the behavioral feedback (recognition accuracy) generated by the patient performing a visual recognition task as an optimization signal to drive the online updating of reinforcement learning environment parameters, forming an adaptive closed loop of stimulus-perception-feedback-optimization. Referring to Figure 5, the essential differences between this embodiment and the retinal topological mapping method in terms of stimulus encoding mechanism and final visual perception effect are illustrated. Specifically, the rigid rule of direct pixel-electrode mapping directly converts the input image into a fixed-frequency pulse sequence. This sequence lacks spatiotemporal dynamics, resulting in a blurred and distorted perception effect in the induced photophanosis. In contrast, this embodiment uses an RLPP generative network to achieve adaptive encoding, transforming the same input image into a biomimetic pulse sequence with rich spatiotemporal information. This sequence simulates the temporal characteristics of the natural firing patterns of neurons, enabling the patient to obtain clear and accurate visual perception. It achieves a self-learning mechanism driven by behavioral feedback, bypassing the problem of missing signal reference. This enables the encoding strategy to automatically adapt to the cortical functional reorganization state caused by long-term blindness to overcome the lack of spatial reference. Furthermore, it maintains the ability to continuously explore the stimulus space through an internal reward mechanism to cope with the loss of perceptual stability caused by photophobia. Ultimately, it achieves high-information-efficiency artificial visual reconstruction that is calibration-free, stable over a long period, and personalized, thereby improving the stability and recognition accuracy of artificial visual perception.
[0038] Referring to Figures 2 and 4, this embodiment further discloses a cortical visual prosthesis control system based on reinforcement learning and behavioral feedback, including an image acquisition module 21, a feature extraction module 22, a pulse generation module 23, a stimulus driving module 24, a feedback acquisition module 25, and a parameter update module 26. The following is a detailed description: The image acquisition module 21 is used to acquire visual images of the external environment. The image acquisition module 21 consists of a miniature CMOS image sensor and a matching optical lens, integrated into the front end of the smart glasses frame or other camera device worn by the patient. Raw image data is transmitted to the system's main processor via an interface. The image acquisition module 21 continuously captures the visual scene in the patient's forward field of vision, providing the raw input source for subsequent processing.
[0039] The feature extraction module 22 is used to extract the feature vector of the visual image. The feature extraction module 22 runs a convolutional neural network model and receives the raw image data output by the image acquisition module 21. The feature extraction module 22 first performs grayscale conversion and contrast adaptive histogram equalization preprocessing on the image, then inputs it into a residual network for forward inference, and finally outputs the feature vector corresponding to the visual image. This feature vector represents the high-level visual content of the image, including edge structure, texture distribution, and semantic category information, and is passed as the state input to the pulse generation module 23 of the reinforcement learning environment.
[0040] The pulse generation module 23 is used to input the feature vector into the reinforcement learning environment to generate electrical stimulation pulse sequences corresponding to multiple microelectrode channels. The pulse generation module 23 includes a reinforcement learning-based point process model and its operating environment. It receives the feature vector output by the feature extraction module 22 and its own cached historical pulse sequences, and calculates the stimulation probability of each microelectrode channel at the current time step using a multi-layer neural network. The calculation formula is as follows: ,in, and These are the network parameters; the dimension of the model output layer is strictly consistent with the number of microelectrode channels. Based on the calculated probability values, a binary pulse sequence is generated through non-homogeneous Bernoulli sampling and output to the stimulation driving module 24.
[0041] The stimulation-driven module 24 is used to convert the electrical stimulation pulse sequence into electrical stimulation commands and apply them to the microelectrode array to induce photophasic perception. Specifically, it receives the stimulation pulse sequence output by the pulse generation module 23, which is a time-series signal composed of binary variables, where each bit corresponds to a stimulation or non-stimulation command for a microelectrode channel at a specific time step. For each stimulation event marked as "1" in the pulse sequence, it is converted into a set of specific electrical stimulation parameters. The electrical stimulation parameters are synchronously applied to the microelectrode array implanted in the patient's primary visual cortex (V1 area). When the command for a channel in the sequence is 1 at a time step, the corresponding electrode tip applies an electrical pulse defined by the above parameters to the surrounding neuronal group. During the stimulation application process, the cumulative charge of each microelectrode channel is monitored in real time. Specifically, for each channel, the total charge of the applied pulse is continuously accumulated. When the cumulative charge of any channel reaches a preset charge safety threshold, the pulse output of that channel is immediately forcibly suppressed, keeping it in a "0" state in subsequent time steps until the cumulative charge decreases below the safe range through natural tissue decay. When electrical stimulation is applied to cortical neurons, patients subjectively perceive discrete spots or patches of light, a phenomenon known as photophany. Synergistic stimulation of multiple electrodes can combine to form photophany patterns with spatial structures, enabling patients to recognize basic visual elements.
[0042] The feedback acquisition module 25 is used to acquire behavioral feedback generated by the patient performing a visual recognition task based on the photopsychoscopic perception. The feedback acquisition module 25 includes an interaction interface unit 251 and a reward calculation unit 252. The interaction interface unit 251 allows the patient to perform a visual recognition task based on photopsychoscopic perception (such as judging line direction, geometric shape, or letter category); the patient's selection result is transmitted to the reward calculation unit 252 via a serial communication interface. The reward calculation unit 252 compares the patient's judgment result with the real image labels stored internally in the system. If a match is successful, an external reward value of +1 is generated; if a match fails, a reward of 0 is generated. Simultaneously, the reward calculation unit 252 statistically analyzes the distribution of perception categories within a preset time window. When the distribution deviates from uniformity, an internal reward value is calculated, and finally, the fused total reward signal is output to the parameter update module 26.
[0043] The parameter update module 26 is used to update the parameters of the reinforcement learning environment based on the behavioral feedback, forming a closed-loop optimization. Specifically, it receives the total reward signal output by the feedback acquisition module 25, and calculates the parameter update amount using a policy gradient algorithm. The gradient calculation formula is as follows: ,in, Indicates in image features Take stimulating actions The probability, Represents the cumulative return value. For the learnable parameters of the point process model, Indicates the parameter The gradient represents the rate of change of the parameters. The parameters θ of the policy model in the pulse generation module 23 are updated along the gradient direction, so that the subsequently generated stimulus policies tend to improve the visual recognition success rate. This update process is executed immediately after each recognition task, forming a complete closed-loop optimization mechanism.
[0044] The image acquisition module 21, feature extraction module 22, pulse generation module 23, stimulus driving module 24, feedback acquisition module 25, and parameter update module 26 mentioned above are interconnected through the system bus to form a closed-loop data flow of image acquisition-feature extraction-pulse generation-stimulation driving-feedback acquisition-parameter update.
[0045] Figure 3 is a schematic diagram of the physical structure of the electronic device provided in the embodiment of the present invention. As shown in Figure 3, the electronic device 50 includes: a processor 501, a memory 502, and a bus 503; wherein, the processor 501 and the memory 502 communicate with each other through the bus 503; the processor 501 is used to call the program instructions in the memory 502 to execute the methods provided in the above-described method embodiments.
[0046] This embodiment provides a non-transitory computer-readable storage medium that stores computer instructions that cause a computer to execute the methods provided in the above-described embodiments.
[0047] Those skilled in the art will understand that all or part of the steps of the above-described method implementation can be implemented by hardware related to program instructions. The aforementioned program can be stored in a computer-readable storage medium. When the program is executed, it performs the steps of the above-described method implementation. The aforementioned storage medium includes various storage media capable of storing program code, such as ROM, RAM, magnetic disk, or optical disk.
[0048] The device embodiments described above are merely illustrative. The units described as separate components may or may not be physically separate, and the components shown as units may or may not be physical units; that is, they may be located in one place or distributed across multiple network units. Some or all of the modules can be selected to achieve the purpose of this embodiment according to actual needs. Those skilled in the art can understand and implement this without any creative effort.
[0049] Through the above description of the embodiments, those skilled in the art can clearly understand that each embodiment can be implemented by means of software plus necessary general-purpose hardware platforms, and of course, it can also be implemented by hardware. Based on this understanding, the above technical solutions, in essence or the part that contributes to the prior art, can be embodied in the form of a software product. This computer software product can be stored in a computer-readable storage medium, such as ROM / RAM, magnetic disk, optical disk, etc., including several instructions to cause a computer device (which may be a personal computer, server, or network device, etc.) to execute the methods of each embodiment or some parts of the embodiments.
[0050] The specific embodiments described above do not constitute a limitation on the scope of protection of this invention. Those skilled in the art should understand that various modifications, combinations, sub-combinations, and substitutions can occur depending on design requirements and other factors. Any modifications, equivalent substitutions, and improvements made within the spirit and principles of this invention should be included within the scope of protection of this invention.
Claims
1. A method for controlling cortical visual prostheses based on reinforcement learning and behavioral feedback, characterized in that, The method includes: acquiring a visual image of the external environment and extracting a feature vector from the visual image; inputting the feature vector into a reinforcement learning environment to generate an electrical stimulation pulse sequence corresponding to multiple microelectrode channels; applying the electrical stimulation pulse sequence to a microelectrode array implanted in the patient's visual cortex to induce photophany; acquiring behavioral feedback generated by the patient performing a visual recognition task based on the photophany; and updating the parameters of the reinforcement learning environment based on the behavioral feedback to form a closed-loop optimization.
2. The cortical visual prosthesis control method based on reinforcement learning and behavioral feedback according to claim 1, characterized in that, The extraction of features from the visual image includes: performing semantic-level feature extraction on the visual image using a deep neural network to obtain feature vectors representing key visual content of the image, wherein the key visual content includes at least one of object edge structure, texture distribution, or semantic category information.
3. The cortical visual prosthesis control method based on reinforcement learning and behavioral feedback according to claim 1, characterized in that, The step of inputting the feature vector into the reinforcement learning environment to generate an electrical stimulation pulse sequence corresponding to multiple microelectrode channels includes: inputting the feature vector and historical pulse sequence as joint inputs into a multivariable point process model; calculating the stimulation probability of each microelectrode channel at a discrete time step through the multivariable point process model; and generating a binarized pulse output of each channel by random sampling based on the stimulation probability to form a multi-channel pulse sequence.
4. The cortical visual prosthesis control method based on reinforcement learning and behavioral feedback according to any one of claims 1-3, characterized in that, During the application of the electrical stimulation pulse sequence, the cumulative charge of each microelectrode channel is monitored in real time. When the cumulative charge of any channel reaches a preset charge safety threshold, the pulse output of that channel is forcibly suppressed until the cumulative charge decays to below the safe range.
5. The cortical visual prosthesis control method based on reinforcement learning and behavioral feedback according to claim 1, characterized in that, The behavioral feedback includes external rewards and internal rewards. The external rewards are generated based on the degree of matching between the patient's judgment result on the visual recognition task and the real image label. A positive reward value is given when the match is successful, and a zero or negative reward value is given when the match fails. The internal rewards are calculated based on the distribution of perceptual categories induced by historical stimuli. When the distribution of perceptual categories deviates from a preset uniformity threshold, additional internal reward weights are given to perceptual categories that occur infrequently.
6. The method for controlling cortical visual prostheses according to claim 1, characterized in that, The step of updating the parameters of the reinforcement learning environment based on the behavioral feedback includes: fusing the external reward and the internal reward into a total reward signal using a weighted summation method; calculating the parameter update amount using a policy gradient algorithm based on the total reward signal; and updating the parameters of the reinforcement learning environment according to the parameter update amount, so that the subsequently generated electrical stimulation pulse sequence tends to improve the success rate of the visual recognition task.
7. A cortical visual prosthesis control system based on reinforcement learning and behavioral feedback, characterized in that, include: The image acquisition module is used to acquire visual images of the external environment; The feature extraction module is used to extract the feature vector of the visual image; The pulse generation module is used to input the feature vector into the reinforcement learning environment to generate an electrical stimulation pulse sequence corresponding to multiple microelectrode channels; The stimulation driving module is used to convert the electrical stimulation pulse sequence into electrical stimulation commands and apply them to the microelectrode array to induce photopharma perception; The feedback acquisition module is used to acquire the behavioral feedback generated by the patient performing a visual recognition task based on the optical illusion perception; The parameter update module is used to update the parameters of the reinforcement learning environment based on the behavioral feedback, forming a closed-loop optimization.
8. A computer-readable storage medium, characterized in that, The storage medium stores a computer program that, when executed by a processor, implements the method of any one of claims 1-6.
9. An electronic device comprising a memory, a processor, and a computer program stored in the memory and executable on the processor, characterized in that, When the processor executes the program, it implements the method of any one of claims 1-6.
10. A computer program product, comprising a computer program, characterized in that, The computer program, when executed by a processor, implements the method of any one of claims 1-6.