A parameter adaptive stochastic resonance system and method based on reinforcement learning
Patent Information
- Application Number
- CN202610838388.5
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2026-06-11
- Publication Date
- 2026-09-22
AI Technical Summary
现有随机共振信号检测系统存在以下缺陷:传统固定参数随机共振系统需人工设定双稳态参数,无法适配水下动态噪声与信道变化,增强效果受限;基于粒子群优化(ParticleSwarm Optimization,PSO)、多策略融合粒子群优化(Multi-strategy Fusion ParticleSwarm Optimization,MFPSO)等算法的自适应随机共振方法,存在计算复杂度高、迭代次数多、运行耗时久的问题,难以满足水下通信实时检测需求;现有随机共振硬件电路参数调节不灵活,无法与算法形成实时闭环,难以部署于边缘计算设备;传统优化算法易陷入局部最优,无法精准搜索最优随机共振参数,弱信号增强效果达不到最优
本公开的实施例中,在信号采集模块、前端调理模块、随机共振模块、强化学习模块和输出模块的协同配合下,实现对带噪信号进行预处理、对预处理后的带噪信号进行随机共振增强、对增强得到的还原信号进行强化学习评价以及共振参数的反馈调节,从而得到最终还原信号;通过随机共振模块可将共振参数直接映射为可调控的数字电位器的等效参数,实现了共振参数的在线闭环调节,以便部署于边缘计算设备。此外,在随机共振模块与强化学习模块的协同配合下,使系统在低信噪比和动态噪声环境下仍具有良好的增强能力、自适应能力和实时性,同时基于强化学习模块将输出信噪比作为共振参数更新的及时奖励,能够更准确地反映共振参数对信号增强质量的影响,确保强化学习模块最终迭代更新得到的最优共振参数能够使随机共振模块对带噪信号的增强效果达到最优。
Smart Images

Figure CN122802039A_ABST
Abstract
Description
Technical Field
[0001] This disclosure relates to the field of underwater wireless optical communication technology, and in particular to a parameter adaptive stochastic resonance system and method based on reinforcement learning. Background Technology
[0002] Underwater Wireless Optical Communication (UWOC) has become a core communication method for marine resource exploration and underwater equipment interconnection due to its advantages such as high transmission rate, low latency, and strong resistance to electromagnetic interference. However, interference from seawater absorption, scattering, turbulence, and detector thermal noise can cause significant attenuation of the received optical signal, resulting in a weak signal with a low signal-to-noise ratio. This makes reliable decision-making and transmission difficult, severely limiting communication distance and system performance.
[0003] Stochastic Resonance (SR), a nonlinear weak signal enhancement technique, can improve the output signal-to-noise ratio through the synergistic effect of noise, weak signal, and bistable systems, and has been applied in the field of weak signal detection in underwater optical communication. Existing stochastic resonance signal detection systems suffer from the following drawbacks: traditional fixed-parameter stochastic resonance systems require manual setting of bistable parameters, which cannot adapt to dynamic underwater noise and channel variations, limiting the enhancement effect; adaptive stochastic resonance methods based on algorithms such as Particle Swarm Optimization (PSO) and Multi-strategy Fusion Particle Swarm Optimization (MFPSO) suffer from high computational complexity, numerous iterations, and long execution times, making it difficult to meet the real-time detection requirements of underwater communication; existing stochastic resonance hardware circuit parameters are inflexible to adjust, unable to form a real-time closed loop with the algorithm, and difficult to deploy on edge computing devices; traditional optimization algorithms are prone to getting trapped in local optima, unable to accurately search for the optimal stochastic resonance parameters, resulting in suboptimal weak signal enhancement.
[0004] Therefore, it is necessary to provide a new technical solution to improve one or more of the problems existing in the above solutions.
[0005] It should be noted that the information disclosed in the background section above is only used to enhance the understanding of the background of this disclosure, and therefore may include information that does not constitute prior art known to those skilled in the art. Summary of the Invention
[0006] The purpose of this disclosure is to provide a parameter adaptive stochastic resonance system and system based on reinforcement learning, which can achieve optimal enhancement effect on noisy signals and is easy to deploy on edge computing devices.
[0007] According to a first aspect of the present disclosure, a parameter adaptive stochastic resonance system based on reinforcement learning is provided, comprising: The signal acquisition module is used to receive noisy signals; A front-end conditioning module is used to preprocess the noisy signal; The random resonance module is used to enhance the noisy signal output by the front-end conditioning module to obtain the restored signal to be evaluated, and to enhance the noisy signal output by the front-end conditioning module under the optimal resonance parameter settings fed back by the reinforcement learning module to obtain the final restored signal. The reinforcement learning module is used to perform statistics on the restored signal to be evaluated, obtain the output signal-to-noise ratio, and update and optimize the resonance parameters based on the output signal-to-noise ratio to obtain the optimal resonance parameters. The output module is used to output the final restored signal.
[0008] In an exemplary embodiment of this disclosure, the stochastic resonance module includes an integrator, an inverter, a first digital potentiometer, a second digital potentiometer, and a multiplication unit for forming a parameter-adjustable bistable closed-loop circuit. An inverter is used to form a primary feedback branch together with the second digital potentiometer and the integrator; The multiplication unit, together with the first digital potentiometer and the integrator, forms a cubic nonlinear feedback term; The integrator is used to receive the noisy signal output by the front-end conditioning module and to perform integration operations on the noisy signal, the primary feedback branch signal, and the tertiary nonlinear feedback branch signal.
[0009] In an exemplary embodiment of this disclosure, the stochastic resonance theoretical model of the stochastic resonance module is as follows: (1) in, Let represent the integral term of the stochastic resonance module, and let x represent the output state of the stochastic resonance module. , Indicates the first equivalent parameter. The parameters representing the dynamic equations, , Indicates the second equivalent parameter. The parameters representing the dynamic equations, Indicates a useful signal and , Indicates noise signal and RC represents the time constant formed by the integrator resistor R and capacitor C in the random resonance module.
[0010] In an exemplary embodiment of this disclosure, the dynamic equation is: (2) Where a and b represent the parameters of the dynamic equation, s(t) represents the useful signal, h(t) represents the noise signal, and x represents the output state of the stochastic resonance module. express Regarding time The first derivative of .
[0011] In an exemplary embodiment of this disclosure, the multiplication unit includes a first multiplier and a second multiplier, the first multiplier and the second multiplier being cascaded between the output and input of an integrator, the first digital potentiometer being electrically connected to the integrator and the second multiplier, and the second digital potentiometer being electrically connected to the integrator and the inverter.
[0012] In an exemplary embodiment of this disclosure, the resonance parameter includes a first equivalent parameter and a second equivalent parameter, wherein the first equivalent parameter is written into the first digital potentiometer and the second equivalent parameter is written into the second digital potentiometer.
[0013] In one exemplary embodiment of this disclosure, the reinforcement learning module includes: An extraction unit is used to extract a first sample sequence and a second sample sequence from the restored signal to be evaluated. The calculation unit is used to perform statistical analysis on the first sample sequence and the second sample sequence to obtain the noise power corresponding to the restored signal, and calculate the output signal-to-noise ratio based on the noise power.
[0014] According to a second aspect of the present disclosure, a parameter adaptive stochastic resonance method based on reinforcement learning is provided, comprising: Receive noisy signals and preprocess the noisy signals; The preprocessed noisy signal is enhanced under the resonance parameter settings during the iterative stage to obtain the restored signal to be evaluated. The restored signal to be evaluated is statistically analyzed to obtain the output signal-to-noise ratio, and the resonance parameters are updated and optimized based on the output signal-to-noise ratio to obtain the optimal resonance parameters; The preprocessed noisy signal is enhanced under the optimal resonance parameter settings to obtain the final restored signal, and the final restored signal is output.
[0015] In an exemplary embodiment of this disclosure, the step of statistically analyzing the restored signal to be evaluated to obtain the output signal-to-noise ratio includes: Extract the first sample sequence and the second sample sequence from the restored signal to be evaluated; Statistical analysis is performed on the first and second sample sequences to calculate the noise power corresponding to the restored signal. (3) in, Indicates noise power, N1 represents the sample length of the first sample sequence, N0 represents the sample length of the second sample sequence, and I 11 I represents the mean of the first sample sequence. 00 This represents the mean of the second sample sequence, where i represents the sample number; Calculate the output signal-to-noise ratio based on noise power: (4) in, Indicates the output signal-to-noise ratio. P represents noise power. s Indicates signal power.
[0016] In an exemplary embodiment of this disclosure, the step of updating and optimizing the resonance parameters based on the output signal-to-noise ratio includes: Based on the preset parameter range and preset quantization step size, the resonance parameters are discretized into a two-dimensional index space, and a reward table is established. The resonance parameters include a first equivalent parameter and a second equivalent parameter. The output signal-to-noise ratio is used as reward feedback and substituted into the Bellman equation to update the reward table, thereby obtaining the updated resonance parameters.
[0017] The technical solution provided in this disclosure may include the following beneficial effects: In the embodiments of this disclosure, the signal acquisition module, front-end conditioning module, stochastic resonance module, reinforcement learning module, and output module work together to preprocess the noisy signal, perform stochastic resonance enhancement on the preprocessed noisy signal, evaluate the enhanced restored signal using reinforcement learning, and adjust the resonance parameters to obtain the final restored signal. The stochastic resonance module directly maps the resonance parameters to the equivalent parameters of an adjustable digital potentiometer, enabling online closed-loop adjustment of the resonance parameters for deployment on edge computing devices. Furthermore, the collaborative work between the stochastic resonance module and the reinforcement learning module ensures that the system maintains good enhancement capabilities, adaptability, and real-time performance even in low signal-to-noise ratio and dynamic noise environments. Simultaneously, the reinforcement learning module uses the output signal-to-noise ratio as a timely reward for resonant parameter updates, more accurately reflecting the impact of the resonance parameters on signal enhancement quality. This ensures that the optimal resonance parameters obtained through the reinforcement learning module's iterative updates optimize the enhancement effect of the stochastic resonance module on the noisy signal.
[0018] It should be understood that the above general description and the following detailed description are exemplary and explanatory only, and are not intended to limit this disclosure. Attached Figure Description
[0019] The accompanying drawings, which are incorporated in and form part of this specification, illustrate embodiments consistent with this disclosure and, together with the description, serve to explain the principles of this disclosure. It is obvious that the drawings described below are merely some embodiments of this disclosure, and those skilled in the art can obtain other drawings based on these drawings without any inventive effort.
[0020] Figure 1 A schematic diagram of a parameter adaptive stochastic resonance system based on reinforcement learning is shown in an exemplary embodiment of this disclosure; Figure 2 This diagram illustrates a simplified illustration of the system applied in an underwater wireless optical communication scenario, as shown in an exemplary embodiment of this disclosure. Figure 3 A circuit diagram of the random resonance module in an exemplary embodiment of this disclosure is shown; Figure 4 A flowchart illustrating the steps of a parameter adaptive stochastic resonance method based on reinforcement learning in an exemplary embodiment of this disclosure is shown. Figure 5 A schematic diagram illustrating the simulation results of the final restored signal in an exemplary embodiment of this disclosure is shown. Figure 6 A schematic diagram illustrating the simulation results of the Q-SR algorithm performance and statistical stability in an exemplary embodiment of this disclosure is shown. Figure 7 A schematic diagram illustrating simulation results of the bit error rate variation characteristics of four algorithms under different input signal-to-noise ratios in an exemplary embodiment of this disclosure is shown. Figure 8 A schematic diagram of the simulation results of the convergence curve of the PSO-SR algorithm in an exemplary embodiment of this disclosure is shown. Figure 9 A schematic diagram of the simulation results of the convergence curve of the MFPSO-SR algorithm in an exemplary embodiment of this disclosure is shown. Figure 10 A schematic diagram of the simulation results of the convergence curve of the Q-SR algorithm in an exemplary embodiment of this disclosure is shown; Figure 11 A schematic diagram showing simulation results comparing the computation time of the algorithms in an exemplary embodiment of this disclosure is provided. Figure 12 The diagram shows the signal recovery obtained under the UWOC experimental system in an exemplary embodiment of this disclosure. Detailed Implementation
[0021] Exemplary embodiments will now be described more fully with reference to the accompanying drawings. However, these exemplary embodiments can be implemented in many forms and should not be construed as limited to the examples set forth herein; rather, they are provided so that this disclosure will be more comprehensive and complete, and will fully convey the concept of the exemplary embodiments to those skilled in the art. The described features, structures, or characteristics may be combined in any suitable manner in one or more embodiments.
[0022] Furthermore, the accompanying drawings are merely illustrative of this disclosure and are not necessarily drawn to scale. The same reference numerals in the drawings denote the same or similar parts, and therefore repeated descriptions of them will be omitted. Some block diagrams shown in the drawings are functional entities and do not necessarily correspond to physically or logically independent entities. These functional entities may be implemented in software, in one or more hardware modules or integrated circuits, or in different network and / or processor devices and / or microcontroller devices.
[0023] This example implementation first provides a parameter-adaptive stochastic resonance system based on reinforcement learning, referencing... Figure 1 As shown, the system includes a signal acquisition module 110, a front-end conditioning module 120, a stochastic resonance module 130, a reinforcement learning module 140, and an output module 150.
[0024] The signal acquisition module 110 is used to receive noisy signals. These noisy signals are weak signals with added noise.
[0025] The front-end conditioning module 120 is used to preprocess the noisy signal. The preprocessing includes impedance matching, amplitude limiting, and DC component correction.
[0026] The random resonance module 130 is used to enhance the noisy signal output by the front-end conditioning module to obtain the restored signal to be evaluated, and to enhance the noisy signal output by the front-end conditioning module under the optimal resonance parameter settings fed back by the reinforcement learning module 140 to obtain the final restored signal. The enhancement processing is a nonlinear enhancement processing.
[0027] The reinforcement learning module 140 is used to perform statistics on the restored signal to be evaluated, obtain the output signal-to-noise ratio, and update and optimize the resonance parameters based on the output signal-to-noise ratio to obtain the optimal resonance parameters.
[0028] Output module 150 is used to output the final restored signal.
[0029] In the embodiments of this disclosure, the signal acquisition module 110, front-end conditioning module 120, random resonance module 130, reinforcement learning module 140, and output module 150 work together to preprocess the noisy signal, perform random resonance enhancement on the preprocessed noisy signal, evaluate the enhanced restored signal through reinforcement learning, and adjust the resonance parameters to obtain the final restored signal. The random resonance module 130 can directly map the resonance parameters to the equivalent parameters of an adjustable digital potentiometer, enabling online closed-loop adjustment of the resonance parameters for deployment in edge computing devices. Furthermore, the collaborative operation of the random resonance module 130 and reinforcement learning module 140 ensures the system maintains good enhancement capabilities, adaptability, and real-time performance even in low signal-to-noise ratio and dynamic noise environments. Simultaneously, the reinforcement learning module 140 uses the output signal-to-noise ratio as a timely reward for resonant parameter updates, more accurately reflecting the impact of the resonance parameters on signal enhancement quality. This ensures that the optimal resonance parameters obtained by the reinforcement learning module 140 through iterative updates optimize the enhancement effect of the random resonance module 130 on the noisy signal.
[0030] The various parts of the system described above in this example embodiment will now be described in more detail.
[0031] In one embodiment, the signal acquisition module 110 is used to receive a noisy signal, which is a weak noisy signal.
[0032] refer to Figure 2 As shown, the noisy signal can come from the downstream electrical signal of a photodiode, laser diode (LD), avalanche photodiode, silicon photomultiplier (SiPM), piezoelectric sensor, electromagnetic sensor, or other analog detection unit.
[0033] For example, when the method of this application is applied to an underwater wireless optical communication scenario, the noisy signal is a unipolar symbol electrical signal obtained by photoelectric conversion of the attenuated optical signal by the photoelectric receiver.
[0034] In one embodiment, the front-end conditioning module 120 is used to preprocess the noisy signal; wherein, the preprocessing includes: impedance matching processing, amplitude limiting processing and DC component correction processing, so that the signal input to the random resonant circuit is in a preset operating range, and avoids the bistable system operating point mismatch caused by excessively high input amplitude, bias drift or high frequency glitches.
[0035] The preprocessing of the noisy signal by the front-end conditioning module 120 ensures that the preprocessed noisy signal is within the preset working range, thus avoiding mismatch of the operating point of the bistable system due to excessively high input amplitude, bias drift, or high-frequency glitches.
[0036] It should be noted that the noisy signal preprocessed by the front-end conditioning module 120 is used for enhancement processing in the random resonance module 130, and also serves as a monitoring signal for displaying system operating status, input statistical analysis, and comparing the signal enhancement effect before and after parameter updates.
[0037] In one embodiment, the random resonance module 130 is used to enhance the noisy signal output by the front-end conditioning module 120 to obtain the restored signal to be evaluated.
[0038] For example, under the resonance parameter settings during the iteration phase, the random resonance module 130 is used to enhance the noisy signal output by the front-end conditioning module 120, resulting in the restored signal to be evaluated.
[0039] After iterative updates, under the optimal resonance parameter settings, the noisy signal output by the front-end conditioning module 120 is enhanced by the random resonance module 130 to obtain the final restored signal.
[0040] It should be noted that at the initial moment of the iteration, the parameters of the random resonance module 130 need to be initialized. Therefore, at the initial moment of the iteration, the preprocessed noisy signal should be enhanced under the initial resonance parameter settings to obtain the restored signal to be evaluated.
[0041] It should be noted that the resonance parameters are written into the random resonance module 130 so that after the system is started, the random resonance module 130 can be used to enhance the noisy signal output by the front-end conditioning module 120.
[0042] In one embodiment, the potential function of the random resonance module 130 can be expressed as: (5) in, Represents the potential function. Indicates the output state of the stochastic resonance module. , The parameters represent the dynamic equations.
[0043] In one embodiment, the stochastic resonance theoretical model of the stochastic resonance module 130 is as follows: (1) in, Let represent the integral term of the stochastic resonance module, and let x represent the output state of the stochastic resonance module. , Indicates the first equivalent parameter. The parameters representing the dynamic equations, , Indicates the second equivalent parameter. The parameters representing the dynamic equations, Indicates a useful signal and , Indicates noise signal and RC represents the time constant formed by the integrator resistor R and capacitor C in the random resonance module.
[0044] The dynamic equation is: (2) Where a and b represent the parameters of the dynamic equation, s(t) represents the useful signal, h(t) represents the noise signal, and x represents the output state of the stochastic resonance module. express Regarding time The first derivative of .
[0045] It should be understood that the above dynamic equation is the same as the dynamic equation corresponding to the potential function represented by formula (5).
[0046] It should be noted that the parameters , These are key parameters used to determine the shape, location, and height of the potential well. By changing these parameters... , The value of can change the ease with which a particle transitions between the two potential wells in a bistable system, allowing the system to form a matching relationship with the useful signal under different noise backgrounds.
[0047] In one embodiment, reference Figure 3 As shown, the random resonance module 130 includes an integrator (IC1), an inverter (IC2), a multiplier unit, a first digital potentiometer (B1), and a second digital potentiometer (A1) for forming a parameter-adjustable bistable closed-loop circuit.
[0048] An inverter (IC2) is used to form a primary feedback branch together with the second digital potentiometer (A1) and the integrator (IC1); The multiplication unit is used to form a cubic nonlinear feedback term together with the first digital potentiometer (B1) and the integrator (IC1); The integrator (IC1) is used to receive the noisy signal output by the front-end conditioning module 120 and to perform integration operations on the noisy signal, the primary feedback branch signal, and the tertiary nonlinear feedback branch signal.
[0049] The multiplication unit includes a first multiplier (D) and a second multiplier (E). The output of the integrator (IC1) is connected to the input of the first multiplier (D), one end of the second digital potentiometer (A1), and the input of the inverter (IC2). The output of the first multiplier (D) is connected to the input of the second multiplier (E), and the output of the second multiplier (E) is fed back to the input of the integrator (IC1) via the first digital potentiometer (B1) to form a triple nonlinear feedback branch related to parameter b. The output of the inverter (IC2) is fed back to the input of the integrator (IC1) via the second digital potentiometer (A1) to form a primary feedback branch related to parameter a. The noisy signal, the primary feedback branch signal, and the triple nonlinear feedback branch signal output by the front-end conditioning module 120 are superimposed at the input of the integrator (IC1), so that the integrator (IC1), the inverter (IC2), the first multiplier (D), the second multiplier (E), the first digital potentiometer (B1), and the second digital potentiometer (A1) together constitute a parameter-adjustable bistable closed-loop circuit.
[0050] The resonance parameters include a first equivalent parameter and a second equivalent parameter. The first equivalent parameter is written into the first digital potentiometer, and the second equivalent parameter is written into the second digital potentiometer.
[0051] In the above-mentioned random resonance module 130, the first digital potentiometer and the second digital potentiometer can respond to external control signals to change the parameters in the above formula (2). , equivalent value , The equivalent value , These are the equivalent resonant parameters of the bistable closed-loop circuit, also known as the first and second equivalent parameters. By adjusting the resistance of the first and second digital potentiometers, the hardware parameters can be adjusted online, allowing the system to dynamically change with the parameter updates of the reinforcement learning module.
[0052] Optional, see reference Figure 3 As shown, Figure 3The circuit parameters can be taken as follows: R=15kΩ, R1=R2=150kΩ, C=180pF, D=E=1; the voltage division adjustment range of the first digital potentiometer (B1) and the second digital potentiometer (A1) is 0~1.00. This range allows the first and second equivalent parameters to be adjustable in stages from zero feedback to full-scale feedback, covering weak feedback, critical resonance, and strong feedback operating states, and avoiding parameter adjustments exceeding the linear operating range of the integrator, inverter, and multiplier, thereby improving parameter search coverage and hardware closed-loop stability. It should be noted that the above circuit parameters can be adjusted according to the input signal frequency range, noise intensity, and hardware bandwidth requirements; this embodiment does not impose specific limitations on this.
[0053] It should be noted that the system of this application is able to convert the parameters in formula (2) , The modulation is directly converted into hardware control quantities and written to the first and second digital potentiometers, so that the random resonance module 130 can both enhance the noisy signal and respond to the parameter adjustment commands fed back by the reinforcement learning module 140, updating the first equivalent parameter. Second equivalent parameter Then, under the optimized and updated resonance parameters, the random resonance module 130 is used to enhance the preprocessed noisy signal.
[0054] It should also be noted that the reconstructed signal to be evaluated is the signal obtained by the random resonance module 130 enhancing the preprocessed noisy signal under the resonance parameter settings during the iteration phase. This reconstructed signal to be evaluated is not the final reconstructed signal, but rather is sent to the reinforcement learning module 140 for evaluation, so that the reinforcement learning module 140 can evaluate the first equivalent parameter using the reconstructed signal to be evaluated. Second equivalent parameter The optimization process is performed, and through the interaction and cooperation of the stochastic resonance module 130 and the reinforcement learning module 140, the optimized resonance parameters after optimization can be determined and fed back to the stochastic resonance module 130.
[0055] In one embodiment, the reinforcement learning module 140 includes an extraction unit and a computation unit.
[0056] The extraction unit is used to extract the first sample sequence and the second sample sequence from the restored signal to be evaluated.
[0057] The calculation unit is used to perform statistical analysis on the first sample sequence and the second sample sequence to obtain the noise power corresponding to the restored signal, and calculate the output signal-to-noise ratio based on the noise power.
[0058] It should be noted that after the reinforcement learning module 140 receives the restored signal to be evaluated, it needs to sample and cache the restored signal before performing the above statistical analysis through the extraction unit and the calculation unit.
[0059] For example, reinforcement learning module 140 is used to extract a first sample sequence of transmitting "1" and a second sample sequence of transmitting "0" from the restored signal based on known reference symbol positions, training sequences, clock synchronization information or pilot information, in order to characterize the degree of statistical separation between the two types of symbols after enhancement.
[0060] For example, the noise power calculated by the computing unit is: (3) in, Indicates noise power, N1 represents the sample length of the first sample sequence, N0 represents the sample length of the second sample sequence, and I 11 I represents the mean of the first sample sequence. 00 Let represent the mean of the second sample sequence, and i represent the sample number.
[0061] The calculated output signal-to-noise ratio is: (4) in, Indicates the output signal-to-noise ratio. P represents noise power. s Indicates signal power.
[0062] It should be noted that the calculation unit is used to perform statistical analysis on the first sample sequence and the second sample sequence. Specifically, the calculation unit needs to calculate the mean I of the first sample sequence. 11 The mean I of the second sample sequence 00 , and the sample variance terms of the first sample sequence and the second sample sequence relative to their respective means; the sample variance terms are the fluctuations of the two types of samples, and are synthesized into noise power according to formula (3).
[0063] It should also be noted that, in reinforcement learning module 140, compared with using peak amplitude, instantaneous voltage or single-point sample value as the timely reward for parameter updates, this application uses the output signal-to-noise ratio as the timely reward for resonant parameter updates, which can more accurately reflect the comprehensive impact of the current parameter combination on signal enhancement quality, symbol separability and subsequent decision reliability.
[0064] In one embodiment, the optimization unit of the reinforcement learning module 140 is used to discretize the resonance parameters into a two-dimensional index space according to a preset parameter range and a preset quantization step size, and to establish a reward table; the output signal-to-noise ratio is used as reward feedback and substituted into the Bellman equation to update the reward table to obtain the updated resonance parameters.
[0065] The resonance parameters include the first equivalent parameter. (Parameters in the dynamic equation) ) and second equivalent parameter (Parameters in the dynamic equation) ).
[0066] The update process of the reward table is represented as follows: (6) Where Q represents the reward table, Indicates the state at time t. Indicates the action at time t. Indicates the learning rate. This represents the immediate reward (output signal-to-noise ratio) obtained from the evaluation of the restored signal. Indicates the discount factor. This represents the candidate action at time t+1. This represents the state at time t+1.
[0067] It should be noted that by repeatedly updating the reward table according to the Bellman equation, the optimal resonance parameters can be approximated through an interactive process of parameter trial, signal restoration feedback, and reward table Q-value correction without the need to establish an accurate noise model or analyze the optimal solution.
[0068] In one embodiment, the optimization unit of the reinforcement learning module 140 is also used to employ an improved ε-greedy strategy and a variable neighborhood search mechanism to select the index of the current best Q-value parameter in the reward table with a probability of 1-ε, while setting a variable neighborhood radius R centered on the historical best resonance parameter.
[0069] It should be noted that in the early stages of iteration, a larger exploration rate ε and initial neighborhood radius are used to broaden the search range and avoid local optima. In the later stages of iteration, the exploration rate ε is gradually reduced, and the neighborhood radius is reduced from the initial value to the minimum value, completing the switch from coarse search to fine search. Finally, the parameter index with the largest current Q value is selected, and this parameter index is mapped to the actual equivalent parameter. , And based on equivalent parameters , Write the corresponding hardware parameters into the stochastic resonance module.
[0070] The aforementioned optimization unit employs an improved ε-greedy strategy and a variable neighborhood search mechanism, which can effectively balance the globality of parameter optimization with search speed, thereby improving the convergence efficiency of parameter optimization.
[0071] For example, before performing a task, the optimization unit of reinforcement learning module 140 needs to set the range and step size of parameters a and b, and initialize the all-zero reward table.
[0072] Optionally, the parameter space is discretized into a 100×100 table, with the initial values of the exploration rate ε set to 0.7, the learning rate α set to 0.02, the discount factor γ set to 0.1, the initial value of the neighborhood radius set to 12, the minimum neighborhood radius set to 2, the stagnation threshold set to 20, and the maximum number of iterations set to 100. It should be noted that the above parameter settings can be adjusted according to actual needs, and this embodiment does not impose specific limitations on them.
[0073] It should be noted that when the reinforcement learning module 140 completes its current iteration, it outputs the currently selected resonance parameter and maps this resonance parameter to the resistance value of a digital potentiometer, thereby writing it into the random resonance module 130. At this time, the random resonance module 130 is used to adjust its bistable potential well structure and nonlinear feedback strength under the newly written resonance parameter settings, and to continue to enhance the same noisy signal under the new resonance parameter settings.
[0074] It should also be noted that the reinforcement learning module 140 is also used to determine whether the current resonance parameters have reached the optimum. When it is determined that the current resonance parameters have not yet reached the optimum, the reinforcement learning module 140 is also used to perform repeated iterations until the convergence condition is met. Specifically, after each statistical calculation of the signal to be evaluated and restored, the reinforcement learning module 140 updates the reward table according to the current output signal-to-noise ratio and outputs the next set of parameters a, b or the corresponding first equivalent parameter a1 and second equivalent parameter b1; the random resonance module 130 updates the resistance value of the digital potentiometer accordingly, and enhances the noisy signal in the next time window again under the updated parameters to obtain a new signal to be evaluated and restored; the reinforcement learning module 140 recalculates the output signal-to-noise ratio and enters the next iteration until the convergence condition is met or the maximum number of iterations is reached.
[0075] Upon reaching convergence, the reinforcement learning module 140 further feeds back the optimal resonance parameters to the stochastic resonance module 130, enabling the stochastic resonance module 130 to enhance the noisy signal output by the front-end conditioning module 120 under the optimal resonance parameter settings, thereby obtaining the final restored signal. The stochastic resonance module 130 also sends the final restored signal to the output module 150, allowing the output module 150 to display the final restored signal for communication demodulation, bit error rate statistics, spectrum analysis, or other application processing.
[0076] It should be noted that when the reinforcement learning module 140 determines that the closed-loop search has reached the convergence state, the stochastic resonance module 130 is used to solidify the historical best resonance parameters as its current working parameters. At the same time, the reinforcement learning module 140 stops the large-scale search and only retains the necessary periodic verification or local fine-tuning.
[0077] In this application, the optimization unit of the reinforcement learning module 140 can be deployed in a microcontroller, digital signal processor (DSP), field-programmable gate array (FPGA), system-on-a-chip, or industrial control computer. The optimization unit iteratively updates the action value in the discrete parameter space by reading the evaluation result of the output signal-to-noise ratio, and sends the current optimal resonant parameter to the system's parameter writing control unit. Preferably, this parameter writing control unit uses a digital potentiometer, but it can also use a variable resistor array driven by a digital-to-analog converter (DAC), an analog switched resistor network, or other equivalent adjustable impedance structures.
[0078] In this application, the random resonance module 130 is used to enhance noisy signals. This enhancement is nonlinear and selective, rather than a simple proportional amplification. Unlike traditional linear amplifiers that amplify both noise and signal simultaneously, this application optimizes the bistable parameters, namely the first and second equivalent parameters, to transfer a portion of the disordered energy in the background noise to the ordered components of the useful signal. This improves the significance of the output main frequency components, enhances the statistical separation between "1" and "0" symbols, and improves the decision margin of the time-domain waveform.
[0079] In this application, the random resonance module 130 does not merely remain at the level of software simulation or theoretical parameter optimization, but uses an adjustable analog circuit as an enhanced execution unit; the resonance parameters output by the reinforcement learning module 140 are converted by the control module of the parameter writing system, which can actually change the first equivalent parameter and the second equivalent parameter of the random resonance module 130, so that the restored signal output by the hardware circuit can be used as the basis for reward evaluation again in the subsequent learning process.
[0080] In one embodiment, the output module 150 is used to output the final restored signal for communication demodulation, bit error rate statistics, spectrum analysis or other application processing.
[0081] For example, in a communication scenario, the output module 150 is used to extract the optimal sampling point and make a threshold decision on the final restored signal based on the symbol rate, sampling rate and synchronization phase, output the "1 / 0" symbol sequence, and further calculate the bit error rate; in a weak feature detection scenario, the output module 150 is also used to output the enhanced amplitude, spectral peak, energy distribution or classification result.
[0082] This example implementation also provides a parameter adaptive stochastic resonance method based on reinforcement learning, referencing... Figure 4 As shown, the method may include steps S101 to S104.
[0083] Step S101: Receive the noisy signal and preprocess the noisy signal.
[0084] It should be noted that the noisy signal is a weak noisy signal; the preprocessing includes impedance matching, amplitude limiting, and DC component correction.
[0085] Step S102: Under the resonance parameter settings in the iterative stage, the preprocessed noisy signal is enhanced to obtain the restored signal to be evaluated.
[0086] It should be noted that the enhancement process is a nonlinear enhancement process.
[0087] Step S103: Perform statistical processing on the restored signal to be evaluated to calculate the output signal-to-noise ratio, and update the resonance parameters based on the output signal-to-noise ratio to obtain the optimal resonance parameters.
[0088] Step S104: Enhance the preprocessed noisy signal under the optimal resonance parameter settings to obtain the final restored signal, and output the final restored signal.
[0089] In the embodiments of this disclosure, the method of this application performs preprocessing on the noisy signal, random resonance enhancement on the preprocessed noisy signal, reinforcement learning evaluation on the enhanced restored signal, and feedback adjustment of the resonance parameters to obtain the final restored signal. Compared with traditional optimization algorithms, this application uses the output signal-to-noise ratio as a timely reward for updating the resonance parameters, which can more accurately reflect the influence of the resonance parameters on the signal enhancement quality and ensure that the optimal resonance parameters obtained by the final iterative update can achieve the best enhancement effect on the noisy signal.
[0090] The steps of the method described above in this example implementation will now be explained in more detail.
[0091] In one embodiment, in step S101, a noisy signal can be received by a signal acquisition module. This noisy signal is a weak noisy signal, as referenced... Figure 2 As shown, the noisy signal can come from the downstream electrical signal of a photodiode, avalanche photodiode, silicon photomultiplier tube, piezoelectric sensor, electromagnetic sensor or other analog detection unit.
[0092] For example, when the method of this application is applied to an underwater wireless optical communication scenario, the noisy signal is a unipolar symbol electrical signal obtained by photoelectric conversion of the attenuated optical signal by the photoelectric receiver.
[0093] In one embodiment, in step S101, the noisy signal can be preprocessed by the front-end conditioning module; wherein, the preprocessing includes: impedance matching processing, amplitude limiting processing and DC component correction processing, so that the signal input to the random resonant circuit is in a preset working range, and avoids the bistable system operating point mismatch caused by excessive input amplitude, bias drift or high frequency glitches.
[0094] By performing the above preprocessing on the noisy signal, the preprocessed noisy signal is placed within the preset working range, thus avoiding mismatch of the operating point of the bistable system due to excessively high input amplitude, bias drift, or high-frequency glitches.
[0095] It should be noted that the preprocessed noisy signal is used for two purposes: firstly, it is subsequently used for enhancement processing in the stochastic resonance module; secondly, it serves as a monitoring signal for displaying system operating status, input statistical analysis, and comparing the signal enhancement effect before and after parameter updates.
[0096] In one embodiment, in step S102, the preprocessed noisy signal can be enhanced by a random resonance module to obtain the restored signal to be evaluated.
[0097] For example, under the resonance parameter settings in the iteration stage, the preprocessed noisy signal is enhanced by the random resonance module, and the resulting signal is the restored signal to be evaluated.
[0098] After the resonance parameters are optimized and updated, the preprocessed noisy signal is enhanced by the random resonance module under the optimal resonance parameter settings, resulting in the final restored signal.
[0099] It should be noted that the parameters of the stochastic resonance module need to be initialized at the initial moment of the iteration. Therefore, at the initial moment of the iteration, the preprocessed noisy signal should be enhanced under the initial resonance parameter settings to obtain the restored signal to be evaluated.
[0100] It should also be noted that the resonance parameters are written into the random resonance module so that after the system starts, the random resonance module can enhance the preprocessed noisy signal under the resonance parameter settings during the iteration phase.
[0101] In one embodiment, the potential function of the stochastic resonance module can be expressed as: (5) in, Represents the potential function. Indicates the system output status. , The parameters represent the dynamic equations.
[0102] The dynamic equation corresponding to this potential function can be expressed as: (2) Where a and b represent the parameters of the dynamic equation, s(t) represents the useful signal, h(t) represents the noise signal, and x represents the output state of the stochastic resonance module. express.
[0103] It should be noted that the parameters , These are key parameters used to determine the shape, location, and height of the potential well. By changing these parameters... , The value of can change the ease with which a particle transitions between the two potential wells in a bistable system, allowing the system to form a matching relationship with the useful signal under different noise backgrounds.
[0104] In one embodiment, the stochastic resonance module includes an integrator (IC1), an inverter (IC2), a first multiplier (D), a second multiplier (E), and a first digital potentiometer (B1) and a second digital potentiometer (A1) controlled by an algorithm.
[0105] The integrator receives the noisy signal output from the front-end conditioning module, integrates the noisy signal to generate a state variable, and completes the state variable integration. The inverter and integrator together form a bistable closed-loop circuit. A first multiplier and a second multiplier are cascaded between the output and input of the integrator to construct a cubic nonlinear feedback term based on the state variable. A first digital potentiometer is electrically connected to the integrator and the second multiplier, and a second digital potentiometer is electrically connected to the integrator and the inverter, used to respond to external control signals to adjust the resonance parameters of the bistable closed-loop circuit.
[0106] It should be noted that the resonance parameters here include a first equivalent parameter and a second equivalent parameter. The first equivalent parameter is written into the first digital potentiometer, and the second equivalent parameter is written into the second digital potentiometer.
[0107] It should be noted that the first and second digital potentiometers in the random resonance module can respond to external control signals to change the parameters in the above formula (2). , equivalent value , The equivalent value , These are the equivalent resonant parameters of the bistable closed-loop circuit, also known as the first and second equivalent parameters. Thus, the hardware parameters can be adjusted online by controlling the resistance of the first and second digital potentiometers.
[0108] Example, reference Figure 3 As shown, when the stochastic resonance module adopts the above circuit structure, the stochastic resonance theoretical model of the stochastic resonance module satisfies the following equation: (1) in, Let x represent the output state of the stochastic resonance module. , Indicates the first equivalent parameter. The parameters representing the dynamic equations, , Indicates the second equivalent parameter. The parameters representing the dynamic equations, , Indicates a useful signal. express, , Indicates noise signal, RC represents the time constant.
[0109] It should be noted that this application is able to include the parameters in formula (2). , The modulation is directly converted into hardware control quantities and written to the first and second digital potentiometers, so that the stochastic resonance module can both enhance noisy signals and respond to parameter adjustment commands from the subsequent reinforcement learning module to update the equivalent value. , Then, under the optimized and updated resonance parameters, the random resonance module is used to enhance the preprocessed noisy signal.
[0110] It should also be noted that the reconstructed signal to be evaluated is the signal obtained by the random resonance module enhancing the preprocessed noisy signal under the resonance parameter settings during the iteration phase. This reconstructed signal to be evaluated is not the final reconstructed signal, but rather is sent to the reinforcement learning module for evaluation, so that the reinforcement learning module can evaluate the equivalent value using the reconstructed signal to be evaluated. , The optimization process is performed, and through the interaction and cooperation of the stochastic resonance module and the reinforcement learning module, the optimized resonance parameters are determined and fed back to the stochastic resonance module.
[0111] In one embodiment, in step S103, the reinforcement learning module can perform statistical analysis on the restored signal to be evaluated to calculate the output signal-to-noise ratio, and update and optimize the resonance parameters based on the output signal-to-noise ratio to obtain the optimal resonance parameters.
[0112] It should also be noted that after receiving the restored signal to be evaluated, the reinforcement learning module samples, caches, and statistically analyzes the restored signal.
[0113] For example, based on known reference symbol positions, training sequences, clock synchronization information, or pilot information, a first sample sequence of transmitting "1" and a second sample sequence of transmitting "0" are extracted from the restored signal to characterize the degree of statistical separation between the two types of symbols after enhancement.
[0114] In one embodiment, statistical analysis is performed on the restored signal to be evaluated to calculate the output signal-to-noise ratio result, including the following steps: Extract the first sample sequence and the second sample sequence from the restored signal to be evaluated; Statistical analysis is performed on the first and second sample sequences to calculate the noise power corresponding to the restored signal. (3) in, Indicates noise power, N1 represents the sample length of the first sample sequence, N0 represents the sample length of the second sample sequence, and I 11 I represents the mean of the first sample sequence. 00 Let represent the mean of the second sample sequence, and i represent the sample number.
[0115] Calculate the output signal-to-noise ratio based on noise power: (4) in, Indicates the output signal-to-noise ratio. P represents noise power, and Ps represents signal power.
[0116] The statistical analysis of the first and second sample sequences includes: calculating the mean I of the first sample sequence. 11 The mean I of the second sample sequence 00 And the corresponding degree of fluctuation.
[0117] It should be noted that, in the reinforcement learning module, compared with using peak amplitude, instantaneous voltage or single-point sample value as the timely reward for parameter updates, this application uses the output signal-to-noise ratio result as the timely reward for resonant parameter updates, which can more accurately reflect the comprehensive impact of the current parameter combination on signal enhancement quality, symbol separability and subsequent decision reliability.
[0118] In one embodiment, updating and optimizing the resonance parameters based on the output signal-to-noise ratio to obtain the optimal resonance parameters can be achieved in the following way: Based on the preset parameter range and preset quantization step size, the resonance parameters are discretized into a two-dimensional index space, and a reward table is established; the resonance parameters include the first equivalent parameter. (Parameters in the dynamic equation) ) and second equivalent parameter (Parameters in the dynamic equation) ).
[0119] The output signal-to-noise ratio is used as reward feedback and substituted into the Bellman equation to update the reward table, thereby obtaining the updated resonance parameters; the update process is expressed as follows: (6) Where Q represents the reward table, Indicates the state at time t. Indicates the action at time t. Indicates the learning rate. This represents the immediate reward (output signal-to-noise ratio) obtained from the evaluation of the restored signal. Indicates the discount factor. This represents the candidate action at time t+1. This represents the state at time t+1.
[0120] It should be noted that by repeatedly updating the reward table according to the Bellman equation, the optimal resonance parameters can be approximated through an interactive process of parameter trial, signal restoration feedback, and reward table Q-value correction without the need to establish an accurate noise model or analyze the optimal solution.
[0121] Furthermore, the action module adopts an improved ε-greedy strategy and a variable neighborhood search mechanism, selecting the index of the current best Q-value parameter in the reward table with a probability of 1-ε, while setting a variable neighborhood radius R centered on the historical best resonance parameter.
[0122] In the early stages of iteration, a larger exploration rate ε and initial neighborhood radius are used to broaden the search range and avoid local optima. In the later stages of iteration, the exploration rate ε is gradually reduced, and the neighborhood radius is reduced from the initial value to the minimum value, completing the transition from coarse search to fine search. Finally, the parameter index with the largest current Q value is selected, and this parameter index is mapped to the actual equivalent parameter. , And based on equivalent parameters , Write the corresponding hardware parameters into the stochastic resonance module.
[0123] The above method employs an improved ε-greedy strategy and a variable neighborhood search mechanism, which can effectively balance the globality of parameter optimization with search speed and improve the convergence efficiency of parameter optimization.
[0124] For example, before performing optimization, the range and step size of parameters a and b must be set, and the all-zero reward table must be initialized.
[0125] Optionally, the parameter space is discretized into a 100×100 table, with the initial values of the exploration rate ε set to 0.7, the learning rate α set to 0.02, the discount factor γ set to 0.1, the initial value of the neighborhood radius set to 12, the minimum neighborhood radius set to 2, the stagnation threshold set to 20, and the maximum number of iterations set to 100. It should be noted that the above parameter settings can be adjusted according to actual needs, and this embodiment does not impose specific limitations on them.
[0126] It should be noted that when the reinforcement learning module completes its current iteration, it outputs the currently selected resonance parameter, mapping this parameter to the resistance value of a digital potentiometer to be written into the stochastic resonance module. At this point, the stochastic resonance module can adjust its bistable potential well structure and nonlinear feedback strength under the newly written resonance parameter settings, and continue to enhance the same noisy signal under the new resonance parameter settings.
[0127] It should also be noted that if the reinforcement learning module determines that the current resonance parameter has not yet reached its optimal value, it needs to perform repeated iterations until the convergence condition is met.
[0128] Upon reaching convergence, the reinforcement learning module feeds back the optimal resonance parameters to the stochastic resonance module. This allows the stochastic resonance module to enhance the noisy signal output from the front-end conditioning module under the optimal resonance parameter settings, resulting in the final restored signal. The final restored signal is then sent to the output module for display and can be used for communication demodulation, bit error rate statistics, spectrum analysis, or other applications.
[0129] It should be noted that when the reinforcement learning module determines that the closed-loop search has reached the convergence state, the historical best resonance parameters can be fixed as the current working parameters of the stochastic resonance module, and the large-scale search can be stopped, retaining only the necessary periodic verification or local fine-tuning.
[0130] In one embodiment, in step S104, the final restored signal can be output through the output module for communication demodulation, bit error rate statistics, spectrum analysis or other application processing.
[0131] For example, in a communication scenario, the output module can extract the optimal sampling point and make a threshold decision on the final restored signal based on the symbol rate, sampling rate and synchronization phase, output the "1 / 0" symbol sequence, and further calculate the bit error rate; in a weak feature detection scenario, it can also output the enhanced amplitude, spectral peak, energy distribution or classification result.
[0132] The method of this application obtains the final restored signal by preprocessing the noisy signal, performing stochastic resonance enhancement on the preprocessed noisy signal, evaluating the restored signal through reinforcement learning, and adjusting the resonance parameters through feedback. This closed-loop design logic is clear. By directly mapping the resonance parameters to the equivalent parameters of an adjustable digital potentiometer, online closed-loop adjustment of the key parameters of stochastic resonance is realized. Through the synergistic cooperation between the stochastic resonance module and the reinforcement learning module, the system still has good enhancement capability, adaptability, and real-time performance in low signal-to-noise ratio and dynamic noise environments.
[0133] To demonstrate the synergistic effect of the reinforcement learning-based parameter adaptive stochastic resonance system and method proposed in this application on noisy signals, the simulation results are presented below.
[0134] The simulation experiment is set up as follows: The test signal for the simulation experiment under low signal-to-noise ratio (SNR) is a unipolar non-return-to-zero code. Twenty symbols are selected for parameter adaptive optimization. Each symbol sequence is sampled by 50 times interpolation to construct a continuous signal sequence containing 1000 sampling points. Gaussian white noise is introduced to simulate actual channel interference.
[0135] When the input SNR is -6dB, the restored signal reference obtained using the system and method proposed in this application Figure 5 As shown, from Figure 5 As can be seen from this, the random resonance module with optimal resonance parameter settings can effectively handle noisy signals (…). Figure 5 After nonlinear enhancement of the superposition of the original signal and noise signal, the noise interference is effectively suppressed and the characteristics of the restored signal are clearly distinguishable.
[0136] The Q-SR algorithm was simulated in 1dB steps within the input SNR range of [-10, 0] dB. The Q-SR algorithm is the abbreviation for the method used in this application's reinforcement learning module to optimize and update parameters. For each SNR, 20 repeated experiments were independently performed using different random input data. (Refer to...) Figure 6 The figure shows the mean output SNR and the statistical results of the 95% confidence interval. From Figure 6As can be seen, the output SNR decreases with the input SNR and shows a steady downward trend. The confidence intervals of each test point remain within a small range, indicating low overall dispersion and good consistency even as noise intensity continues to deteriorate. Even in extremely low SNR environments, the experimental results based on different input data remain tightly distributed, and fluctuations in the input data have a negligible impact on the performance of the Q-SR algorithm. This fully verifies that the method proposed in this application possesses excellent statistical stability over a wide SNR range, providing reliable performance support for the UWOC weak signal detection system.
[0137] Furthermore, to verify the performance of the Q-SR algorithm, four algorithms—Q-SR, MFPSO-SR, parameter-fixed stochastic resonance (F-SR), and no-stochastic-resonance (NO-SR)—were compared. The specific metric was the bit error rate (BER) under different input signal-to-noise ratios. MFPSO-SR refers to an adaptive stochastic resonance method based on Multi-strategy Fusion Particle Swarm Optimization (MFPSO), which uses MFPSO to iteratively optimize the stochastic resonance parameters before performing stochastic resonance enhancement.
[0138] refer to Figure 7 As shown, the NO-SR algorithm has the worst performance, with a bit error rate consistently around 10%. -2 The above describes the F-SR algorithm, which achieves signal enhancement through stochastic resonance and has a lower overall bit error rate than the NO-SR algorithm. However, its statically fixed parameters make it difficult to adaptively match the noise distribution characteristics under different input SNR scenarios, preventing the system from achieving optimal performance across the entire SNR range. Both the MFPSO-SR and Q-SR algorithms can adaptively adjust parameters, significantly improving the bit error rate. As the input SNR increases, the bit error rate decreases rapidly, approaching 10 when the input SNR is 0dB. -5 While both algorithms have comparable performance in terms of bit error rate, the MFPSO-SR algorithm is based on population iteration and has a long running time. The running time of the Q-SR algorithm is determined only by the number of iteration steps and the time spent updating the Q value, making it more real-time.
[0139] refer to Figures 8 to 10 The diagram shows simulation results for the Q-SR, PSO, and MFPSO algorithms with an input SNR of –6dB. The optimal parameters are shown below. =0.12、 =0.22 Output SNR is 6.385dB. Figure 8 The PSO-SR algorithm shown in the figure achieves the optimal SNR after 44 iterations; Figure 9 The MFPSO-SR algorithm shown in the figure reduces the number of convergences to 26 by improving the particle update strategy; Figure 10 The Q-SR algorithm shown in the figure finds the optimal SNR in 7 iterations, outperforming swarm intelligence algorithms in convergence speed. In terms of computational complexity, swarm intelligence algorithms require iterating and optimizing the positions of individuals in a multi-dimensional space, and their complexity increases exponentially with the scale of underwater channel-related problems. In contrast, the Q-SR algorithm optimizes parameters based on the interaction mechanism of the agent's state and actions, eliminating computational redundancy in the population dimension. Its complexity is polynomial, and its growth rate with increasing problem size is lower, significantly outperforming swarm intelligence algorithms in real-time performance and computational efficiency.
[0140] Furthermore, this simulation experiment also performed 50 repeated tests each on the Q-SR, PSO, and MFPSO algorithms under the MATLAB R2022b environment, based on an AMD Ryzen 7 6800H processor. (Reference) Figure 11 As shown in the figure, the PSO-SR and MFPSO-SR algorithms have a processing time concentrated between 80 and 92 seconds. This level of processing delay cannot adapt to the rapid changes in underwater optical communication channels and will significantly reduce the overall communication performance of the system. In contrast, the method proposed in this application (Q-SR algorithm) only requires 0.79 seconds, which is more than 98% more efficient than the former two.
[0141] Furthermore, this simulation experiment also verified the results by building a UWOC experimental system. The transmitter of the experimental system used an arbitrary waveform generator (DLGOL DG5352, AWG) to read a pseudo-random code with a total order of 214, which drove a green laser diode (LD) with a wavelength of 520nm and a transmission power of 1.46mW to generate an optical signal. The LD beam was attenuated sequentially by attenuators of 50%×10%×1% before being incident on a water tank with a volume of 1m×0.5m×0.25m and an effective transmission distance of 1m for transmission. The receiver uses four silicon photomultiplier tubes (SiPMs) to receive optical signals and uses an oscilloscope (MSO5204B Tektronix) to acquire two signals: one is a noisy signal, which is the superposition of the four received optical signals; the other is a recovered signal, which is the recovered signal output by the random resonance module after processing by the system of this application. The recovered signal is obtained by the random resonance module performing nonlinear enhancement on the noisy signal under the optimal resonance parameter settings.
[0142] The system and method proposed in this application can yield a set of example optimal hardware parameters. =0.04、 =0.16, at which point the symbol rate is 1MHz and the sampling rate is 5MHz. (Reference) Figure 12As shown in the figure, when the input signal SNR is approximately -4.9dB, the system can increase the output signal SNR to approximately 6.23dB and increase the output amplitude from approximately 84mV to approximately 1.679V. At this point, the bit error rate is 3.47 × 10⁻⁶. -3 The experimental results show that the system and method proposed in this application can achieve real-time enhancement of weak signals under low signal-to-noise ratio conditions.
[0143] Of course, in addition to underwater wireless optical communication, the system and method proposed in this application can also be used for acoustic sensing, vibration fault diagnosis, weak electromagnetic signal detection, and other scenarios that require parameter adaptive nonlinear enhancement under strong noise background.
[0144] Regarding the system in the above embodiments, the specific manner in which each unit performs operations has been described in detail in the embodiments related to the method, and will not be elaborated here.
[0145] It should be noted that although several units of the system for executing actions have been mentioned in the detailed description above, this division is not mandatory. In fact, according to embodiments of this disclosure, the features and functions of two or more units described above can be embodied in one unit. Conversely, the features and functions of one unit described above can be further divided and embodied by multiple units. Some or all of the units can be selected to achieve the purpose of this disclosure according to actual needs. Those skilled in the art can understand and implement this without any inventive effort.
[0146] Other embodiments of this disclosure will readily occur to those skilled in the art upon consideration of the specification and practice of the invention disclosed herein. This application is intended to cover any variations, uses, or adaptations of this disclosure that follow the general principles of this disclosure and include common knowledge or customary techniques in the art not disclosed herein. The specification and examples are to be considered exemplary only, and the true scope and spirit of this disclosure are indicated by the appended claims.
Claims
1. A parameter-adaptive stochastic resonance system based on reinforcement learning, characterized in that, include: The signal acquisition module is used to receive noisy signals; A front-end conditioning module is used to preprocess the noisy signal; The random resonance module is used to enhance the noisy signal output by the front-end conditioning module to obtain the restored signal to be evaluated, and to enhance the noisy signal output by the front-end conditioning module under the optimal resonance parameter settings fed back by the reinforcement learning module to obtain the final restored signal. The reinforcement learning module is used to perform statistics on the restored signal to be evaluated, obtain the output signal-to-noise ratio, and update and optimize the resonance parameters based on the output signal-to-noise ratio to obtain the optimal resonance parameters. The output module is used to output the final restored signal.
2. The parameter adaptive stochastic resonance system based on reinforcement learning according to claim 1, characterized in that, The stochastic resonance module includes an integrator, an inverter, a first digital potentiometer, a second digital potentiometer, and a multiplication unit for forming a parameter-adjustable bistable closed-loop circuit. An inverter is used to form a primary feedback branch together with the second digital potentiometer and the integrator; The multiplication unit, together with the first digital potentiometer and the integrator, forms a cubic nonlinear feedback term; The integrator is used to receive the noisy signal output by the front-end conditioning module and to perform integration operations on the noisy signal, the primary feedback branch signal, and the tertiary nonlinear feedback branch signal.
3. The parameter adaptive stochastic resonance system based on reinforcement learning according to claim 2, characterized in that, The stochastic resonance theoretical model of the stochastic resonance module is as follows: (1) in, Let represent the integral term of the stochastic resonance module, and let x represent the output state of the stochastic resonance module. , Indicates the first equivalent parameter. The parameters representing the dynamic equations, , Indicates the second equivalent parameter. The parameters representing the dynamic equations, Indicates a useful signal and , Indicates noise signal and RC represents the time constant formed by the integrator resistor R and capacitor C in the random resonance module.
4. The parameter adaptive stochastic resonance system based on reinforcement learning according to claim 3, characterized in that, The dynamic equation is: (2) Where a and b represent the parameters of the dynamic equation, s(t) represents the useful signal, h(t) represents the noise signal, and x represents the output state of the stochastic resonance module. express Regarding time The first derivative of .
5. The parameter adaptive stochastic resonance system based on reinforcement learning according to claim 2, characterized in that, The multiplication unit includes a first multiplier and a second multiplier, which are cascaded between the output and input of the integrator. The first digital potentiometer is electrically connected to the integrator and the second multiplier, and the second digital potentiometer is electrically connected to the integrator and the inverter.
6. The parameter adaptive stochastic resonance system based on reinforcement learning according to claim 2 or 3, characterized in that, The resonance parameters include a first equivalent parameter and a second equivalent parameter. The first equivalent parameter is written into the first digital potentiometer, and the second equivalent parameter is written into the second digital potentiometer.
7. The parameter adaptive stochastic resonance system based on reinforcement learning according to claim 1, characterized in that, The reinforcement learning module includes: An extraction unit is used to extract a first sample sequence and a second sample sequence from the restored signal to be evaluated. The calculation unit is used to perform statistical analysis on the first sample sequence and the second sample sequence to obtain the noise power corresponding to the restored signal, and calculate the output signal-to-noise ratio based on the noise power.
8. A parameter adaptive stochastic resonance method based on reinforcement learning, characterized in that, include: Receive noisy signals and preprocess the noisy signals; The preprocessed noisy signal is enhanced under the resonance parameter settings during the iterative stage to obtain the restored signal to be evaluated. The restored signal to be evaluated is statistically analyzed to obtain the output signal-to-noise ratio, and the resonance parameters are updated and optimized based on the output signal-to-noise ratio to obtain the optimal resonance parameters; The preprocessed noisy signal is enhanced under the optimal resonance parameter settings to obtain the final restored signal, and the final restored signal is output.
9. The parameter adaptive stochastic resonance method based on reinforcement learning according to claim 8, characterized in that, The process of statistically analyzing the restored signal to be evaluated to obtain the output signal-to-noise ratio includes: Extract the first sample sequence and the second sample sequence from the restored signal to be evaluated; Statistical analysis is performed on the first and second sample sequences to calculate the noise power corresponding to the restored signal. (3) in, Indicates noise power, N1 represents the sample length of the first sample sequence, N0 represents the sample length of the second sample sequence, and I 11 I represents the mean of the first sample sequence. 00 This represents the mean of the second sample sequence, where i represents the sample number; Calculate the output signal-to-noise ratio based on noise power: (4) in, Indicates the output signal-to-noise ratio. P represents noise power. s Indicates signal power.
10. The parameter adaptive stochastic resonance method based on reinforcement learning according to claim 8, characterized in that, The process of updating and optimizing the resonance parameters based on the output signal-to-noise ratio includes: Based on the preset parameter range and preset quantization step size, the resonance parameters are discretized into a two-dimensional index space, and a reward table is established. The resonance parameters include a first equivalent parameter and a second equivalent parameter. The output signal-to-noise ratio is used as reward feedback and substituted into the Bellman equation to update the reward table, thereby obtaining the updated resonance parameters.